A digital twin is a predictive model of a real consumer group, grounded in real human data. It is not a generic chatbot given a persona. Every twin starts from actual consumer behavior and real review language.
Twins are built in three stages.
Fragmented inputs are ingested, normalized across formats and taxonomies, quality-scored, and filtered before anything is generated. Segment granularity is preserved throughout. Three tiers of data feed in:
Both structured and unstructured data matter. Structured means demographics, scaled ratings, purchase channels, and transactional data. Unstructured means open-ended responses, call verbatims, and social comments. Verbatim language carries the most weight, because it is what makes a twin sound and reason like a person rather than a row in a table.
Domain-specific models analyze the transformed data to map decision logic and build profiles across psychographics, demographics, and values. Your data is then enriched against a global repository of more than 300 million pre-established twins using matching from contextual patterns rather than on identifiers.
HumanOS Validate benchmarks outputs against human studies, ground-truth signals, and ongoing production data, then surfaces when recalibration is needed.
The result is a queryable, reusable representation of a customer population, rather than a one-off study commissioned and archived.
Yes, in three specific ways. And no in the way that usually prompts the question.
Your data is never used to train or fine-tune any model. Native AI does not fine-tune underlying models at all, so your inputs cannot alter model behavior for anyone, inside your organization or outside it. Your material stays private, proprietary, and isolated.
Yes. The audience selector lets you check as many saved audiences as you want on a single thread.
You get two layers of output from one question.
Once a simulation renders, you can click a specific response group in the chart and spin up a sub-thread containing only those twins. The conversation is then locked to that cohort, and you can probe them with open-ended follow-ups while the broader audience stays out of it.
So a single workflow moves from measurement to qualitative depth: ask the broad question, see where the audience splits, click into the group that matters, and ask them why.
These are two different operations, and the difference is the whole answer.
| Ask in a thread | Run a simulation |
|---|---|
| Searches the audience's full review corpus, retrieves the most relevant context, and returns one synthesized response. Reads like a well-run focus group summary: themes, motivations, barriers, sentiment, in consumers' own words. | Puts the question to each twin individually across your chosen sample. Every twin answers on its own. Responses are aggregated into a statistical distribution, charted, and exportable as respondent-level rows with Twin IDs. |
Running a quantitative simulation against a statistically robust sample, such as n = 1,000, returns nearly identical distributions and statistical confidence to a run against the full audience, with a shorter processing turnaround.
This is the system's grounding working as designed, not a fault. A twin answers only where its underlying data supports an answer.
A general-purpose model will produce a plausible guess when it lacks information. A digital twin will not. Where the grounding data does not carry what your question needs, that twin is filtered out rather than allowed to invent, which protects the integrity of your distribution.
The percentage surviving these filters is the Answer Rate, and it is diagnostic rather than disappointing. A 20 percent rate says four in five of your customers have not discussed or experienced the topic in their reviews or records. That is information about your category and your audience. Forcing the other 800 to answer would produce numbers with nothing behind them.
A low rate usually means widen the audience or broaden the question.
The Synthetic Output setting controls how tightly a twin's response stays anchored to its grounding data, and how much inference it is permitted to apply.
| Setting | How it behaves | Use it for |
|---|---|---|
| High Fidelity | Holds to the source data, drawing exact quotes and mirroring known distributions. Lowest risk of an unsupported claim. Any answer traces directly back to root data. | Validation against human panels, and any output that must be citable and auditable. |
| Balanced Default | Mixes exact quotes with synthesized phrasing for a natural conversational read. Still firmly grounded, with a measured amount of inference so the twin connects patterns across your data. | Most everyday work. Concept testing, message and claim testing, packaging. |
| High Creativity | Reaches further, paraphrasing and inferring from weaker signals in the data. Deliberately bolder. Answers are harder to trace to a single source record. | Early ideation, trend spotting, and thin-data segments where you need the audience to be interactive at all. |
Let data density decide. Where your grounding data is rich, stay on High Fidelity or Balanced. Where it is sparse and you want forward-looking direction, move toward Balanced or High Creativity.
The trade is not accuracy against invention. Every setting requires evidence. What changes is how strong that evidence must be before a twin will commit to an answer. High Fidelity will tell you it has no information. High Creativity will answer from a hint.
During validation work, running the same questions across all three settings shows which level of inference gives the highest parity against your human benchmark for your particular question set.
Both run against the same twins and the same underlying data, and in both you select the audience you are putting the questions to. Threads additionally accept several audiences at once. The two modes suit different kinds of work.
Survey Runner is an agentic workflow for fielding a complete research instrument. You upload a questionnaire as a .docx file, and the agent programs and executes the study end to end, at any length.
Every item runs in a single pass, which consolidates setup into one upload and makes long instruments practical to field.
A thread is a single conversation with your audience. Everything you ask stays in one context window, so a follow-up builds directly on the answer before it. You start broad, see where the audience divides, and ask why.
Verbatim language surfaces in volume here, giving you the vocabulary consumers use and the reasoning behind their choices. A thread also supports structured exercises, such as conducting a forced-choice comparison between competing concepts or message stimuli.
Clicking a response group in a chart opens a sub-thread containing only those twins, so the conversation narrows to the cohort you picked.
| Survey Runner | Threads |
|---|---|
| A finished instrument fielded end to end in one pass | One continuous conversation, where follow-ups compound |
| Respondent-level export with demographic columns, concept ratings, and verbatims aligned by Twin ID | Verbatim depth on motivation, vocabulary, and barriers |
| Every item answered on its own terms | Qualitative and quantitative work side by side in one thread |
| Scale, at any questionnaire length | Sub-threads into any cohort you spot in a result |
New data ingests on a set cadence, ranging from daily to monthly.
The pipeline supports daily refreshes for fast-moving public signals such as retailer reviews and forum discussion. Bespoke client data integrations, including CRM and loyalty updates, are set on a calendar co-designed with your team.
Rather than relying only on the calendar, twins are continuously checked against human benchmarks and real-world outcomes such as sales data. As real behavior shifts, the gap between twin output and observed truth widens. When that gap passes the agreed tolerance, the validation monitor flags the audience as Watch or Act, so your team learns an audience is going stale from the system rather than from a surprising result.
Think of it as a precision lens that applies itself.
The catalog is the master directory connecting your portfolio — brands, flavors, models, SKUs — to the real customer data that belong to each one. Rather than having twins search all of your feedback for every question, the catalog keeps them focused on the exact product you asked about.
Ask what customers disliked about a specific flavor and the system recognizes that product from your catalog, then restricts the twin to the reviews attached to it. Ask a general question and the scope broadens automatically to the wider category or audience.
Twins are built without storing personally identifiable information.
High-compliance organizations can run simulations with zero PII sharing, keeping sensitive records inside their own environment.
Answer Rate is covered above. Full definitions of relevance and faithfulness, including how to read each one and what to do when either is low, are being finalized with our technical team and will follow shortly.
In outline: relevance asks whether an answer actually addresses the question asked. Faithfulness asks whether every claim in an answer is supported by the retrieved source material.
Coming soon.