Questions your team asked, answered.

How is a digital twin formed, and what data goes into one?

A digital twin is a predictive model of a real consumer group, grounded in real human data. It is not a generic chatbot given a persona. Every twin starts from actual consumer behavior and real review language.

Twins are built in three stages.

1. Transform the data

Fragmented inputs are ingested, normalized across formats and taxonomies, quality-scored, and filtered before anything is generated. Segment granularity is preserved throughout. Three tiers of data feed in:

  • First-party. Your own de-identified data. CRM records, loyalty programs, internal surveys, service transcripts, interviews, and past custom studies.
  • Second-party. Enriched datasets from trusted partners, syndicated panels, and behavioral datasets.
  • Third-party and public signals. Category-level context and natural consumer language. Retailer and category reviews, public forums, open social discourse, and census demographics.

Both structured and unstructured data matter. Structured means demographics, scaled ratings, purchase channels, and transactional data. Unstructured means open-ended responses, call verbatims, and social comments. Verbatim language carries the most weight, because it is what makes a twin sound and reason like a person rather than a row in a table.

2. Model the behavior

Domain-specific models analyze the transformed data to map decision logic and build profiles across psychographics, demographics, and values. Your data is then enriched against a global repository of more than 300 million pre-established twins using matching from contextual patterns rather than on identifiers.

3. Calibrate and validate

HumanOS Validate benchmarks outputs against human studies, ground-truth signals, and ongoing production data, then surfaces when recalibration is needed.

  • Human benchmark. Compare against controlled studies or human panels.
  • Ground truth. Check against observed behavior and known outcomes.
  • Production checks. Monitor live usage, drift, and segment-level changes.
  • Validated twin asset. Benchmarked, monitored, and ready for governed use.

The result is a queryable, reusable representation of a customer population, rather than a one-off study commissioned and archived.

Do digital twins learn from the information put into the system?

Yes, in three specific ways. And no in the way that usually prompts the question.

What does not happen

Your data is never used to train or fine-tune any model. Native AI does not fine-tune underlying models at all, so your inputs cannot alter model behavior for anyone, inside your organization or outside it. Your material stays private, proprietary, and isolated.

What does happen

  • In-context retrieval, per question. When you ask something, the platform retrieves the most relevant reviews, records, and responses and supplies them alongside your question. The twin reasons from that retrieved material. The learning is momentary and scoped to your request.
  • Database refreshes. When new data is added and indexed, twins reflect it immediately, because future questions retrieve the newer records.
  • Thread memory. Inside a single conversation, the platform reconstructs the thread history and supplies it as context, so twins remember what was already discussed. This is why follow-up questions work and why you can narrow from a broad topic to a specific detail without repeating yourself.
Can I ask multiple audiences the same question in one thread?

Yes. The audience selector lets you check as many saved audiences as you want on a single thread.

You get two layers of output from one question.

  • A synthesized readout. Twins across every selected audience process the prompt, and the thread returns one cohesive response covering the overarching takeaways and noting where cohorts differ.
  • Per-audience visualizations. Beneath the text, charts break the response data out by individual audience segment so you can compare them side by side.

Going deeper: sub-threads

Once a simulation renders, you can click a specific response group in the chart and spin up a sub-thread containing only those twins. The conversation is then locked to that cohort, and you can probe them with open-ended follow-ups while the broader audience stays out of it.

So a single workflow moves from measurement to qualitative depth: ask the broad question, see where the audience splits, click into the group that matters, and ask them why.

Asking 1,000 twins versus the entire audience — and when to use each

These are two different operations, and the difference is the whole answer.

Ask in a threadRun a simulation
Searches the audience's full review corpus, retrieves the most relevant context, and returns one synthesized response. Reads like a well-run focus group summary: themes, motivations, barriers, sentiment, in consumers' own words. Puts the question to each twin individually across your chosen sample. Every twin answers on its own. Responses are aggregated into a statistical distribution, charted, and exportable as respondent-level rows with Twin IDs.

Use a thread when

  • You are exploring the why. Motivations, unmet needs, barriers, in unfiltered consumer vocabulary.
  • You are refining phrasing before committing to a large run.
  • You are trend spotting or mapping an under-researched category.

Run a simulation when

  • You need a numerically backed winner across concepts, claims, or packaging.
  • You are benchmarking against a controlled human panel.
  • You want subgroup cuts and crosstabs across cohorts.
  • You need row-level data for downstream systems.

Why a sample rather than the full audience

Running a quantitative simulation against a statistically robust sample, such as n = 1,000, returns nearly identical distributions and statistical confidence to a run against the full audience, with a shorter processing turnaround.

Why do twins drop out of a simulation? I selected 1,000 and got 200 responses back

This is the system's grounding working as designed, not a fault. A twin answers only where its underlying data supports an answer.

A general-purpose model will produce a plausible guess when it lacks information. A digital twin will not. Where the grounding data does not carry what your question needs, that twin is filtered out rather than allowed to invent, which protects the integrity of your distribution.

Three filters remove twins

  • Sentiment compatibility. Your question is classified as positive, negative, or neutral, and twins whose grounding reviews are incompatible are screened out first. Ask what frustrated people about a checkout process and twins with entirely positive histories are excluded, so you hear from customers who actually experienced friction.
  • Unknown and refusal responses. Where a twin's data does not mention your topic, it answers that it lacks context. Those responses are removed in post-processing rather than left to skew the charts as blanks.
  • Strict answer schemas. On a binary question, answers are forced into yes, no, or unknown, and the unknowns are dropped. If only 200 twins hold explicit evidence for a firm answer, the other 800 fall away.

Read the Answer Rate as a finding

The percentage surviving these filters is the Answer Rate, and it is diagnostic rather than disappointing. A 20 percent rate says four in five of your customers have not discussed or experienced the topic in their reviews or records. That is information about your category and your audience. Forcing the other 800 to answer would produce numbers with nothing behind them.

A low rate usually means widen the audience or broaden the question.

High Fidelity, Balanced, High Creativity — what differs and when to use each

The Synthetic Output setting controls how tightly a twin's response stays anchored to its grounding data, and how much inference it is permitted to apply.

SettingHow it behavesUse it for
High FidelityHolds to the source data, drawing exact quotes and mirroring known distributions. Lowest risk of an unsupported claim. Any answer traces directly back to root data.Validation against human panels, and any output that must be citable and auditable.
Balanced
Default
Mixes exact quotes with synthesized phrasing for a natural conversational read. Still firmly grounded, with a measured amount of inference so the twin connects patterns across your data.Most everyday work. Concept testing, message and claim testing, packaging.
High CreativityReaches further, paraphrasing and inferring from weaker signals in the data. Deliberately bolder. Answers are harder to trace to a single source record.Early ideation, trend spotting, and thin-data segments where you need the audience to be interactive at all.

The practical rule

Let data density decide. Where your grounding data is rich, stay on High Fidelity or Balanced. Where it is sparse and you want forward-looking direction, move toward Balanced or High Creativity.

The trade is not accuracy against invention. Every setting requires evidence. What changes is how strong that evidence must be before a twin will commit to an answer. High Fidelity will tell you it has no information. High Creativity will answer from a hint.

During validation work, running the same questions across all three settings shows which level of inference gives the highest parity against your human benchmark for your particular question set.

When should I use Survey Runner instead of a normal simulation?

Both run against the same twins and the same underlying data, and in both you select the audience you are putting the questions to. Threads additionally accept several audiences at once. The two modes suit different kinds of work.

Survey Runner

Survey Runner is an agentic workflow for fielding a complete research instrument. You upload a questionnaire as a .docx file, and the agent programs and executes the study end to end, at any length.

  • Prompt translation. The agent converts traditional survey items into twin-ready prompts, optimizing phrasing while preserving the intent and context of your original question.
  • Study sequencing. The agent programs the full instrument and queues it for execution.
  • Human approval. The agent presents the programmed survey for audit and fields it once you approve.

Every item runs in a single pass, which consolidates setup into one upload and makes long instruments practical to field.

Threads

A thread is a single conversation with your audience. Everything you ask stays in one context window, so a follow-up builds directly on the answer before it. You start broad, see where the audience divides, and ask why.

Verbatim language surfaces in volume here, giving you the vocabulary consumers use and the reasoning behind their choices. A thread also supports structured exercises, such as conducting a forced-choice comparison between competing concepts or message stimuli.

Clicking a response group in a chart opens a sub-thread containing only those twins, so the conversation narrows to the cohort you picked.

What each mode gives you

Survey RunnerThreads
A finished instrument fielded end to end in one passOne continuous conversation, where follow-ups compound
Respondent-level export with demographic columns, concept ratings, and verbatims aligned by Twin IDVerbatim depth on motivation, vocabulary, and barriers
Every item answered on its own termsQualitative and quantitative work side by side in one thread
Scale, at any questionnaire lengthSub-threads into any cohort you spot in a result
How frequently are audiences refreshed?

New data ingests on a set cadence, ranging from daily to monthly.

Refresh cadence is configurable

The pipeline supports daily refreshes for fast-moving public signals such as retailer reviews and forum discussion. Bespoke client data integrations, including CRM and loyalty updates, are set on a calendar co-designed with your team.

The system also tells you when a refresh is due

Rather than relying only on the calendar, twins are continuously checked against human benchmarks and real-world outcomes such as sales data. As real behavior shifts, the gap between twin output and observed truth widens. When that gap passes the agreed tolerance, the validation monitor flags the audience as Watch or Act, so your team learns an audience is going stale from the system rather than from a surprising result.

What is the purpose of the product catalog?

Think of it as a precision lens that applies itself.

The catalog is the master directory connecting your portfolio — brands, flavors, models, SKUs — to the real customer data that belong to each one. Rather than having twins search all of your feedback for every question, the catalog keeps them focused on the exact product you asked about.

How it works

Ask what customers disliked about a specific flavor and the system recognizes that product from your catalog, then restricts the twin to the reviews attached to it. Ask a general question and the scope broadens automatically to the wider category or audience.

Why it matters

  • No cross-contamination. Feedback about one product in your portfolio does not bleed into findings about another.
  • Sharper answers. Narrowing to the right reviews makes synthesized responses and direct quotes far more specific.
  • Less setup. You avoid building and toggling between dozens of micro-audiences, one per flavor or variant. Query one broad audience and let the catalog route each question to the right data.
How is personal information protected?

Twins are built without storing personally identifiable information.

  • De-identified on the way in. Enterprise datasets arrive de-identified, and automated sensitive-data detection runs during transformation to catch anything that slipped through.
  • Matching without identifiers. Enrichment against the global twin repository uses matching from contextual patterns — demographic variables, transactional habits, verbatim language — rather than names, emails, or hashed identifiers.
  • No fine-tuning on your data. Because the platform uses in-context retrieval rather than fine-tuning, your inputs never enter any model.
  • Infrastructure. SOC 2 compliant AWS environment, multi-factor authentication, role-based access controls, encryption in transit and at rest, and separated development, staging, and production environments.
  • Regulatory alignment. Architected to align with GDPR, CCPA, and SOC 2.

High-compliance organizations can run simulations with zero PII sharing, keeping sensitive records inside their own environment.

What do relevance and faithfulness measure?Coming

Answer Rate is covered above. Full definitions of relevance and faithfulness, including how to read each one and what to do when either is low, are being finalized with our technical team and will follow shortly.

In outline: relevance asks whether an answer actually addresses the question asked. Faithfulness asks whether every claim in an answer is supported by the retrieved source material.

What does the Import Feedback tab do?Coming

Coming soon.