Ask. Never guess.Introducing Digital Twins →
GuidesCustomer researchCustomer digital twin

Are digital twins reliable? A practitioner’s guide to trusting AI customers


Can you trust a digital twin? It depends almost entirely on two things: what the twin was built from, and what you’re asking it to do. A twin built from a language model’s general knowledge of “enterprise buyers” will give you fluent, agreeable answers that researchers have repeatedly caught contradicting what real customers do. A twin built from your own calls, tickets, and interviews—and asked about the things those sources actually cover—can be trusted in the way you’d trust a well-organized research repository, with the added ability to question it directly.

That distinction gets lost in most discussions of “synthetic customers,” which treat every AI simulation of a customer as the same technology with the same failure modes. They’re not. This guide covers what reliability actually means for a customer digital twin, what the published research shows, the specific ways twins fail, where their limits sit, and when a real interview is still the only right answer.

Can you trust a digital twin?

A digital twin is reliable when three conditions hold. It draws on evidence that genuinely covers the question, it can show you which pieces of evidence it drew on, and you’re asking it for direction rather than a precise number. Remove any one of those and trust should drop accordingly.

Most skepticism about AI customers targets the version where none of those conditions hold: a model prompted to “act like a customer” and asked to predict how people will respond to something new. That skepticism is well founded, and the research below backs it up. But it doesn’t follow that every customer simulation is unreliable, any more than a bad survey proves surveys don’t work. The useful question isn’t whether AI customers can be trusted in general. It’s whether a specific twin, built a specific way, can be trusted for a specific decision.

What “accurate” means for a digital twin

“Accuracy” covers several different claims, and vendors often blur them. Before judging whether a twin is reliable, it helps to separate the kind of accuracy you actually need from the kind you’re being promised.

Type of accuracyWhat it meansWhat to expect from an evidence-grounded twin
Exact matchReproducing the precise figure a real survey would return—say, 62% preferring option ADon’t expect it. A twin summarizing qualitative evidence isn’t a statistical panel and shouldn’t output percentages as if it were one
Directional correctnessGetting the ranking or the reasoning right—which option customers would favor, and whyReasonable, when the attached evidence covers the question. This is what most product and research decisions need
TraceabilityBeing able to verify the answer against the source it came fromRequired. An answer you can check against the original call, ticket, or transcript is worth more than a precise-sounding one you can’t

Most published accuracy figures for synthetic customers measure correlation with real survey results across many questions, which is closest to directional correctness. That’s a meaningful benchmark, but it says nothing about whether any individual answer is right—which is why traceability matters more in practice. A twin that’s directionally right 85% of the time is still wrong on the question you happen to be asking roughly one time in seven, and without citations you can’t tell which time that is.

What the research actually shows

The strongest skeptical case comes from Nielsen Norman Group’s 2024 study of synthetic users, which Radical Product and many others cite. NN/g (Nielsen Norman Group) researchers Maria Rosala and Kate Moran compared answers from a synthetic-user product and ChatGPT against three real interview studies. In one, real participants had completed only some of the online courses they’d started, citing work and competing priorities, while the synthetic users claimed to have finished all of them. Synthetic users also described actively participating in discussion forums that real participants rarely used, and gave equal weight to every need rather than separating the critical from the nice-to-have. NN/g’s conclusion was that synthetic-user findings are hypotheses to test, not evidence.

Two other lines of research explain why. Bisbee and colleagues, writing in Political Analysis, found that ChatGPT prompted with demographic personas could approximate average survey scores but produced far less variation than real respondents—the model reproduced the center of the distribution and lost the spread. And Anthropic researchers found that five leading AI assistants consistently behaved sycophantically, tailoring answers toward what the user seemed to want, a tendency the authors traced to training on human preference judgments.

The research also points to what fixes the problem. A Stanford-led study of 1,052 people built agents from each participant’s own interview, survey answers, or both, then tested how well the agents predicted that person’s responses to the General Social Survey, personality tests, and economic games. Interview-grounded agents reached 83% of the consistency participants showed with their own answers two weeks later, and combined agents reached 86%. Agents given demographics alone reached 74%, and showed larger accuracy gaps across racial and ideological groups.

The pattern across all four studies is consistent. Simulations built from a model’s general idea of a type of person fail in predictable ways, and grounding the simulation in real people’s own words measurably narrows the gap. That’s the case for building twins from first-party customer evidence rather than generating them.

Five ways digital twins fail

Knowing the specific failure modes makes a twin easier to test and its answers easier to read. Each of these applies to generic synthetic customers by default, and each can still affect a grounded twin that’s scoped or instructed poorly.

Sycophancy

A twin that agrees with your concept is the easiest kind to believe and the one you should check hardest. Language models lean toward the answer the person asking seems to want, so a question framed as “Would enterprise admins find this feature valuable?” invites a yes. Guard against it by phrasing questions neutrally, asking the twin for objections before benefits, and writing instructions that tell it to push back when the evidence doesn’t support an idea—Dovetail’s Digital Twins documentation recommends exactly this.

The average trap

A twin that represents everyone represents no one. Blend churned SMB customers, enterprise champions, and prospects into one twin and its answers regress to a composite that no real customer would recognize—the same loss of variance Bisbee’s team measured. Dovetail’s documentation warns that blending segments averages away the differences a twin exists to surface, and recommends scoping each twin to one account, segment, persona, or behavior instead.

Idealized behavior

Generic synthetic customers describe what a conscientious person would do: finish the course, read the release notes, compare three vendors before buying. Real customers do less, later, and for messier reasons. A grounded twin is less exposed because it draws on what customers actually reported, but it still inherits the gap between what people say and what they do. Interview and survey evidence captures stated behavior, so pair a twin’s answers with product usage data before treating them as a description of actual behavior.

Filling gaps with plausible answers

When a generic model lacks information, it produces the most plausible continuation—which is how synthetic users end up inventing quotes and experiences. A well-built twin should do the opposite: say the evidence doesn’t cover the question. Dovetail’s documentation is explicit that a Digital Twin can’t answer questions nobody in its attached sources has discussed. A twin that hedges on a question isn’t broken; it’s telling you where your evidence is thin, and pointing at where the next study should go.

Answering from stale evidence

A twin is only as current as the sources attached to it. If its newest interviews predate a repositioning, a pricing change, or a major release, it will answer confidently about a product that no longer exists. Dovetail recommends reviewing a twin’s sources quarterly, adding new interview rounds as they finish, and detaching sources that no longer reflect the segment.

Digital twins vs. synthetic users

The terms get used interchangeably, but the construction differs, and the construction determines reliability. For a fuller comparison that also covers traditional and AI personas, see customer digital twin vs. AI persona vs. synthetic customer.

DimensionEvidence-grounded digital twinGeneric synthetic user
Built fromYour own calls, tickets, interviews, surveys, and researchA language model’s training data, shaped by a demographic or persona prompt
RepresentsA defined account, segment, persona, or behavior you chooseA generic type of person
When it lacks evidenceShould say so, and point to the gapProduces a plausible answer anyway
Verifying an answerCite the source call, ticket, or transcript and check itNo source exists to check—the answer is the only output
FreshnessUpdates as new evidence reaches its attached sourcesReflects the model’s training data until the model changes
Documented weaknessesInherits gaps and biases in your data; summarizes rather than quotesSycophancy, low variance, idealized behavior (NN/g; Bisbee et al.)
Best used forReusing past research, pressure-testing ideas against known objections, answering recurring stakeholder questionsRehearsing an interview guide, brainstorming questions for a new market

Generic synthetic users aren’t useless. NN/g itself suggests they can help teams get oriented in an unfamiliar domain or pilot an interview guide before real sessions. The risk comes from treating their output as evidence about your customers, when it’s really evidence about how customers are usually described on the internet.

How Dovetail grounds Digital Twins in real customer evidence

In Dovetail, a Digital Twin is a type of AI Agent that stays in character as one customer account, segment, or persona and answers questions in Dovetail Chat. Several design choices address the failure modes above directly.

  • First-party evidence only. A Digital Twin draws only on customer data already connected to your Dovetail workspace—sales calls, support tickets, interviews, surveys, and feedback—not public web data or general model knowledge. You attach it to specific Channels (streams of high-volume feedback like tickets, app reviews, and NPS or CSAT comments), Projects (collections of interviews, transcripts, and highlights), Docs, and Folders.
  • Citations you can check. Answers include traceable citations, and the twin will show the specific quote, transcript, or ticket it drew on when asked. Because a twin summarizes rather than quoting verbatim, Dovetail’s own guidance is to open the source before acting on a high-stakes answer.
  • No retraining lag. When new evidence reaches a twin’s attached sources, its knowledge updates automatically.
  • Permissions carried through. A twin can only reach data the person using it already has permission to view. Managers and Contributors can create and edit twins; Viewers can chat with them but can’t change what they’re grounded in.

That combination is what makes the trust question answerable at all. A generic synthetic user asks you to take its answer on faith because there’s nothing underneath it to inspect. A grounded twin gives you a claim and the evidence behind it, and leaves the judgment where it belongs—with the person making the decision.

Limitations of digital twins

Grounding fixes the worst failure modes of synthetic customers, but it introduces its own constraints. These are the ones worth planning around.

Data dependency. A twin can’t know more than its sources. A segment represented by four interviews and a handful of tickets produces a twin that sounds as confident as one built on four hundred, and a vocal minority in thin evidence can pass for consensus. Dovetail won’t build a Digital Twin without connected customer data, and that’s the right constraint—but it means a new market, a new segment, or a new product line starts with no reliable twin at all.

Recency bias. The flip side of continuous updates is that a twin’s answers reflect whatever evidence dominates its sources right now. A spike in support tickets after a bad release can make a segment twin sound more frustrated than the segment is over a full year. Check the date range behind important answers, and weigh a burst of recent complaints against the longer record.

Edge-case blindness. Twins reflect the customers you’ve already heard from. Accessibility needs, rare-but-costly workflows, and customers who churned silently without ever filing a ticket are underrepresented in most evidence, so they’re underrepresented in the twin. A twin can tell you nothing about the customers whose perspective was never captured.

Privacy constraints. A twin concentrates customer evidence into an interface more people can query, so the consent and retention terms governing that evidence still apply. An account-level twin built on a single customer’s calls carries more privacy weight than a segment twin built on hundreds of anonymized survey responses. Scope what each twin can access to the purpose it serves, and review access the same way you would for the underlying project or channel.

When real interviews still win

A twin makes existing research reusable; it doesn’t replace the act of learning something new. Some questions still need a person on the other end.

  • Unknown unknowns. A twin can only surface themes already present in its evidence. The problem nobody has mentioned yet—the workaround customers built because they assumed you’d never fix it—shows up only when a researcher asks an open question and follows an unexpected answer.
  • Emotional context. Hesitation, frustration, and the pause before an honest answer carry meaning that a summary of a transcript flattens. Decisions involving trust, loss, or organizational politics—why a champion stopped defending your product internally, for instance—need that context firsthand.
  • Early-stage discovery. For a new market, a new persona, or a genuinely novel concept, there’s no evidence base to ground a twin in. Interviews build that base; a twin can only draw on it afterward.
  • Pricing and willingness to pay. Stated preferences diverge from purchasing behavior, and neither a twin nor a synthetic panel can close that gap. Price tests and real purchase data answer that question.

The practical pattern is to use a twin before and between studies—to screen ideas, sharpen the questions worth asking, and answer the recurring questions past research has already covered—then spend real customer time on the questions only customers can answer. For examples of that split across Product, Sales, Marketing, and Customer Success, see nine customer digital twin use cases for B2B teams.

How to test whether a digital twin is reliable

Reliability is a property you check, not one you assume. Before letting a twin inform real decisions, run it through a short validation pass.

  1. Ask questions you already know the answers to. Pull three or four findings from a recent study and ask the twin about them without leading. If it contradicts evidence you’ve verified, fix its scope or sources before trusting it on open questions.
  2. Ask for sources on every answer that matters. Open the cited call, ticket, or transcript and confirm it says what the twin claims. A twin that can’t point to a source for a confident claim is guessing.
  3. Try to make it agree with a bad idea. Pitch a concept your evidence clearly argues against. A twin that endorses it is being sycophantic, and its instructions need to tell it to push back.
  4. Ask about something outside its evidence. A reliable twin says it doesn’t know. One that answers anyway will fill gaps elsewhere too.
  5. Compare segments. Ask the same question of two narrowly scoped twins. If their answers are indistinguishable, the scope is too broad or the evidence too thin to separate them.

Repeat the pass whenever you attach new sources or rewrite the twin’s instructions. For the full build process, including how to scope and instruct a twin from the start, see the complete guide to customer digital twins.

Build digital twins you can check

The honest answer to “are digital twins reliable?” is that reliability belongs to specific twins, not the category. A twin generated from a model’s idea of your customer inherits every failure mode researchers have documented. A twin grounded in your own customer evidence, scoped to one segment, instructed to push back, and cited down to the source transcript gives you something closer to a research repository you can question—useful for direction, verifiable when it matters, and clear about what it doesn’t know.

Dovetail’s Digital Twins are built from the calls, tickets, surveys, and research already in your workspace, and every answer traces back to the evidence behind it.

FAQs

Are synthetic customers reliable?

Synthetic customers generated from a language model’s general knowledge are reliable for rough direction at best. Nielsen Norman Group’s tests found them overly positive, shallow, and wrong about real behavior, and political-science research found their answers far less varied than real people’s. Reliability improves sharply when a simulation is grounded in real people’s own words: a Stanford-led study found interview-grounded agents reached 83% of participants’ own test-retest consistency, against 74% for agents given demographics alone. Treat any synthetic output as a hypothesis unless you can trace it back to real customer evidence.

Do digital twins replace customer research?

No. A digital twin can only reflect evidence your team has already collected, so it’s strongest at making past research reusable—answering the question a stakeholder asks for the third time, or pressure-testing a concept against objections customers have already raised. New research is still the only way to learn about problems nobody has mentioned yet, emotionally charged decisions, and customers whose perspective was never captured. Twins make research go further; they don’t produce new research.

Can a customer digital twin predict behavior?

Not in the forecasting sense. A twin grounded in real evidence can tell you what a segment has said, done, and objected to, and which of two options is more consistent with that record—that’s directional decision support. It can’t produce a conversion rate, forecast willingness to pay, or tell you with confidence how people will react to something genuinely new. Dovetail’s own documentation says Digital Twins don’t produce statistical forecasts and can’t predict future behavior.

Can ChatGPT create a customer digital twin?

ChatGPT can role-play a customer type or reason over data you paste into a conversation, which is useful for rehearsal. A reliable digital twin needs more than a prompt: a connected, permissioned body of first-party customer evidence, a defined scope, updates as new evidence arrives, and answers that cite the specific call, ticket, or interview behind them. Without those, you’re getting a plausible characterization of a customer, not a representation of yours.

Do customer digital twins replace personas?

Usually not. A persona is a static summary that gives teams shared language about who they serve. A digital twin adds a way to ask new questions of the current evidence behind that persona and get answers you can check. Many teams keep the persona for alignment and use the twin when they need an answer to a specific question.

Is an AI persona a digital twin?

Not by default. An AI persona is usually a prompted character drawing on a model’s training data, which is why it tends toward the averaged, agreeable answers researchers have documented. A customer digital twin represents a defined real-world counterpart, draws only on that counterpart’s evidence, updates as the evidence changes, and can show you the source behind each answer. If you can’t check where an answer came from, it isn’t functioning as a twin.

Editor's picks↘

Latest articles↘

Turn customer feedback into product innovation