Synthetic Users: When to Trust the Simulation
Cicero Campelo, CISSP
August 7, 2026 · 16 min read
Part of our guide to AI for startups.

Table of contents
- What are synthetic users?
- Why a prompted persona is not a synthetic user
- How accurate are synthetic users?
- Which questions can you trust synthetic users to answer?
- Where synthetic users pay off first
- Where synthetic users still fail
- Synthetic users do not replace talking to customers
- What should you settle before feeding customer data into a simulation?
- What to do this week
- Sources
- Frequently asked questions
Synthetic users are AI-generated stand-ins for real people, built to answer research questions the way a specific population would. Ask a thousand of them whether they would switch plans at a higher price and you get a distribution back in minutes, not a fielded study in six weeks.
The category stopped being a curiosity this year. In July 2026, TechCrunch reported that Simile raised $200 million at a $2 billion valuation, five months after a $100 million Series A. Simile came out of the Stanford research this field grew from, work Joon Sung Park did with Percy Liang and Michael Bernstein, his two Stanford advisers and now his co-founders. Park, the company's co-founder and CEO, gave a Sequoia interview that is the most useful thing a founder can read on the subject, mostly because he spends so much of it explaining what simulation cannot do.
That is the framing worth taking. The question is not whether synthetic users work. It is which questions they can answer, and there is a real test for that.
What are synthetic users?
A synthetic user is an AI agent that stands in for a person in research. It holds a background, answers questions, makes choices, and can be run thousands of times in parallel. The goal is not a chatbot that sounds human. The goal is a population whose aggregate answers track what the real population would have said.
The line runs back to 2022, when Park's team built Social Simulacra, a simulated subreddit populated with thousands of what they then called personas rather than agents. He calls it "the precursor to the agent paper that we ended up writing." The project that made the idea legible came in April 2023, and it was called Smallville. Park's team built 25 generative agents and let them live in a small simulated town. They "basically created generative agents that is paired with generative AI model with memory, planning, and reflection to basically create this lived experience of agents living in this small town," he says. The agents woke up, went to work, formed relationships, and did things nobody scripted. One agent who ran a cafe decided to throw a Valentine's Day party, spent the previous day gathering supplies and inviting customers, and the party happened.
Smallville proved believability. It did not prove accuracy, and for a founder those are different products. A believable agent makes a good demo. An accurate agent changes a decision.
The lineage here is older than the models. Thomas Schelling published his segregation model in 1971, using agents so simple that, as Park describes them, an "individual agent was simply red dot or blue dot" that looked at its neighbors each turn and decided whether to move. That was enough to show something real about macro behavior, and Schelling went on to share the 2005 Nobel Memorial Prize in Economic Sciences with Robert Aumann. What changed is the resolution. "But now, we can actually create real agents that replicate the full richness of individuals and run the same kind of simulations," Park says.
Why a prompted persona is not a synthetic user
A prompted persona models what people say. A synthetic user models what people do, and closing the distance between the two is the entire product. The cheapest version of this idea is to tell a model to act like a 34-year-old woman in a coastal city and ask it questions. It is the obvious first thing to try, and it is the version Park's team rejected, and the reason is worth understanding before you buy or build anything.
The problem is the say-do gap. "And a lot of the large language models are trained on attitudinal data. Fundamentally, it is the things that people have said online," Park says. A model trained on what people said is a good model of what people say. It is a much weaker model of what they do, and closing that distance is the entire job of product research.
So the pipeline starts with real humans. Simile works with polling and panel vendors, including a partnership with Gallup, to reach a representative sample and collect data on those people before any agent exists. The interview is the interesting part. Instead of asking about a product, they ask people to "tell me the story of your life," because a life story carries the long-tail detail that generalizes past the original topic. Park describes the objective as training an interviewer to work out "how can you spend the minimum amount of time to get the maximum amount of visibility about this person." On top of that sits behavioral data, including a repository of randomized controlled trials, so the model learns from what people did under real conditions and not only from what they reported.
A founder building in the adjacent space describes the same shape from the other direction. Listen Labs, whose product runs AI market research through interviews with real people, is layering simulation on top of that corpus. "The way we do simulation is essentially you have one person do model really really well and then you scale it up with a thousand people. So, you have a representative sample," says founder and CEO Alfred Wahlforss. Two teams approaching from opposite ends land on the same constraint: model an individual from real data, then scale.
That gives you a clean buying test. If a vendor cannot tell you which real people the simulation is grounded in and how that data was collected, you are not buying synthetic users. You are buying a prompted persona with a dashboard.
How accurate are synthetic users?
The best published number for synthetic users is 85 percent of human self-consistency, and what it is measured against matters more than the figure itself. Park's team built agents for a population of about a thousand US adults, 1,052 in the published study, and tested whether those agents could reproduce the same people's answers on the General Social Survey. The agents matched participants' responses with 85 percent of the accuracy the participants themselves reached when they retook the survey two weeks later.
Read that number carefully, because it inflates easily in retelling. It is not 85 percent correct. The yardstick is human self-consistency, and people are not consistent with themselves. Park names that ceiling directly when asked about the theoretical limit: "if you ask me the same question, I'll actually answer the question slightly differently."
For live customer work the measurement is different again. On quantitative questions they compare the shape of the simulated answers against the real ones using total variation distance, and hold a threshold for decisions. "So, TVD of let's say less than 0.15, we believe is actually quite strong evidence for making decision," Park says.
He is also candid that the field has not settled its standards yet. "I see simulation as a field as akin to developing your day one of inferential statistics," he says, pointing at how long it took for researchers to agree that a p-value below 0.05 counted as evidence strong enough for science. Nobody has agreed on the equivalent line for simulations.
The practical version for a founder: when a vendor quotes an accuracy figure, ask two questions. Accurate against what ground truth, and measured how. A number with no benchmark attached means nothing, in the same way an eval score means nothing without the eval.
Which questions can you trust synthetic users to answer?
Simulations split into two kinds, and the split decides what you can trust one with. Convergent simulations land on a stable shape even when individual agents are wrong. Divergent ones do not.
"One simulation is what I would consider to be simulations that converge. The other categories of simulations are the simulations that diverge," Park says.
Convergent questions have outcomes pulled toward a stable shape whatever the small errors. His example is network structure: simulate a group of people and a hub always forms, the pattern network scientists call a scale-free network. He points at PageRank, where "the core observation of PageRank was doesn't matter how these networks actually get formulated, you actually see some web pages that get exponentially more links that are attached to it." When the pull toward that shape is strong, agent error stops being fatal. Park says you are fine even if errors compound over time, because the pull toward convergence is strong enough that you still learn where everything lands.
Divergent questions do not behave. "It's like your classical questions like was World War I inevitable or was it not?" Park says, or an election: "Will the same person win the election every time?" Run it twice and you can get two different worlds, because every decision inside the run changes the ones after it.
The move is not to avoid divergent questions. It is to change what you take from them. For those, the output is a distribution, not an answer. "So, you imagine you run the simulation 100 times, how many of those times do the results come out to be X?" Park says, and the value is partly in seeing the spread of futures and the mechanism that produced each one, so you can prepare for more than the modal case.
Wahlforss raises the same worry from the engineering side, saying that when agent interactions compound it becomes hard to predict how things interact, which is why his team has not yet put its simulated panel into open debate with itself.
Sorting your own backlog against this takes about ten minutes:
- Convergent, safe to simulate for an answer. Which of forty positioning lines lands with this segment. Whether a benefit ordering changes stated preference. Where in the onboarding copy a specific persona gets confused. Whether a price increase pushes a segment's stated intent past a threshold. Which of three concepts a sub-population prefers, and why.
- Divergent, simulate for a range only. Whether this launch makes you the category leader. How a specific competitor responds. Whether a feature spreads. What the market looks like in three years if you take this positioning.
The failure mode is not running the divergent simulation. It is reporting its top result to your board as a forecast.
Where synthetic users pay off first
The first win synthetic users deliver is volume on the questions you already ask. Park describes what customers see almost immediately: "right now we're very much in the practice of testing five to 10 different ideas a month," and then the obvious follow-up, "But what does it look like for us to test instantly thousands of different ideas across thousands of different sub populations?" Concept testing is where most engagements start because it is the most mechanical, and because the current cadence is so obviously constrained by fielding time rather than by ideas.
The second win is the question nobody currently asks because it has no method. Park's example is a car company launching an electric vehicle. Concept testing tells you whether the EV lands. What it never tells you is the knock-on effect: "But what does that do to the perception of, let's say, non-electric vehicle? Does it change the market perception? Then what does it mean for the rest of the product line?" His summary of the status quo is blunt: "Today, there's no way to test for this." Second-order effects on your existing line are exactly the thing founders discover after launch.
The third is rehearsal. "Some of our customers very routinely actually ask us to simulate their earnings call," Park says, which surprised him at first and then turned out to be common. The startup version is a pricing announcement, a repositioning, or a difficult customer email to your ten largest accounts.
There is also a reach argument that has nothing to do with cost. Online experiments only capture the people who respond to online experiments. Simulation, done on a properly grounded population, is "also much more representative because only certain groups of people will actually respond to the online experiments," Park says. Wahlforss frames the same opportunity as the long tail: simulation is how "you're able to unlock the 99% of use cases where you would never have time to talk to real people."
For an early startup, that long tail is where the value sits. You are not replacing the four customer calls you run each week. You are getting answers to the forty questions you currently resolve by arguing in a meeting. That is the same reason AI prototyping earns its place: it is worth most when it kills bad ideas before they consume a sprint.
Where synthetic users still fail
The most important limitation is one you would not guess: the frontier models are being optimized away from this problem, not toward it.
Park is direct about where the big labs are pointed. Their North Star is something like a superintelligent machine, and "These machines are meant to be rational. And these machines are supposed to be really amazing at technical problems that have an objective answer." People are not that. They carry subjective values, preferences, and taste. The result is a split between raw model capability and human-simulation quality: "we have sort of plateaued with current modeling paradigm, our ability to really simulate humans," he says.
His framing of the fix is the memorable part. Today's models are like the CPU of an intelligence unit, one model trained on rational data and excellent at objective problems. What simulation needs is closer to a GPU, many units each faithfully representing the viewpoints of a different sub-population. "In fact, we want model that's as human as possible," he says, which is the opposite of the direction the frontier is racing in.
For a founder, that has a concrete implication. The usual bet, that the next model release makes your product better for free, does not apply cleanly here. A larger, more rational model may be no better at predicting your customers, and the improvements that matter for simulation will come from grounding data and architecture rather than from the next general-purpose release.
Two smaller limits on synthetic users are worth holding onto. There is a randomness floor, since people genuinely answer the same question differently on different days, so no amount of model progress gets you to perfect prediction. And a simulation can only reflect the data it was grounded in, so ask when that data was collected. A population modeled from interviews collected last year has no way to account for a competitor that launched last month, which is precisely when you most want to ask it something.
Synthetic users do not replace talking to customers
The people building these systems are the ones least likely to claim otherwise, because the real interviews are an input to the product, not a stage it eliminates. Framer co-founder and CEO Jorn van Dijk gives the founder version of it on YC's Root Access, describing it as the one thing he keeps repeating to founders and still has to remind his own team of at a much later stage: talk more to users, "Give yourself exposure to the user," and build a feedback loop with real people using your product.
The sequencing that works looks like this. Real conversations define which questions matter and supply the grounding data. Simulation scales the questions you already know how to ask, especially the long tail. Real people come back for anything expensive or irreversible, and when they do, the method for those conversations is the whole subject of The Mom Test.
It is worth noting where Park's motivation started, because it argues for both halves. He came to this from social computing, where the hard problem is not testing a UI but predicting what happens when millions of people interact inside your design. "The only way we test it today is you basically field test it. You release your prototype, see what happens, and sometimes it actually comes at a real cost," he says. Shipping and watching is a real test. It is also the most expensive one available, and by the time it returns an answer, the harm is already in the world.
One prerequisite sits underneath all of it: you cannot simulate a population you have not defined. If you have not written down which customer you are actually serving, a synthetic panel will give you a confident average of the wrong people, which is worse than no answer at all.
What should you settle before feeding customer data into a simulation?
An agent grounded in one person's interviews is closer to that person's likeness than to a row in a database, and this is where most pilots get careless. The Stanford team flagged it themselves when the 1,000-person study came out, noting reasonable concerns about deepfakes and co-option of individuals' likenesses, and putting guardrails around the agent bank they built.
The pressure shows up fast on the commercial side too. Customers with large first-party datasets want their own data folded in to make the simulation specific to their population. Park describes it as an open conversation about "how can we, in a responsible and ethical way, leverage existing data that is also in-house for our customers," which is the honest way to put it: the question is live, not settled.
Answer these four in writing before any customer data leaves your systems, and put the answers in the contract, not the kickoff deck:
- Consent scope. Did the people whose data grounds these agents consent to being modeled, or only to being surveyed? Those are different permissions, and the second does not imply the first.
- Minimization and re-identification. What is the smallest slice of your data that makes the simulation useful? A life-story interview plus a purchase history is often enough to re-identify someone even with names removed.
- Retention and deletion. If a person withdraws consent, can the derived agent and anything trained on it be deleted, and can the vendor demonstrate it rather than assert it?
- Provenance in your own systems. A simulated response must never land in your warehouse looking like a real customer statement. Label it at write time. Six months later, nobody remembers which panel a number came from.
That last one is the failure I would bet on. It is the same discipline that makes clean data provenance the foundation of any AI feature, and it costs nothing to get right on day one and a great deal to fix afterward.
What to do this week
- Take your open research questions and sort them into convergent (safe for synthetic users) and divergent. Anything where small differences compound into different worlds goes in the divergent pile and never gets reported as a single number.
- Pick one convergent question you have been deciding by argument, such as which positioning line lands with your core segment, and make it the pilot.
- Ask any vendor the two questions that matter: which real people is this grounded in and how was that data collected, and what accuracy benchmark are you quoting against what ground truth. Vague answers to either one are the answer.
- Run one divergent question a hundred times and read the spread instead of the winner. Then write down what would have to be true for each cluster of outcomes.
- Settle the four data questions above in writing before the pilot starts, and label simulated responses in your own systems from the first row.
- Keep a small standing loop of real customer conversations running regardless. The simulation is downstream of it, and it degrades as the grounding data ages.
Synthetic users are a real instrument, not a replacement for judgment. Used on the questions they suit, they turn research from a scheduled project into something you can run inside an ordinary decision. Used on the questions they do not suit, they give you a confident number with nothing behind it. Knowing the difference is the skill, and it is one piece of the broader operating discipline covered in the AI for startups founder guide. Building that discipline across your whole company is what this course teaches: AI Operating System for Startups.
Sources
- Simulating Humans at Scale: Simile's Joon Sung Park (Sequoia Capital), the interview this article distills, on grounding, accuracy thresholds, convergence and divergence, and the model plateau.
- Knowing What Your Customers Want, All the Time: Listen Labs' Alfred Wahlforss (Sequoia Capital), cross-source, on grounding simulation in real interviews and the long tail of research questions.
- Generative Agent Simulations of 1,000 People (arXiv), the study behind the 85 percent figure. Note that this links version 1 deliberately: the paper has since been revised and retitled, and the current version reports interview-only, survey-only and combined agents reaching 83, 82 and 86 percent of the same test-retest benchmark. Stanford HAI's write-up covers the 1,052-participant population and the researchers' own guardrails on likeness.
- Generative Agents: Interactive Simulacra of Human Behavior (arXiv), the April 2023 Smallville paper.
- Synthetic-user startup Simile raises $200M at $2B valuation (TechCrunch) for the round and the valuation, and Index Ventures on the Series B, an investor's account of the round and the enterprise deployments behind it.
- Social Simulacra (arXiv), the 2022 simulated-subreddit paper Park calls the precursor to the generative agents work.
- The Prize in Economic Sciences 2005 (NobelPrize.org), for Thomas Schelling and Robert Aumann, whose agent-based work predates the models by decades.
- CEO of Framer: Why Designers Should Become Founders (YC Root Access), for the case for real user exposure.
Frequently asked questions
What are synthetic users?
Synthetic users are AI agents that stand in for real people in research. Each one holds a background, answers questions, and makes choices, and thousands can be run in parallel so a team gets a distribution of responses in minutes instead of fielding a study over weeks. The serious versions are not prompted personas. They are grounded in data collected from real, representative humans first, then generalized, which is what separates a research instrument from a model repeating what people say online.
How accurate are synthetic users?
A widely cited Stanford study of 1,052 US adults gives the clearest public benchmark, where agents built from interviews reproduced participants' General Social Survey answers with 85 percent of the accuracy those participants achieved when retaking the same survey themselves two weeks later. That is not 85 percent correct. The yardstick is human self-consistency, since people do not perfectly reproduce their own answers either. For live decisions, Simile co-founder and CEO Joon Sung Park says his team measures total variation distance between the simulated and real response distributions and treats a TVD below 0.15 as strong enough evidence to decide on.
Can synthetic users replace talking to real customers?
No, and the people building them say so. Simulations have to be grounded in data collected from real humans, so the real interviews are an input, not a step you skip. The practical split is that real conversations define the questions worth asking and supply the grounding data, while synthetic users scale questions you already know how to ask, especially the long tail you would never spend a week of calls on. For an expensive or irreversible decision, go back to real people.
Which questions should you not ask synthetic users?
Avoid asking a simulation for a single confident answer to a divergent question, meaning one where small differences compound into different outcomes. Joon Sung Park splits simulations into those that converge, where the result is pulled toward a stable shape even if the agents are slightly wrong, and those that diverge, like whether the same candidate wins an election every time. Convergent questions such as which of forty positioning lines lands with a segment tolerate error well. For divergent questions, run the simulation many times and read the spread of outcomes rather than the top result.
Build your AI Operating System
A practical course to grow with AI, build internal tools, and operate safely. Join the waitlist and you'll be first in when the course opens.