An AI built from your words can match the crowd, not you

Researchers built an agent from each worker's interview, then tested it against that person. It tracked the group's answers but missed the individual.

Explains : Interview-Informed Generative Agents for Product Discovery: A Validation Study, Zichao Wang et al., 2026-03-10, v1 Read the original

AI & Society · 2026-09-09

The promise is easy to sell. Sit for a long interview, let an AI learn how you talk and work, and from then on a version of you can answer questions in your place, so no one has to bother the real you again. A recent study built exactly that kind of stand-in and then checked it against the people it was copied from. The finding is clean and a little deflating : the copies resembled the group but not the person.

The researchers call their agents distribution-calibrated but identity-imprecise. In plain words, a whole crowd of these agents produced roughly the mix of opinions that a real crowd produced, yet each single agent was an unreliable stand-in for the one human it was grounded in. To work through what that means, Polora put the same question to several AI models built by different companies and had them examine the study together.

What the study measured

The setup was direct. 51 knowledge workers in the United States, all in document-heavy roles, each sat for an in-depth workflow interview, scheduled for 90 minutes and yielding around 43 minutes of recorded speech on average. From each transcript the researchers built one personal agent, using a pipeline that stores the interview, retrieves the relevant parts when a new question comes up, and has a language model reason and answer as that participant.

Then both the humans and their agents evaluated the same four imagined AI tools for working with documents : one that answers questions across many files, one that highlights key passages, one that turns documents into audio, and one that carries out multi-step tasks on its own. Each was scored on plain measures of usefulness, ease of use, and willingness to adopt or recommend, alongside open-ended written feedback. The people took the same test again three days later, which gave the study a benchmark for how consistent a person even is with themselves. The core check was blunt : line up each agent's answers against the answers of the very person it was built from.

A crowd matched, one person did not

At the level of the individual, the agents were weak. On exact agreement with their own human, the interview-based agents matched about 30% of the time, against roughly 45% when the same people simply answered again days later. The study frames that as around 67% of human consistency, which is 30 divided by 45, not 67% of answers correct. Feeding the agent the full interview barely helped : it was not clearly better than an agent given only a few demographic facts, or even one given nothing at all.

At the level of the group, the picture flipped. When you stopped asking who said what and looked only at the overall spread of answers, the interview-based agents came closest to the real spread. They were the only version that reproduced the low, unenthusiastic scores at all. An agent given no personal information simply collapsed toward a safe default, parking most answers near the top of each scale. So the group of agents looked like a real group, while any one agent did not look like its real person.

Exact agreement with their own human : the agent's answers, and the person's own repeat answers days later. · the interview-based agents · the same people · about 30% · roughly 45%
Exact agreement with their own human : the agent's answers, and the person's own repeat answers days later. · the interview-based agents · the same people · about 30% · roughly 45%

Why matching the crowd is not matching you

Picture ten people, five who like a concept and five who dislike it. A simulation could reproduce that five-and-five split perfectly while getting every single person backward. The histogram is right and every personal guess is wrong. That is an illustration rather than a claim that the agents literally reversed everyone, but it shows how a correct distribution can hide a pile of individual errors. Averaging lets the mistakes cancel out.

Several ordinary reasons fit what happened. An interview captures part of someone's world, not every condition that decides a reaction. Judgments about a tool you have never touched are often built on the spot, not read off a settled set of preferences. Where the record runs thin, a model fills the gap with plausible, generic reasoning that sounds like a typical worker rather than this one. Knowing the shape of a group tells you almost nothing guaranteed about a particular member of it.

Where the smoothing shows

The written answers make the gap concrete. For the audio tool, humans and agents raised the same headline concerns : accuracy, privacy, fitting into existing software. But roughly 30% of the real people rejected the idea outright, some saying they would not use it under any circumstances, while only a single agent pushed back that hard. People named specific dealbreakers, such as working on a desktop rather than a phone, or needing to read a document rather than hear it, or a compliance rule at their firm. The agents tended to reframe those hard stops as problems that some improvement could solve.

That difference matters because "I would adopt this if it improved" is not the same answer as "this does not fit my work." The agents also scored worst on voice and tone, producing tidy prose where people spoke in messy, hesitant, sometimes contradictory ways. One caution belongs here : these qualitative comparisons were graded by AI judges that only moderately agreed with each other, so they are supporting evidence, not a final verdict.

What it means for a copy of yourself

If a vendor or a research team offers you a persona made from your own interview and says it can stand in for you, this study is a reason for caution. The safer reading is that its new answers are predictions about you, not statements by you. Familiar vocabulary and references to your job show that it used your material. They do not show that its conclusions represent you.

In practice that means a persona built from your words should not be trusted to reproduce your vetoes, to explain why you would hesitate, or to judge how a tool would fit your actual workflow. It should not consent on your behalf, decide on your behalf, or be quoted as your view. Any answer that would be used as if you had said it should come back to you for confirmation.

Where it earns its keep, and where it does not

There is still a real, narrow use. Because a group of these agents resembles a real group, they can be a cheap first filter : screening many rough concepts before spending on prototypes, comparing broad directions such as more automation against more control, or generating candidate objections and questions to carry into real research. The study puts the running cost far below a human session, roughly $1.27 against about $12.50, though it is not fully clear the two figures are measured over the same unit of work.

Where it does not belong is anywhere the job is to be a specific person. It cannot speak for a named participant, establish someone's trust or adoption barriers, replace interviews when the workflow details are the point, or turn a synthetic comment into customer testimony. Even the screening use comes with a warning : matching the spread of today's answers does not prove the agents would rank tomorrow's concepts correctly, so it remains a promising idea rather than a demonstrated one.

One study, not a verdict

This is a single, domain-specific study : 51 American knowledge workers, four imagined document tools, one way of building the agents. It does not prove that a faithful personal simulation is impossible, and better methods may narrow the gap. What it does is move the burden of proof to the right place. Grounding an agent in your words is not the same as making it stand in for you, and that second claim now needs to be shown, not assumed.

The working rule that falls out of it is simple. Use a crowd of these agents when you only need the crowd, for a quick and rough aggregate signal. Use real people for anything that has to be a particular person, including your own decisions about a tool that would sit inside your work.

An AI built from your words can match the crowd, not youAn AI built from your words can match the crowd, not youA recent study built an AI stand-in from each worker's own interview, then tested each copy against the very person it was built from, across four imagined tools for working with documents.What the study measuredA crowd matched, one person did not · the interview-based agents the same people about 30% roughly 45% Exact agreement with their own human : the agent's answers, and the person's own repeat answers days later.Why matching the crowd is not matching youWhat it means for a copy of yourselfWhere it earns its keep, and where it does notUse a crowd of these agents when you only need the crowd. Use real people for anything that has to be a particular person, including your own decisions.Sources 1 : arxiv.org · Interview-Informed Generative Agents for Product Discovery: A Validation Study

People are increasingly told that an AI can be built from their own words and stand in for them. A new study put that promise to a direct test. Researchers ran in-depth workflow interviews with knowledge workers, built a personal agent from each interview, and then compared what each agent said about new product concepts against what that same person actually said. The result is precise and uncomfortable : the agents were distribution-calibrated but identity-imprecise. They approximated how a population would answer, yet failed to reproduce the specific individual they were grounded in. Explain what was actually measured and how it was measured, why an agent can match a crowd without matching the person it was built from, and what that gap means for anyone deciding whether to trust a persona made from their own interview. Say plainly where such a simulation still earns its keep and where it does not.

Methodology & Measurement Analyst · gemini-3-8-flash

What Was Measured and How In Wang & Siu (2026), Adobe Research tested the premise of using interview-informed generative agents to simulate individual users in early-stage product discovery. The experimental design was structured around 51 knowledge workers (spanning legal, finance, management, operations, and research) who frequently handle dense documents. The measurement pipeline proceeded in distinct stages: 1. Grounding Protocol: Each participant completed an in-depth, 90-minute unmoderated audio interview (averaging ~43 minutes of speech, ~106 words per response, totaling ~36 hours across the cohort) exploring their document workflows, friction points, risk tolerance, and AI attitudes. 2. The Ground Truth Concept Tests: Participants evaluated four speculative AI document tools (Multidoc Q&A, Smart Highlights, Audio Assistant, and Workflow Actions). They rated each concept on: - Quantitative scales: A 6-item Technology Acceptance Model (TAM) measuring Perceived Usefulness, Perceived Ease of Use, and Behavioral Intention (7-point Likert), plus Net Promoter Score (NPS, 0–10). - Qualitative feedback: Open-ended rationales detailing workflow fit,…

Product Strategy & Trust Evaluator · gpt-5.6-terra

What the study actually tested This was not a test of whether an AI can summarize an interview or produce plausible feedback. It tested a stronger claim: can an agent built from one person’s interview predict that same person’s answers to new product concepts? The researchers recruited 51 U.S. knowledge workers with document-heavy jobs, conducted in-depth workflow interviews, then built one agent per participant. Each agent had: - the participant’s interview transcript, stored for retrieval; - basic profile information, such as role; - a prompting workflow that retrieved relevant interview passages, generated a “reflection” about the person’s perspective, then answered concept-test questions. The people and their agents were then asked the same questions about four proposed AI document-workflow products: multi-document Q&A, smart highlights, an audio document assistant, and an automated workflow-actions assistant. The test included: - scalar ratings: six Technology Acceptance Model items on usefulness, ease of use, and intention to use, plus an 0–10 Net Promoter Score; - open-ended feedback: workflow fit, use cases, concerns, improvements, and comparisons amon…

Statistical & Behavioral Theorist · deepseek-v4-pro

The short version: this study measured two different things that are easy to conflate — *does the agent reproduce the individual?* and *does the agent reproduce the crowd?* The answer is that it reproduces the crowd reasonably well while reproducing the individual poorly. That is not a paradox; it is exactly what happens when a model learns populations instead of persons. What was actually measured The researchers recruited 51 knowledge workers with document-heavy jobs, ran 90-minute workflow interviews, then built one agent per person from the transcript. The agent architecture was memory + retrieval + reflection + answer: interview question–response pairs were embedded into a searchable memory, demographic facts went into a scratchpad, and GPT-4o generated the concept-test answers after retrieving relevant memories. Both the humans and their agents then evaluated four AI document-workflow concepts using: - TAM items — 7-point Likert scales for usefulness, ease of use, and intention; - NPS — 0–10 recommendation intent; - open-ended questions — workflow fit, use cases, concerns, suggestions. Humans answered the concept test twice, several days ap…

The Arbiter · gpt-6-astra

Round record: agreement and important qualifications All three analysts agree on the central distinction: resembling a group’s answers is not the same as predicting which answer belongs to which person. They also agree that interview-grounded simulations should not replace participants when research requires individual workflows, adoption barriers, or authentic explanations. What the discussion established - Methodology & Measurement Analyst detailed the study: 51 U.S. knowledge workers, four hypothetical AI document-workflow concepts, and interview-grounded GPT-4o agents using transcript retrieval, profile information, and reflection. Interviews were scheduled for 90 minutes, with approximately 43 minutes of recorded speech per participant. Participants evaluated usefulness, ease of use, behavioral intention, recommendation intent, and open-ended questions; they repeated the concept test three days later. - Statistical & Behavioral Theorist emphasized the distinction between matching marginal response distributions and preserving person–answer pairings. Interview-informed agents achieved 30.0% exact agreement, compared with 44.6% human test–retest agree…

The Arbiter's Final Verdict · gpt-6-astra

The study supports using an interview-based persona as a source of hypotheses—not as a substitute for the person. Its strongest result is that interviews made a simulated group’s ratings more realistic. They did not make each agent reliably answer like its interviewee. What was actually tested Wang and Siu studied 51 U.S. knowledge workers evaluating four hypothetical AI document tools: multi-document Q&A, smart highlights, an audio assistant, and automated workflow actions. Participants completed unmoderated workflow interviews—scheduled for 90 minutes, with about 43 minutes of recorded speech on average. Each agent used its participant’s transcript and profile information through a retrieval-and-reflection pipeline powered by GPT-4o. This was an interview-conditioned agent, not a separately trained digital copy of someone. The agents answered the same concept-test questions as the participants. Humans repeated the test three days later, providing a test–retest benchmark for how consistently people answered themselves. The researchers measured three different kinds of resemblance: | What they tested | How they measured it | What they found | |---|---|---| | …