연구진은 일하는 사람 한 명 한 명을 인터뷰해 그 사람을 대신할 에이전트를 만든 뒤, 실제 본인과 견주어 보았다. 에이전트는 집단의 답은 따라갔지만 개인은 놓쳤다.
해설 대상 : Interview-Informed Generative Agents for Product Discovery: A Validation Study, Zichao Wang et al., 2026-03-10, v1 원문 보기
AI와 사회 · 2026-09-09
팔기 쉬운 약속이다. 긴 인터뷰에 한 번 응하면 AI가 당신이 말하고 일하는 방식을 익히고, 그다음부터는 당신을 본뜬 판본이 당신 대신 질문에 답한다. 실제 당신을 다시 성가시게 할 일이 없다는 것이다. 최근 한 연구가 바로 그런 대역을 만들어, 그것이 본뜬 사람들과 견주어 보았다. 결과는 깔끔하면서도 조금 맥이 빠진다. 그 복제본은 집단은 닮았지만 개인은 닮지 않았다.
연구진은 자기들이 만든 에이전트를 두고 '분포는 맞지만 개인은 어긋난다'고 말한다. 쉽게 풀면, 이런 에이전트를 잔뜩 모아 놓으면 실제 사람들의 무리가 내놓는 의견 분포를 대략 재현했지만, 에이전트 하나하나는 그것이 바탕으로 삼은 사람 한 명을 믿을 만하게 대신하지 못했다는 뜻이다. 이 결과가 무엇을 말하는지 따져 보려고, 폴로라는 같은 질문을 여러 회사가 만든 여러 AI 모델에 던지고 그 모델들이 이 연구를 함께 살펴보게 했다.
연구가 무엇을 쟀나
실험 설계는 단순했다. 미국의 지식 노동자 51명, 모두 문서를 많이 다루는 일을 하는 사람들이 저마다 자기 업무 흐름을 놓고 깊이 있는 인터뷰에 응했다. 인터뷰는 90분으로 잡혔고, 녹음된 발언은 평균 43분쯤이었다. 연구진은 인터뷰 기록 하나하나로 개인 에이전트를 하나씩 만들었다. 이 에이전트는 인터뷰를 저장해 두었다가 새 질문이 들어오면 관련된 대목을 꺼내 오고, 언어 모델이 그 참가자인 것처럼 따져서 답하는 방식으로 돌아간다.
그런 다음 사람과 그 에이전트가 똑같이, 문서 작업을 돕는 가상의 AI 도구 넷을 평가했다. 하나는 여러 파일에 걸쳐 질문에 답하는 도구, 하나는 중요한 대목을 짚어 주는 도구, 하나는 문서를 음성으로 바꿔 주는 도구, 하나는 여러 단계로 이어진 일을 스스로 해내는 도구다. 각 도구는 얼마나 쓸모 있는지, 쓰기 쉬운지, 받아들이거나 남에게 권할 마음이 있는지를 재는 단순한 척도로 점수를 매겼고, 자유롭게 쓴 글도 함께 받았다. 사람들은 사흘 뒤에 같은 평가를 한 번 더 했는데, 이로써 한 사람이 자기 자신과 얼마나 일관되는지를 재는 잣대가 생겼다. 핵심 확인은 단순했다. 각 에이전트의 답을 그 에이전트를 만든 바로 그 사람의 답과 나란히 놓고 견주는 것이다.
개인 수준에서 에이전트는 약했다. 자기가 본뜬 사람과 정확히 일치하는 비율을 보면, 인터뷰로 만든 에이전트는 30%쯤 맞혔다. 같은 사람이 며칠 뒤 그저 다시 답했을 때의 45%쯤에 견주면 낮다. 연구는 이를 사람이 보이는 일관성의 67%쯤이라고 표현하는데, 이는 30을 45로 나눈 값이지 답의 67%를 맞혔다는 뜻이 아니다. 에이전트에게 인터뷰 전체를 넣어 줘도 별 도움이 되지 않았다. 인구통계 정보 몇 가지만 준 에이전트나 아무것도 주지 않은 에이전트보다 뚜렷하게 낫지 않았다.
집단 수준에서는 그림이 뒤집혔다. 누가 무슨 말을 했는지를 묻지 않고 답 전체가 어떻게 퍼져 있는지만 보면, 인터뷰로 만든 에이전트가 실제 분포에 가장 가까웠다. 낮고 시큰둥한 점수를 조금이라도 재현한 것도 이 판본뿐이었다. 개인 정보를 하나도 주지 않은 에이전트는 무난한 기본값으로 주저앉아, 대부분의 답을 각 척도의 꼭대기 근처에 몰아 놓았다. 그래서 에이전트의 무리는 실제 무리처럼 보였지만, 그중 어느 하나도 자기가 본뜬 실제 사람처럼 보이지는 않았다.
각자가 본뜬 사람의 원래 답과 정확히 일치한 비율 · 인터뷰로 만든 에이전트 · 같은 사람이 며칠 뒤 그저 다시 답했을 때 · 30%쯤 · 45%쯤
열 사람을 그려 보자. 다섯은 어떤 구상을 좋아하고 다섯은 싫어한다. 어떤 시뮬레이션은 이 다섯 대 다섯의 갈림을 완벽하게 재현하면서도 한 사람 한 사람은 모조리 거꾸로 맞힐 수 있다. 막대그래프는 맞지만 개개인에 대한 짐작은 다 틀린 것이다. 이것은 에이전트가 실제로 모두를 뒤집었다는 주장이 아니라 하나의 예시일 뿐이지만, 맞는 분포가 어떻게 개개인의 오답 더미를 가릴 수 있는지 보여 준다. 평균을 내면 실수들이 서로 상쇄된다.
무슨 일이 있었는지는 평범한 이유 몇 가지로 설명된다. 인터뷰는 한 사람의 세계 가운데 일부를 담을 뿐, 어떤 반응을 좌우하는 모든 조건을 담지는 못한다. 한 번도 써 보지 않은 도구에 대한 판단은 대개 정해져 있던 취향에서 읽어 내는 것이 아니라 그 자리에서 즉석으로 지어진다. 기록이 얇은 대목에서는 모델이 그 빈틈을, 이 사람이 아니라 흔한 노동자처럼 들리는 그럴듯하고 일반적인 논리로 메운다. 한 무리의 생김새를 안다고 해서 그 안의 특정한 한 사람에 대해 보장되는 것은 거의 없다.
매끄럽게 다듬은 자국이 드러나는 곳
자유롭게 쓴 답을 보면 그 틈이 구체적으로 드러난다. 음성 도구를 두고 사람과 에이전트는 같은 걱정거리를 짚었다. 정확성, 사생활, 이미 쓰는 소프트웨어에 들어맞는지였다. 그런데 실제 사람의 30%쯤은 그 구상을 대놓고 거부했고, 어떤 이는 어떤 경우에도 쓰지 않겠다고까지 말했다. 그만큼 세게 맞선 에이전트는 딱 하나뿐이었다. 사람들은 구체적인 결격 사유를 댔다. 휴대전화가 아니라 데스크톱에서 일한다든가, 들어서가 아니라 읽어야 한다든가, 회사의 규정 준수 규칙 같은 것이다. 에이전트는 그런 단호한 거절을, 무언가 개선하면 풀릴 문제로 바꿔 놓기 일쑤였다.
이 차이가 중요한 까닭은, '개선되면 받아들이겠다'는 말이 '이건 내 일에 안 맞는다'는 말과 같은 답이 아니기 때문이다. 에이전트는 목소리와 말투에서도 가장 낮은 점수를 받았다. 사람들이 어수선하고 머뭇거리며 때로 앞뒤가 안 맞게 말할 때, 에이전트는 말끔한 글을 내놓았다. 여기에 한 가지 단서를 달아야 한다. 이렇게 자유롭게 쓴 답을 매긴 것은 서로 어느 정도만 의견이 맞은 AI 심판들이라, 최종 판정이 아니라 뒷받침하는 증거일 뿐이다.
어떤 업체나 연구팀이 당신 자신의 인터뷰로 만든 페르소나를 내놓으며 당신을 대신할 수 있다고 말한다면, 이 연구는 조심하라는 근거가 된다. 더 안전하게 보자면, 그것이 내놓는 새 답은 당신이 한 말이 아니라 당신에 대한 예측이다. 낯익은 어휘와 당신 일에 대한 언급은 그것이 당신의 재료를 썼다는 것을 보여 준다. 그 결론이 당신을 대변한다는 것을 보여 주지는 않는다.
실제로 이는, 당신의 말로 만든 페르소나에게 당신의 거부를 그대로 재현하거나, 당신이 왜 망설일지 설명하거나, 어떤 도구가 당신의 실제 업무 흐름에 어떻게 들어맞을지 판단하는 일을 맡겨서는 안 된다는 뜻이다. 그것이 당신을 대신해 동의하거나, 당신을 대신해 결정하거나, 당신의 견해로 인용되어서는 안 된다. 당신이 말한 것처럼 쓰일 답이라면 무엇이든 당신에게 돌아와 확인을 받아야 한다.
값을 하는 곳과 못 하는 곳
그래도 좁게나마 진짜 쓸모는 있다. 이런 에이전트의 무리는 실제 무리를 닮으므로, 값싼 1차 거름망으로 쓸 수 있다. 시제품에 돈을 들이기 전에 거친 구상 여럿을 걸러 내거나, 자동화를 더 할지 통제를 더 할지 같은 큰 방향을 견주거나, 실제 조사로 가져갈 반론과 질문의 후보를 뽑아 내는 식이다. 연구는 이 운영 비용을 사람과의 한 세션보다 훨씬 낮게, 약 12.50달러에 견주어 1.27달러쯤으로 잡는다. 다만 이 두 값이 같은 작업 단위를 두고 잰 것인지는 온전히 분명하지 않다.
특정한 한 사람이 되어야 하는 일이라면 어디에도 쓸 자리가 없다. 그것은 이름이 있는 참가자를 대변할 수 없고, 누군가의 신뢰나 도입을 막는 장벽이 무엇인지 밝힐 수 없으며, 업무의 세부가 요점일 때 인터뷰를 대신할 수 없고, 지어낸 논평을 고객의 증언으로 바꿔 놓을 수 없다. 거름망으로 쓸 때조차 경고가 따른다. 오늘의 답이 퍼진 모양을 맞혔다고 해서 그 에이전트가 내일의 구상을 올바로 줄 세우리라는 증명은 아니다. 그러니 이것은 유망한 착상일 뿐 입증된 것은 아니다.
같은 도구가 무리에는 값을 하고 한 사람에게는 못 한다. · 값을 하는 곳 · 못 하는 곳 · 이런 에이전트의 무리는 실제 무리를 닮으므로, 값싼 1차 거름망으로 쓸 수 있다. · 특정한 한 사람이 되어야 하는 일이라면 어디에도 쓸 자리가 없다.
판결이 아니라 한 편의 연구
이것은 특정 분야에 국한된 하나의 연구다. 미국의 지식 노동자 51명, 가상의 문서 도구 넷, 에이전트를 만드는 한 가지 방식이다. 이 연구는 충실한 개인 시뮬레이션이 불가능하다고 증명하지 않으며, 더 나은 방법이 그 틈을 좁힐 수도 있다. 이 연구가 하는 일은 입증의 책임을 제자리에 옮겨 놓는 것이다. 에이전트를 당신의 말에 뿌리내리게 하는 것과 그것이 당신을 대신하게 만드는 것은 다른 일이고, 그 두 번째 주장은 이제 당연하게 여길 것이 아니라 보여 줘야 한다.
여기서 나오는 실천 원칙은 간단하다. 무리만 있으면 되는 일, 빠르고 거친 총합 신호가 필요한 일에는 이런 에이전트의 무리를 쓴다. 특정한 한 사람이어야 하는 일에는 실제 사람을 쓴다. 당신 일 안에 자리 잡을 도구를 두고 당신 자신이 내리는 결정도 여기 든다.
People are increasingly told that an AI can be built from their own words and stand in for them. A new study put that promise to a direct test. Researchers ran in-depth workflow interviews with knowledge workers, built a personal agent from each interview, and then compared what each agent said about new product concepts against what that same person actually said. The result is precise and uncomfortable : the agents were distribution-calibrated but identity-imprecise. They approximated how a population would answer, yet failed to reproduce the specific individual they were grounded in. Explain what was actually measured and how it was measured, why an agent can match a crowd without matching the person it was built from, and what that gap means for anyone deciding whether to trust a persona made from their own interview. Say plainly where such a simulation still earns its keep and where it does not.
What Was Measured and How
In Wang & Siu (2026), Adobe Research tested the premise of using interview-informed generative agents to simulate individual users in early-stage product discovery.
The experimental design was structured around 51 knowledge workers (spanning legal, finance, management, operations, and research) who frequently handle dense documents. The measurement pipeline proceeded in distinct stages:
1. Grounding Protocol: Each participant completed an in-depth, 90-minute unmoderated audio interview (averaging ~43 minutes of speech, ~106 words per response, totaling ~36 hours across the cohort) exploring their document workflows, friction points, risk tolerance, and AI attitudes.
2. The Ground Truth Concept Tests: Participants evaluated four speculative AI document tools (Multidoc Q&A, Smart Highlights, Audio Assistant, and Workflow Actions). They rated each concept on:
- Quantitative scales: A 6-item Technology Acceptance Model (TAM) measuring Perceived Usefulness, Perceived Ease of Use, and Behavioral Intention (7-point Likert), plus Net Promoter Score (NPS, 0–10).
- Qualitative feedback: Open-ended rationales detailing workflow fit,…
What the study actually tested
This was not a test of whether an AI can summarize an interview or produce plausible feedback. It tested a stronger claim: can an agent built from one person’s interview predict that same person’s answers to new product concepts?
The researchers recruited 51 U.S. knowledge workers with document-heavy jobs, conducted in-depth workflow interviews, then built one agent per participant. Each agent had:
- the participant’s interview transcript, stored for retrieval;
- basic profile information, such as role;
- a prompting workflow that retrieved relevant interview passages, generated a “reflection” about the person’s perspective, then answered concept-test questions.
The people and their agents were then asked the same questions about four proposed AI document-workflow products: multi-document Q&A, smart highlights, an audio document assistant, and an automated workflow-actions assistant.
The test included:
- scalar ratings: six Technology Acceptance Model items on usefulness, ease of use, and intention to use, plus an 0–10 Net Promoter Score;
- open-ended feedback: workflow fit, use cases, concerns, improvements, and comparisons amon…
The short version: this study measured two different things that are easy to conflate — *does the agent reproduce the individual?* and *does the agent reproduce the crowd?* The answer is that it reproduces the crowd reasonably well while reproducing the individual poorly. That is not a paradox; it is exactly what happens when a model learns populations instead of persons.
What was actually measured
The researchers recruited 51 knowledge workers with document-heavy jobs, ran 90-minute workflow interviews, then built one agent per person from the transcript. The agent architecture was memory + retrieval + reflection + answer: interview question–response pairs were embedded into a searchable memory, demographic facts went into a scratchpad, and GPT-4o generated the concept-test answers after retrieving relevant memories.
Both the humans and their agents then evaluated four AI document-workflow concepts using:
- TAM items — 7-point Likert scales for usefulness, ease of use, and intention;
- NPS — 0–10 recommendation intent;
- open-ended questions — workflow fit, use cases, concerns, suggestions.
Humans answered the concept test twice, several days ap…
The Arbiter · gpt-6-astra
Round record: agreement and important qualifications
All three analysts agree on the central distinction: resembling a group’s answers is not the same as predicting which answer belongs to which person. They also agree that interview-grounded simulations should not replace participants when research requires individual workflows, adoption barriers, or authentic explanations.
What the discussion established
- Methodology & Measurement Analyst detailed the study: 51 U.S. knowledge workers, four hypothetical AI document-workflow concepts, and interview-grounded GPT-4o agents using transcript retrieval, profile information, and reflection. Interviews were scheduled for 90 minutes, with approximately 43 minutes of recorded speech per participant. Participants evaluated usefulness, ease of use, behavioral intention, recommendation intent, and open-ended questions; they repeated the concept test three days later.
- Statistical & Behavioral Theorist emphasized the distinction between matching marginal response distributions and preserving person–answer pairings. Interview-informed agents achieved 30.0% exact agreement, compared with 44.6% human test–retest agree…
The Arbiter's Final Verdict · gpt-6-astra
The study supports using an interview-based persona as a source of hypotheses—not as a substitute for the person. Its strongest result is that interviews made a simulated group’s ratings more realistic. They did not make each agent reliably answer like its interviewee.
What was actually tested
Wang and Siu studied 51 U.S. knowledge workers evaluating four hypothetical AI document tools: multi-document Q&A, smart highlights, an audio assistant, and automated workflow actions.
Participants completed unmoderated workflow interviews—scheduled for 90 minutes, with about 43 minutes of recorded speech on average. Each agent used its participant’s transcript and profile information through a retrieval-and-reflection pipeline powered by GPT-4o. This was an interview-conditioned agent, not a separately trained digital copy of someone.
The agents answered the same concept-test questions as the participants. Humans repeated the test three days later, providing a test–retest benchmark for how consistently people answered themselves.
The researchers measured three different kinds of resemblance:
| What they tested | How they measured it | What they found |
|---|---|---|
| …