あなたの言葉から作ったAIは、集団には似ても、あなたには似ない

研究者は働き手ひとりずつのインタビューからエージェントを作り、その本人と照らし合わせて検証した。エージェントは集団の回答は追えたが、個人は捉えられなかった。

解説の対象 : Interview-Informed Generative Agents for Product Discovery: A Validation Study, Zichao Wang et al., 2026-03-10, v1 原文を読む

AIと社会 · 2026-09-09

売り込みやすい約束だ。長いインタビューを受け、話し方や仕事の進め方をAIに学ばせれば、それ以降はあなたの代わりに「あなたの分身」が質問に答えてくれる。もう本人をわずらわせる必要はない、というわけだ。最近のある研究は、まさにこの種の代役を作り、それが写し取った本人たちと照らし合わせて確かめた。結果は明快で、少しばかり気勢をそぐものだった。分身は集団には似ていたが、個人には似ていなかった。

研究者は自分たちのエージェントを「分布としては整っているが、個人としては不正確」と呼ぶ。かみくだいて言えば、こうしたエージェントを大勢集めると、現実の集団が示す意見の割合をおおむね再現するのに、一体一体のエージェントは、その土台となったたったひとりの人間の代役としては当てにならない、ということだ。これが何を意味するのかを突き止めるため、Poloraは同じ問いを別々の会社が作った複数のAIモデルに投げかけ、この研究をともに検討させた。

研究は何を測ったか

仕組みは単純だ。米国の51人のナレッジワーカー、いずれも書類を多く扱う職種の人たちが、それぞれ業務についての詳しいインタビューを受けた。予定時間は90分で、録音された発話は平均でおよそ43分ぶんになった。研究者は各人の書き起こしからひとつずつ個人用エージェントを作った。使ったのは、インタビューを保存しておき、新しい質問が来たら関連する部分を取り出し、言語モデルにその参加者になりきって考えさせ答えさせる、という一連の仕組みだ。

続いて、人間とそのエージェントの双方が、書類作業のための架空のAIツール4つを同じように評価した。多数のファイルをまたいで質問に答えるもの、重要な箇所を強調するもの、書類を音声に変えるもの、そして複数の手順からなる作業を自力で進めるものだ。それぞれについて、役に立つか、使いやすいか、導入したり勧めたりしたいかといった素朴な尺度で採点し、あわせて自由記述の感想も書いてもらった。人間には3日後に同じテストをもう一度受けてもらい、これによって、ある人が自分自身とどれだけ一貫しているかという基準が得られた。核心の確認はそっけないものだった。各エージェントの回答を、それを作る土台となったまさにその本人の回答と並べて突き合わせるのだ。

集団は一致し、個人は一致しなかった

個人の水準では、エージェントは弱かった。自分の土台となった本人との完全一致で見ると、インタビューに基づくエージェントが一致したのはおよそ30%で、これに対し同じ人たちが数日後にただ答え直したときの一致はおよそ45%だった。研究はこれを「人間の一貫性の約67%」と表現しているが、これは30を45で割った値であって、回答の67%が正しいという意味ではない。エージェントにインタビューの全文を与えても、ほとんど効き目はなかった。いくつかの属性情報だけを与えたエージェントや、何も与えなかったエージェントと比べて、はっきり優れているとは言えなかったのだ。

集団の水準では、様相が逆転した。誰がそう言ったかを問うのをやめ、回答全体の散らばりだけに目を向けると、インタビューに基づくエージェントが現実の散らばりに最も近づいた。低く、気乗りしない点数を少しでも再現できたのは、この版だけだった。個人情報を何も与えられなかったエージェントは、無難な既定値へと崩れ落ち、ほとんどの回答を各尺度の上端あたりに寄せてしまった。つまり、エージェントの集団は現実の集団のように見えたが、そのうちのどの一体も、現実の本人には見えなかった。

本人との完全一致の割合を、エージェントと数日後に答え直した本人で比べた。 · インタビューに基づくエージェント · 同じ人たち · およそ30% · およそ45%
本人との完全一致の割合を、エージェントと数日後に答え直した本人で比べた。 · インタビューに基づくエージェント · 同じ人たち · およそ30% · およそ45%

集団に似ることが、あなたに似ることではない理由

10人を思い浮かべてほしい。ある構想を気に入る人が5人、気に入らない人が5人だ。あるシミュレーションが、この5対5の割れ方を完璧に再現しながら、一人ひとりについては全員を取り違える、ということもありうる。全体の割合は正しいのに、個人ごとの推測はすべて外れている。これはエージェントが文字どおり全員を反転させたという主張ではなく、あくまで例え話だが、正しい分布が個人の誤りの山をどう覆い隠しうるかを示している。平均を取ることで、誤りが打ち消し合ってしまうのだ。

起きたことには、ありふれた理由がいくつも当てはまる。インタビューが捉えるのはその人の世界の一部であって、反応を左右するすべての条件ではない。触れたこともないツールについての判断は、確立した好みから読み取るというより、その場で組み立てられることが多い。記録が薄いところでは、モデルはその隙間を、もっともらしく当たり障りのない理屈で埋める。それはこの本人というより、ありふれた働き手らしく聞こえる理屈だ。集団の形を知っていても、そのなかの特定の一人について確実なことは、ほとんど何も分からない。

ならされた跡が見えるところ

自由記述の回答が、その隔たりをはっきりさせる。音声ツールについて、人間とエージェントは同じ主だった懸念を挙げた。正確さ、プライバシー、既存のソフトへのなじみやすさだ。ところが、現実の人々のおよそ30%はこの案をきっぱり退け、なかにはどんな状況でも使わないと言う人もいたのに対し、そこまで強く反対したエージェントはたった一体だけだった。人々は、電話ではなくデスクトップで作業していること、聞くのではなく書類を読む必要があること、勤め先のコンプライアンス規則といった、具体的な断りの理由を挙げた。エージェントのほうは、そうした決定的な断りを、何らかの改良で解決できる問題へと言い換えがちだった。

この違いが重要なのは、「改良されたら導入する」というのは「これは自分の仕事に合わない」という答えとは別物だからだ。エージェントはまた、声や口調の面で最も低い評価となった。人々が乱雑で、ためらいがちで、ときに矛盾する言い方をするのに対し、エージェントはこぎれいな文章を生み出したからだ。ここでひとつ注意がいる。こうした定性的な比較を採点したのはAIの判定役で、その判定役どうしの一致もほどほどにとどまった。だからこれは裏づけとなる証拠であって、最終的な結論ではない。

あなたの分身にとって何を意味するか

ベンダーや研究チームが、あなた自身のインタビューから作ったペルソナを差し出し、あなたの代わりを務められると言うなら、この研究は慎重になるべき理由になる。より安全な受け止め方は、その新しい回答は、あなたについての予測であって、あなたによる発言ではない、というものだ。見慣れた語彙や、あなたの仕事への言及は、それがあなたの素材を使ったことを示す。だが、その結論があなたを代表していることを示すわけではない。

具体的には、あなたの言葉から作ったペルソナに、あなたの拒否を再現させたり、あなたがなぜためらうのかを説明させたり、あるツールが実際の業務にどう合うかを判断させたりしてはならない、ということだ。それはあなたに代わって同意することも、あなたに代わって決めることも、あなたの見解として引用されることもあってはならない。あなたが言ったものとして使われかねない回答は、どれもあなたのところへ戻し、確認を取るべきだ。

役に立つ場面と、立たない場面

それでも、狭いながらも本物の使い道はある。こうしたエージェントの集団は現実の集団に似るので、安価な最初のふるいになりうる。試作にお金をかける前に多くの粗い構想をふるいにかけたり、より多くの自動化か、より多くの制御かといった大まかな方向性を比べたり、本格的な調査に持ち込むための反論や疑問の候補を出したり、といった使い方だ。研究は、その運用費を人間相手のセッションよりずっと低い、およそ1.27ドルに対して約12.50ドルと見積もっている。ただし、この2つの数字が同じ作業単位で測られているのかは、はっきりしない。

向かないのは、特定の一人であることが仕事の中身になる場面のすべてだ。名前のある参加者を代弁することも、誰かの信頼や導入の障壁を突き止めることも、業務の細部こそが要点であるときにインタビューの代わりを務めることも、作り物のコメントを顧客の証言に変えることもできない。ふるいとしての使い道にすら、注意がついてまわる。今日の回答の散らばりに一致したからといって、明日の構想を正しく順位づけできると証明されたわけではない。だからこれは、実証されたものというより、有望な発想にとどまる。

ひとつの研究であって、結論ではない

これは、ひとつの分野に限られた単独の研究だ。米国のナレッジワーカー51人、架空の書類ツール4つ、エージェントの作り方はひと通りだけ。忠実な個人シミュレーションが不可能だと証明したわけではなく、より優れた手法がこの隔たりを縮めるかもしれない。この研究がしたのは、証明の責任をあるべき場所へ移すことだ。エージェントをあなたの言葉に根づかせることは、それをあなたの代役に仕立てることと同じではない。そして後者の主張は、いまや当然のものとするのではなく、示されなければならない。

そこから導かれる実務上の原則は単純だ。集団だけがあればよいとき、つまり手早くおおまかな全体の傾向がほしいときには、こうしたエージェントの集団を使えばよい。だが、特定の一人でなければならないことには、実在の人間を使う。あなた自身の仕事の内側に入り込むツールについての、あなた自身の判断も、それに含まれる。

あなたの言葉から作ったAIは、集団には似ても、あなたには似ないあなたの言葉から作ったAIは、集団には似ても、あなたには似ない働き手ひとりずつのインタビューからAIエージェントを作り、その本人と照らし合わせて検証した研究がある。米国のナレッジワーカー51人が対象で、架空の書類ツールをどう評価するかを比べた。分身は集団の意見の割合は追えたが、土台となった個人の代役としては当てにならなかった。 · ※ ナレッジワーカー : 書類やデータなど、知識を扱う仕事をする人集団は一致し、個人は一致しなかった · インタビューに基づくエージェント 同じ人たち およそ30% およそ45% 本人との完全一致の割合を、エージェントと数日後に答え直した本人で比べた。集団に似ることが、あなたに似ることではない理由ならされた跡が見えるところ · ※ コンプライアンス : 法令や社内の規則を守ることあなたの分身にとって何を意味するか · ※ ペルソナ : ここでは、本人を模して作られたAIの人物像役に立つ場面と、立たない場面集団だけでよいとき、手早くおおまかな全体の傾向がほしいときには、こうしたエージェントの集団を使えばよい。だが、特定の一人でなければならないことには、実在の人間を使う。出典 1 : arxiv.org · Interview-Informed Generative Agents for Product Discovery: A Validation Study

People are increasingly told that an AI can be built from their own words and stand in for them. A new study put that promise to a direct test. Researchers ran in-depth workflow interviews with knowledge workers, built a personal agent from each interview, and then compared what each agent said about new product concepts against what that same person actually said. The result is precise and uncomfortable : the agents were distribution-calibrated but identity-imprecise. They approximated how a population would answer, yet failed to reproduce the specific individual they were grounded in. Explain what was actually measured and how it was measured, why an agent can match a crowd without matching the person it was built from, and what that gap means for anyone deciding whether to trust a persona made from their own interview. Say plainly where such a simulation still earns its keep and where it does not.

Methodology & Measurement Analyst · gemini-3-8-flash

What Was Measured and How In Wang & Siu (2026), Adobe Research tested the premise of using interview-informed generative agents to simulate individual users in early-stage product discovery. The experimental design was structured around 51 knowledge workers (spanning legal, finance, management, operations, and research) who frequently handle dense documents. The measurement pipeline proceeded in distinct stages: 1. Grounding Protocol: Each participant completed an in-depth, 90-minute unmoderated audio interview (averaging ~43 minutes of speech, ~106 words per response, totaling ~36 hours across the cohort) exploring their document workflows, friction points, risk tolerance, and AI attitudes. 2. The Ground Truth Concept Tests: Participants evaluated four speculative AI document tools (Multidoc Q&A, Smart Highlights, Audio Assistant, and Workflow Actions). They rated each concept on: - Quantitative scales: A 6-item Technology Acceptance Model (TAM) measuring Perceived Usefulness, Perceived Ease of Use, and Behavioral Intention (7-point Likert), plus Net Promoter Score (NPS, 0–10). - Qualitative feedback: Open-ended rationales detailing workflow fit,…

Product Strategy & Trust Evaluator · gpt-5.6-terra

What the study actually tested This was not a test of whether an AI can summarize an interview or produce plausible feedback. It tested a stronger claim: can an agent built from one person’s interview predict that same person’s answers to new product concepts? The researchers recruited 51 U.S. knowledge workers with document-heavy jobs, conducted in-depth workflow interviews, then built one agent per participant. Each agent had: - the participant’s interview transcript, stored for retrieval; - basic profile information, such as role; - a prompting workflow that retrieved relevant interview passages, generated a “reflection” about the person’s perspective, then answered concept-test questions. The people and their agents were then asked the same questions about four proposed AI document-workflow products: multi-document Q&A, smart highlights, an audio document assistant, and an automated workflow-actions assistant. The test included: - scalar ratings: six Technology Acceptance Model items on usefulness, ease of use, and intention to use, plus an 0–10 Net Promoter Score; - open-ended feedback: workflow fit, use cases, concerns, improvements, and comparisons amon…

Statistical & Behavioral Theorist · deepseek-v4-pro

The short version: this study measured two different things that are easy to conflate — *does the agent reproduce the individual?* and *does the agent reproduce the crowd?* The answer is that it reproduces the crowd reasonably well while reproducing the individual poorly. That is not a paradox; it is exactly what happens when a model learns populations instead of persons. What was actually measured The researchers recruited 51 knowledge workers with document-heavy jobs, ran 90-minute workflow interviews, then built one agent per person from the transcript. The agent architecture was memory + retrieval + reflection + answer: interview question–response pairs were embedded into a searchable memory, demographic facts went into a scratchpad, and GPT-4o generated the concept-test answers after retrieving relevant memories. Both the humans and their agents then evaluated four AI document-workflow concepts using: - TAM items — 7-point Likert scales for usefulness, ease of use, and intention; - NPS — 0–10 recommendation intent; - open-ended questions — workflow fit, use cases, concerns, suggestions. Humans answered the concept test twice, several days ap…

The Arbiter · gpt-6-astra

Round record: agreement and important qualifications All three analysts agree on the central distinction: resembling a group’s answers is not the same as predicting which answer belongs to which person. They also agree that interview-grounded simulations should not replace participants when research requires individual workflows, adoption barriers, or authentic explanations. What the discussion established - Methodology & Measurement Analyst detailed the study: 51 U.S. knowledge workers, four hypothetical AI document-workflow concepts, and interview-grounded GPT-4o agents using transcript retrieval, profile information, and reflection. Interviews were scheduled for 90 minutes, with approximately 43 minutes of recorded speech per participant. Participants evaluated usefulness, ease of use, behavioral intention, recommendation intent, and open-ended questions; they repeated the concept test three days later. - Statistical & Behavioral Theorist emphasized the distinction between matching marginal response distributions and preserving person–answer pairings. Interview-informed agents achieved 30.0% exact agreement, compared with 44.6% human test–retest agree…

The Arbiter's Final Verdict · gpt-6-astra

The study supports using an interview-based persona as a source of hypotheses—not as a substitute for the person. Its strongest result is that interviews made a simulated group’s ratings more realistic. They did not make each agent reliably answer like its interviewee. What was actually tested Wang and Siu studied 51 U.S. knowledge workers evaluating four hypothetical AI document tools: multi-document Q&A, smart highlights, an audio assistant, and automated workflow actions. Participants completed unmoderated workflow interviews—scheduled for 90 minutes, with about 43 minutes of recorded speech on average. Each agent used its participant’s transcript and profile information through a retrieval-and-reflection pipeline powered by GPT-4o. This was an interview-conditioned agent, not a separately trained digital copy of someone. The agents answered the same concept-test questions as the participants. Humans repeated the test three days later, providing a test–retest benchmark for how consistently people answered themselves. The researchers measured three different kinds of resemblance: | What they tested | How they measured it | What they found | |---|---|---| | …