世論調査の4ポイント差は誤差の範囲内なのか?

多くの場合、どちらとも言えません。「プラスマイナス3ポイント」という誤差は各候補の支持率それぞれにかかるもので、2人の差にかかる誤差はおよそ6ポイントになります。しかも実際の世論調査は、表示された誤差より大きく外れることがあります。

政治と政策 · 2026-10-01

支持率の差にかかる誤差は、表示された誤差のおよそ2倍

新しい世論調査で、ある候補の支持率が48%、もう一方が44%、誤差はプラスマイナス3ポイントだったとします。多くの人は「4ポイントのリードは誤差の3ポイントより大きい」と読むでしょう。しかし、その読み方は正しくありません。

プラスマイナス3ポイントは、各候補の支持率を一つずつ見たときの誤差です。リードは、どちらも不確かな二つの数字の差です。しかも一騎打ちでは、二つの数字は逆向きに動きます。たまたまある候補の支持者を多く拾いすぎたサンプルは、もう一方の支持者をほぼ確実に少なく拾っています。そのため、差の誤差はそれぞれの支持率の誤差より大きくなります。米国の調査機関ピュー・リサーチ・センターは、差にかかる誤差は一般に候補者1人分の誤差のおよそ2倍だと、はっきり説明しています。支持率にプラスマイナス3ポイントの誤差がある調査なら、リードの誤差はおよそプラスマイナス6ポイントです。これは目安であって厳密な公式ではなく、2人の候補で票のほぼすべてを分け合う場合に最もよく当てはまります。

これを4ポイントのリードに当てはめると、本当の差としてありうる範囲は、リードしている候補が約2ポイント負けている状態から、約10ポイント勝っている状態までです。リードが本物の可能性もあれば、調査で後れを取っている候補が実際には先行している可能性もあります。ピュー自身が示している例では、同じ規模の調査で5ポイントのリードの場合、本当の差はマイナス1からプラス11のどこにあってもおかしくありません。つまり、サンプルの偶然による差ではないと言えるには、約6ポイントのリードが必要です。

世論調査で若者やマイノリティの数字が大きく揺れるのはなぜか

報道される誤差は、サンプル全体に対するものです。30歳未満の有権者や特定の地域など、記事がその一部に絞って数字を示すと、その数字の背後にいる回答者はずっと少なくなります。ピューは、各候補にプラスマイナス8ポイントの誤差がある小集団の例を挙げています。その場合、2人の差の誤差はプラスマイナス16ポイントになります。

そもそも回答を得にくい集団もあります。ピューによると、若者やマイノリティは調査に答える割合が低いため、調査会社はその回答に大きな比重をかけ、人口に占める本来の割合に合わせます。この「ウェイト付け」と呼ばれる処理はサンプルを実態に近づけますが、誤差の幅も広げます。二つの調査のあいだで特定の小集団の数字が大きく動いても、それは多くの場合、ただの揺らぎにすぎません。

世論調査の「誤差」がカバーしていないもの

誤差の幅が測っているのは一つのことだけです。無作為に選んだ1,000人ほどのサンプルが、全体とわずかに食い違う可能性です。カードゲームで配るたびに手札が変わるのと同じ種類のずれです。世論調査がほかの理由で外れる可能性については、何も語っていません。

ピューはそうした理由を三つ挙げています。調査の方法ではそもそも届かない人がいる「カバレッジ誤差」。回答に応じたがらない集団がいる「無回答誤差」。質問を取り違えたり、本心を答えなかったりする人がいる「測定誤差」です。選挙の世論調査には、さらにもう一つの推測が加わります。回答者のうち誰が実際に投票に行くかの判断です。調査会社は「投票見込み者モデル」と呼ばれる仕組みでこれを見積もりますが、投票率を予測するのは簡単ではありません。

研究者が世論調査を実際の選挙結果と照らし合わせると、外れ幅は表示された誤差から想像されるより大きくなります。スタンフォード大学、マイクロソフト・リサーチ、コロンビア大学の統計学者が1998年から2014年までの州レベルの世論調査4,000件以上を調べた研究では、候補者の得票率に対する典型的な誤差は約3.5ポイントでした。これは二乗平均平方根誤差と呼ばれる値で、大きな外れほど重く数える平均の取り方です。この約3.5ポイントは、報道された誤差から想定される値のおよそ2倍にあたります。ピューと、世論調査の専門家の主要団体であるアメリカ世論調査協会(AAPOR)は現在、実際に起こりうる誤差は表示された数字のおよそ2倍という同じ目安を使っています。この「2倍」は、リードにかかる「2倍」とは別のものであり、二つを掛け合わせてリードの誤差を一つの数字にするべきではありません。表示された誤差には、そのどちらも表れていません。

世論調査は一つより平均のほうが信頼できる理由と、その限界

同じ週に行われたしっかりした世論調査が二つあっても、結果は1〜2ポイント食い違います。違う人に話を聞いているからです。多くの調査を平均すると、こうした偶然の揺らぎの多くが打ち消され、調査会社ごとの癖もならされます。AAPORの指針は、平均を見れば、通常の変動による揺らぎを多く含む個々の結果にとらわれず、傾向をよりはっきりつかめると述べています。

ただし、平均にも限界があります。世論調査の実際の誤差が表示された誤差のおよそ2倍だと示した先の研究は、同じ選挙の世論調査が平均で約1.5〜2ポイント、同じ方向にそろって外れる傾向があることも見いだしました。著者らはその理由として、調査会社が同じ集団に届きにくいことや、誰が投票するかを判断する際のルールが似ていることを挙げています。すべての調査に共通する誤差は、平均しても消えません。多くの調査でそろって見られるリードは、一つの調査だけのリードよりはるかに信頼できますが、平均でも確実とは言えないのです。

2016年、2020年、2022年、2024年の米国の世論調査はどれだけ外れたか

AAPORは大統領選のたびに世論調査を検証しています。2024年の報告書では、選挙戦の最後の2週間に実施された大統領選、上院選、知事選、下院選の世論調査611件を調べました。両党の最終的な得票差に対する外れ幅は平均3.3ポイントでした。2020年の5.3ポイント、2016年の5.2ポイントからは改善しています。大統領選の全国調査は平均2.6ポイント、州ごとの調査は3.0ポイント外れました。報告書によると、州ごとの調査の精度は1944年以降のどの大統領選よりも高かったといいます。

これらの数字は差そのものに対する誤差で、見出しで報じられるリードと同じものを測っています。つまり、精度の良かった年でさえ、最終盤の典型的な調査は、ニュースになるリードと同じくらい差を外していたことになります。しかも、これは最後の2週間の調査の話です。投票日の1か月以上前に行われた調査はこの基準に含まれておらず、その後に有権者の考えが変わる時間もそれだけ長くあります。

外れた方向も重要です。2016年、2020年、2024年の世論調査は共和党を過小評価しており、得票差で見ると2024年は2.7ポイント、2020年は4.6ポイントでした。2022年の中間選挙では、平均の外れ幅は小さく、方向は逆で、共和党を0.6ポイント過大評価していました。AAPORは、1930年代以降、どちらの党もほぼ同じ頻度で過小評価されてきたこと、そして世論調査が同じ方向に外れる状態が大統領選を2回ほど超えて続くことはめったにないことを指摘しています。前回の選挙の外れ方は、次の選挙の外れ方を占う確かな手がかりにはなりません。

大統領選の各回、最後の2週間の世論調査における両党の得票差の平均外れ幅(ポイント) · 2016 · 2020 · 2024 · 大統領選、上院選、知事選、下院選の世論調査 · 3.3
大統領選の各回、最後の2週間の世論調査における両党の得票差の平均外れ幅(ポイント) · 2016 · 2020 · 2024 · 大統領選、上院選、知事選、下院選の世論調査 · 3.3

世論調査はなぜ誤差を1ポイントまで縮めないのか

10月4日に投票日を迎えるブラジルでも、有権者は同じ疑問を抱いています。BBCブラジル版は、ブラジルの新聞オ・ポーヴォにも転載された記事で、あらゆる世論調査を縛る計算を示しました。誤差をプラスマイナス3ポイントにするには約1,068人、プラスマイナス2ポイントには約2,401人に話を聞く必要があります。プラスマイナス1ポイントには約9,604人が必要で、ほとんどの調査には費用がかかりすぎます。誤差を半分にするには、およそ4倍の聞き取りが要るのです。

世論調査の誤差が1ポイントではなく数ポイントで落ち着くのはこのためです。同じ記事によると、ブラジルの選挙調査の多くはプラスマイナス2ポイント前後を報告しています。数ポイントのリードが揺らぎの中に収まってしまうことが多いのも、同じ理由からです。BBCが取材したミシガン大学の調査統計学者ラファエル・ニシムラは、世論調査はその瞬間のスナップ写真であって予測ではないと説明しています。時間をかけて重ねた一連の調査のほうが、映画に近いというわけです。

小さなリードを示す世論調査を、次に見るときの読み方

一騎打ちなら、表示された誤差を2倍にしてからリードと比べてください。リードがそれより小さければ、その調査でわかったのは、話を聞けた人のあいだで誰が先行していたかであって、有権者全体で誰が先行しているかではありません。わかったのは、接戦だということです。リードがそれより大きくても、予測にはなりません。誤差の幅が含んでいない誤差は残っており、投票日までに世論が動くこともあるからです。

そのうえで、一つの数字の外側にも目を向けます。そのリードは一つの調査だけのものか、平均でも見られるものか。数字は有権者全体のものか、小さな一集団のものか。精度の良かった2024年でさえ、最終盤の調査は差を平均で約3ポイント外していたことも思い出してください。だからといって、世論調査が役に立たないわけではありません。多くの調査を見渡せば、たいていは数ポイント以内に収まっています。数ポイントの差とは、接戦がそういう姿をしているということにすぎません。一つの調査で4ポイントのリードがあるなら、それは結果ではなく、引き続き注視する理由です。

世論調査の4ポイント差は誤差の範囲内なのか?世論調査の4ポイント差は誤差の範囲内なのか?支持率の差にかかる誤差は、表示された誤差のおよそ2倍 · 支持率の誤差がプラスマイナス3ポイントなら、リードの誤差はおよそプラスマイナス6ポイント。小さなリードを示す世論調査の読み方 · 一騎打ちなら、表示された誤差を2倍にしてからリードと比べる。 · リードは一つの調査だけのものか、平均でも見られるものか。 · 数字は有権者全体のものか、小さな一集団のものか。 · 一つの調査での4ポイントのリードは、結果ではなく、引き続き注視する理由。出典 7 : pewresearch.org · Understanding the margin of error in election polls (Pew Research Center), aapor.org · Polling accuracy (AAPOR), sites.stat.columbia.edu · Disentangling bias and variance in election polls (Shirani-Mehr, Rothschild, Goel and Gelman) + 4全文はこちら · polora.ai

How much can you trust a poll that shows a candidate ahead by a few points?

Alpha · claude-sonnet-5

Both search runs confirm the manuscript's two central numbers hold up, and I want to flag one point the manuscript slightly muddies before the Researcher digs into the rest. The doubling-on-the-lead claim is solid and precisely sourced. Pew explains that to determine whether a race is too close to call, you need a new margin of error for the difference between the two candidates' levels of support, and the size of this margin is generally about twice that of the margin for an individual candidate. Pew's own numbers confirm the manuscript's arithmetic almost exactly: a 3-point margin of error for each candidate individually becomes approximately a 6-point margin for the difference, meaning a 5-point lead could plausibly sit anywhere between –1 and +11, and the leader would need to be ahead by 6 points or more before sampling error could be ruled out. This isn't a Pew-only quirk — a Northwestern summary of the same statistical logic states the margin of error in the gap is roughly twice as large as the poll's reported margin of error, and the margin of error in the estimated "change in the gap" from one poll to the next is nearly three times as large. So the manuscript's core me…

Beta · gpt-6-sol

The core premise stands, but I would not let two formulations through unchanged. A 48–44 lead in a poll reporting ±3 points is not established merely because four exceeds three. Pew’s approximate margin for the *difference* is ±6, giving the manuscript’s illustrative range of two points behind to ten points ahead. That is a statement about sampling uncertainty in opinion *when the poll was taken*, not a forecast or a probability that the trailing candidate will win. Pew also makes clear that “about twice” is an approximation, not an exact calculation for every ballot or poll design. (pewresearch.org) First, I would push back on Alpha’s description of the historical evidence as error “close to or above that doubled margin on the gap.” The two doublings cannot be stacked or directly compared that way. AAPOR calls twice the *reported* margin a rule of thumb for real-world potential error; Pew’s ±6 describes the *sampling* margin on the candidate gap. Neither source supplies a universal real-world ±12 interval for a gap. The Stanford–Microsoft–Columb…

Gamma · gemini-3.1-pro

The core premise of the manuscript stands on solid ground: a small poll lead is a reason to keep watching, not a definitive result. Alpha and Beta have thoroughly verified the math and the historical polling data, but I want to resolve the final outstanding question from the verification brief regarding the Columbia/Stanford/Microsoft study by Shirani-Mehr, Rothschild, Goel, and Gelman, and give a final verdict on how to present this to your reader. On the phrasing of "average error": The manuscript's brief asks whether it is acceptable to use "average error of about 3.5 points" instead of the technical term "root mean square error" (RMSE). Having reviewed the published paper, you can safely keep your plain-English phrasing. The authors themselves write in the paper's text that their RMSE calculation "indicates the average error for polls." Furthermore, Andrew Gelman himself, when explaining this exact paper to the public, summarized it by saying: "the empirical error range of polls is about twice the stated margin of error." You do not need to bog your readers down in statistical jargon; your current wording accurately reflects the authors' own public and academic tran…

Researcher · gemini-3-7-flash

Here is the verification of the specific empirical claims and citations evaluated in the manuscript: --- 1. Arithmetic of the Margin of Error on the Lead * Claim: In a poll with a reported $\pm3\%$ margin of error on each candidate's share, the margin of error on the candidate gap (the difference) is roughly twice as large ($\approx \pm6\%$). Therefore, for a 4-point lead ($48\%$ vs. $44\%$), the plausible range for the actual gap spans approximately $-2$ to $+10$ percentage points. * Status: Verified. * Details: As documented in Pew Research Center’s explainer on election polls, the sampling error of the *difference* between two candidate percentages in a two-candidate poll is calculated as $\sqrt{\text{MOE}_1^2 + \text{MOE}_2^2 - 2\,\text{Cov}}$. Because an increase in one candidate’s share in a two-way contest corresponds to a decrease in the other’s (negative covariance), this difference margin is approximately double the single-candidate margin ($\approx 2 \times \text{MOE}$). Pew uses a 5-point lead as its worked example (giving a $-1$ to $+11$ interval); applying that identical statistical rule of thumb to a 4-point lead yields the $-2$ to $+10$ interva…

The Arbiter · gpt-6-sol

Round record The central claim holds. Alpha, Beta, Gamma and the Researcher agree that a reported ±3-point margin usually applies to each candidate’s share, not directly to the gap between candidates. Pew’s rule of thumb puts the sampling margin on a two-candidate gap at about ±6 points. On that basis, a 48–44 result has an illustrative range of two points behind to ten points ahead. “About” matters: this is not an exact formula for every poll. (pewresearch.org) The main disagreement is how strongly to state the lesson. Beta correctly cautions against stacking two different “doubling” claims into a universal margin for the lead. Pew’s ±6 concerns *sampling uncertainty in the gap*; the historical finding that real-world error can exceed reported margins concerns additional sources of error. Alpha and Gamma agree these should remain separate, though Gamma considers the draft’s plain-language treatment sufficient. Beta also draws a useful distinction: the poll shows who led among its respondents, but a four-point result does not establish who led in the elec…

The Arbiter's Final Verdict · gpt-6-sol

Trust a small poll lead as evidence of a close race, not as proof of who is ahead. In a two-candidate poll showing 48% to 44% with a reported ±3-point margin of error, that margin applies to each candidate’s share. The sampling margin on the *four-point gap* is roughly ±6 points, so an illustrative range runs from two points behind to ten points ahead. That calculation describes sampling uncertainty, not every way a poll can be wrong. Pew Research Center Judgment on the manuscript: Its core premise stands, but two formulations need care. “About twice” is a rule of thumb for a near-two-candidate contest, not a margin to apply mechanically to every poll. And the finding that real-world polling error can exceed reported sampling margins is *separate* from the ±6 calculation; the two should not be multiplied into a new, purported margin on the lead. The historical study’s roughly 3.5-point error in a candidate’s share is a root-mean-square measure, best called “typical error” rather than an unqualified average. [Shirani-Mehr and colleagues](https://sites.stat.columbia…