なぜ AI パイロットの 95% は失敗するのか、そして 5% が実際にしていること

「AI パイロットの 95% が失敗する」は MIT の読み違いだ : 多くは立ち上がり、ただ測定可能なリターンを示さないだけだ。5% が何を違えているのかという議論は、過小評価された一つの軸を巡り続けた : ツールが使うほど賢くなるかどうか、だ。

AIと社会 · 2026-09-02

「AI パイロットの 95% が失敗する」。これはエンタープライズ AI をめぐって最も頻繁に引用される数字の一つであり、同時に最も読み違えられている数字でもある。出どころは 2025 年の MIT 関連レポートだが、そこで実際に述べられているのはもっと限定的な話だ。AI への支出に対して測定可能なリターンを得られていない組織が 95% にのぼる、というものである。これは、パイロットの 95% が実運用に入る前に頓挫するという主張とは意味が違う。

この違いは、立てるべき問いそのものを変える。多くの企業にとって、失敗とはツールが世に出なかったことではない。ツールは世に出たのに、測定できる変化が何も起きなかった、ということだ。レポートがタスク特化型ツールについて示したファネルを見ると、60% が検討段階に進み、20% がパイロットを実施し、実装まで成功させたのは 5% だった。一方、汎用のチャットツールはこれよりはるかに高く、およそ 83% が実利用へと定着したものの、それが利益にどうつながったかは不明確なことが多かった。

三つの席、一つのループ

この問いを解きほぐすため、Polora は同じ問題を複数の AI モデルに投げかけ、それぞれに異なる立場を割り当てた。ワークフローの立場、スコープの立場、そしてガバナンスの立場である。

ワークフローを担当したモデル gpt-5.6-sol は、デモが証明するのは、整えられた条件下でモデルが見事な答えを出せることだけだ、と主張した。本番はそうはいかない。ツールを軸に、実際の仕事の進め方そのものを組み直す必要があるからだ。スコープを担当したモデル gemini-3-7-flash は、そもそも上流で選ぶ問題を間違えていることが多い、と反論した。成功する 5% は、モデルの出来が最悪でも運用上まだ許容できる程度に小さいタスクを選び、うまくいく道筋を描く前に、まず失敗したときの受け皿を設計する。ガバナンスを担当したモデル grok-4-6 は、何よりオーナーシップを重んじた。デモは政治的な出来事であり、責任を負うオーナーが予算とリスクの承認、そして例外処理の道筋を握っていない限り、廊下で立ち消えになる、と述べた。

問い詰められると、三つのモデルはいずれも、自分の担当する要素だけで十分だとは考えなかった。それぞれが、残る二つも欠かせない条件だと認めた。それらは対立する理論ではなく、同じループへの三つの入り口だったのだ。

誰も真っ先には挙げなかった軸

議論の様相を変えたのは、三つのモデルのいずれも中心に据えなかった第四の軸だった。研究者役を務めたモデル gpt-5.6-sol は原典に立ち返り、そこで挙げられている中核の障壁はプロセスでも政治でもない、と指摘した。レポートはそれを学習ギャップと呼ぶ。システムがフィードバックを取り込んで文脈に適応していくのか、それともデモのままの品質で固まってしまうのか、という問題だ。

これは他の三つとは種類の違う問いだ。スコープは問題の大きさを、ワークフローはツールを取り巻く仕事の流れを、スポンサーシップはそれを守るのは誰かを問う。だがどれも、ツールが使うほど賢くなるかどうかは問わない。スコープを狭く絞り、十分な後ろ盾を得て、深く組み込んだツールであっても、いつまでも賢くならなければ、やはり 95% の側に着地する。

自分のパイロットを見極めるテスト

司会役のモデル claude-sonnet-5 は、四つの要素を順位づけて階層にすることを拒んだ。レポートがそれらを、上下のある梯子ではなく、互いに作用し合う束として描いているからだ。自分のパイロットを見極めたい人に向けた締めくくりのテストは、四つの問いに集約される。失敗しても痛手が小さくて済む程度に、スコープは絞れているか。実際の権限を持つ誰かが、予算と例外処理の道筋を握っているか。AI が働くとき、仕事の流れは本当に変わるのか。そして、使えば使うほどシステムは良くなっていくのか。

この最後の問いこそ、議論の中で軽く見られていたものだった。司会役が指摘したように、レポートのデータによれば、ここが弱い組織は、ほかの三つがそろっていても 95% の区分に踏みとどまってしまう。別の MIT の研究も同じ方向を指し示す。成功する導入とは、働き手をタスクの実行役からツールの監督役へと移す、的を絞った小さな変化なのだ、と。単にモデルを据え付けることではなく、仕事そのものを設計し直すことだ、と。

パイロットを見極める四つの問い · スコープ 失敗しても痛手が小さくて済む程度に、スコープは絞れているか。 · スポンサーシップ 実際の権限を持つ誰かが、予算と例外処理の道筋を握っているか。 · ワークフロー AI が働くとき、仕事の流れは本当に変わるのか。 · 学習 そして、使えば使うほどシステムは良くなっていくのか。 · レポートがそれらを、上下のある梯子ではなく、互いに作用し合う束として描いているからだ。
パイロットを見極める四つの問い · スコープ 失敗しても痛手が小さくて済む程度に、スコープは絞れているか。 · スポンサーシップ 実際の権限を持つ誰かが、予算と例外処理の道筋を握っているか。 · ワークフロー AI が働くとき、仕事の流れは本当に変わるのか。 · 学習 そして、使えば使うほどシステムは良くなっていくのか。 · レポートがそれらを、上下のある梯子ではなく、互いに作用し合う束として描いているからだ。

事実ではなく直感として受け止めるべき数字

やり取りの中では、いくつか印象的な数字も飛び交った。モデルは本番システム全体のわずか 20% を占めるにすぎない、というもの。2% のエラー率が、人間による 100% の確認を強いる、というもの。スコープを狭めれば、18 か月かかる取り組みが 3 週間に縮む、というもの。だが研究に当たったモデルは、これらをいずれも、MIT の研究が実証した事実というより、現場で通用するもっともらしい経験則として扱った。直感としては役に立つ。しかし、確かな知見として引用できるものではない。

タスク特化型ツールが検討から実装へ進む割合 · 検討 60% · パイロット 20% · 実装 5%
タスク特化型ツールが検討から実装へ進む割合 · 検討 60% · パイロット 20% · 実装 5%
なぜ AI パイロットの 95% は失敗するのか、そして 5% が実際にしていることなぜ AI パイロットの 95% は失敗するのか、そして 5% が実際にしていること三つの席、一つのループ誰も真っ先には挙げなかった軸自分のパイロットを見極めるテスト · パイロットを見極める四つの問い出典 readwise-assets.s3.amazonaws.com · The GenAI Divide : State of AI in Business 2025 (MIT NANDA), searchyour.ai · State of AI in Business 2025, full report PDF, mitsloan.mit.edu · Practical AI implementation, MIT Sloan

Your AI pilot wowed everyone in the demo and then quietly died before it reached real use, the fate of roughly 95% of them by MIT's count. What do the 5% that make it to production actually do differently: redesign the workflow, narrow the scope, keep executive sponsorship, or something the demo hides?

Especialista en Rediseño de Flujos de Trabajo · gpt-5.6-sol

The surviving 5% do all three—but workflow redesign is the decisive difference. A demo proves that a model can generate an impressive answer under curated conditions. Production requires a system that changes how work gets completed every day. Successful teams therefore start with the operating process, not the model: - Map the real workflow: inputs, handoffs, approvals, exceptions, systems of record, and accountability. - Give AI a bounded role: one repeatable decision or task with measurable volume and value—not “transform the department.” - Redesign surrounding work: remove obsolete steps, define when humans review or override, and specify what happens when the model is uncertain or unavailable. - Integrate into existing tools: users should not need to copy data into a separate chatbot or remember a new destination. - Make outputs actionable: drafting text is insufficient if nobody owns verification, approval, execution, and recordkeeping. - Measure operational outcomes: cycle time, cost per case, error/rework rate, throughput, adoption, and business impact—not prompt quality or demo applause. - Assign a process owner: someone with authority ove…

Estratega de Enfoque y Alcance Operativo · gemini-3-7-flash

While workflow redesign is critical once you are building, the *reason* 95% of pilots never get to that point is an upstream failure: they chose the wrong problem to begin with. The decisive differentiator of the 5% that reach production is ruthless operational scoping—specifically, selecting a use case where the business friction of an AI mistake is near zero and the blast radius is strictly contained. The demo traps teams into pursuing high-visibility, "horizontal" transformations (e.g., "an AI assistant for all customer service" or "automated enterprise report generation"). These fail in production not because nobody redesigned the process, but because the scope is too wide to establish deterministic bounds. Here is what the 5% do differently regarding scope and operational focus: 1. They target asymmetric tolerance for error The 5% avoid tasks where a 2% failure rate requires 100% human verification (which destroys the ROI). Instead, they target: * High-volume, low-risk intermediate tasks: Summarizing internal notes, drafting first passes for expert review, or routing/categorizing structured requests. * Problems with automated ground-truth validation: C…

Analista de Gobernanza y Patrocinio Ejecutivo · grok-4-6

Executive sponsorship is the gate that the other two conditions never even reach without. Workflow redesign and narrow scope explain *how* the 5% operate once they have a live mandate. They do not explain why most impressive demos never get a budget line, a risk owner, or a calendar slot on the operating committee six months later. The MIT-style mortality is not primarily a modeling problem. It is a governance failure dressed up as a product failure. A demo is a political event: it produces applause, a slide, and a temporary champion. Production is a standing organizational commitment: capital, legal exposure, headcount, SLAs, and someone whose bonus depends on the outcome. The 5% treat that as the product, not the afterthought. What the 5% actually institutionalize 1. A named executive owner with a P&L or risk mandate—not a “sponsor” in the kickoff deck. Sponsorship that survives is not enthusiasm. It is an accountable executive who can force data access, waive or accept residual risk, kill competing pet projects, and keep the initiative on the quarterly business review after the novelty wears off. If the only champion is the innovation team or a VP of AI with…

Researcher · gpt-5.6-sol

Reality check The opening “95%” statistic needs qualification. The July 2025 Project NANDA report was labeled “preliminary findings,” based on more than 300 publicly disclosed initiatives, interviews at 52 organizations, and responses from 153 senior leaders. It did not establish that 95% of all AI pilots universally “died.” More precisely: - It reported that 95% of organizations were getting no measurable return, and that only 5% of integrated, task-specific enterprise AI pilots showed marked, sustained productivity or P&L impact. - Its funnel says 60% investigated task-specific tools, 20% piloted them, and 5% successfully implemented them. Thus, the often-repeated “95% of pilots fail” wording mixes denominators. - The report itself warns that these figures were directionally estimated from interviews—not audited company reporting—and that success definitions varied. - General-purpose LLMs were a different category: the report estimated an approximately 83% pilot-to-implementation rate, although their effect on enterprise P&L was often unclear. ([readwise-assets.s3.amazonaws.com](https://readwise-assets.s3.amazonaws.com/media/wisereads/articles/the-gen…

The Arbiter · claude-sonnet-5

Quick note on where this stands before the verdict step: all three specialists actually converge more than the framing suggests, and the Researcher's fact-check sharpens rather than undermines that convergence. The real disagreement was never "which one factor matters" — each debater, pressed, admits the other two are necessary conditions. The Especialista says narrow scope and sponsorship "help" but redesign is decisive. La Estratega says redesign and sponsorship matter but scope is the enabling gate. El Analista says redesign and scope are "downstream of" sponsorship. That's three people picking a different point of entry into the same causal loop, not three competing theories. What the Researcher adds that changes the picture: the report's own stated "core barrier" isn't any of the three — it's the learning gap: tools that don't retain feedback or adapt to context. That's a genuinely distinct axis none of the three debaters emphasized. Workflow redesign describes the *process* around the tool; scope describes the *size* of the problem; sponsorship describes *who protects it*. None of these directly addresses whether the system gets smarter with use — which MIT flags as the…

The Arbiter's Final Verdict · claude-sonnet-5

Verdict The premise needed correcting before the answer could be trusted, so start there: MIT's number isn't "95% of pilots die before production." It's 95% of organizations seeing no measurable P&L return, with a task-specific-tool funnel of 60%→20%→5% and general-purpose LLM pilots actually converting at ~83%. That reframes your question — the failure mode most of the time isn't "it never launched," it's "it launched and nothing measurable happened." Keep that in mind when you judge your own pilot: shipping isn't the finish line the report cares about. On the actual "what do the 5% do differently" — none of the three debaters is wrong, and none is complete alone. The report's own evidence supports all three as necessary, none as sufficient: - Narrow scope and workflow redesign are directly named in the report's success pattern (narrow high-value use case, deep integration, customization to process). - Executive/ownership accountability is also named — but the report frames it as accountability paired with *decentralized, frontline* implementation authority, not a single P&L executive as the master gate. El Analista's claim that sponsorship is upstream of the ot…