AI 파일럿의 95% 는 왜 실패하고, 그 5% 는 실제로 무엇을 하는가

'AI 파일럿의 95% 가 실패한다'는 말은 MIT 를 잘못 읽은 것이다 : 대부분은 출시는 되었으나 측정 가능한 수익을 전혀 보여주지 못했을 뿐이다. 그 5% 가 무엇을 다르게 하는지를 둘러싼 논쟁은 저평가된 한 축, 즉 도구가 쓰일수록 더 똑똑해지는가로 자꾸 되돌아왔다.

AI와 사회 · 2026-09-02

"AI 파일럿의 95% 가 실패한다"는 기업 AI 에서 가장 자주 인용되는 수치 중 하나다. 그런데 이 수치는 대개 잘못 읽힌다. 출처는 2025년 MIT 와 연계된 보고서인데, 실제로 말하는 바는 그보다 좁다. 조직의 95퍼센트가 자사의 AI 지출에서 측정 가능한 수익을 전혀 보지 못하고 있었다는 것이다. 이는 파일럿의 95퍼센트가 실제 사용 단계에 이르기 전에 무산된다는 주장과는 다르다.

이 구분은 우리가 던져야 할 질문 자체를 바꾼다. 대부분의 기업에서 실패는 도구가 끝내 출시되지 못했다는 뜻이 아니다. 도구는 출시되었는데 측정 가능한 어떤 일도 일어나지 않았다는 뜻이다. 보고서가 그린 과제 특화형 도구의 깔때기를 보면, 60퍼센트가 검토 단계에 머물고, 20퍼센트가 파일럿까지 가며, 5퍼센트가 성공적인 구현에 이른다. 범용 챗 도구는 약 83퍼센트가 실제 사용으로 이어져 이보다 훨씬 높았지만, 이익에 미친 효과는 불분명한 경우가 많았다.

세 개의 자리, 하나의 순환

이 질문을 뜯어보기 위해 Polora 는 같은 문제를 여러 AI 모델 앞에 놓고, 각 모델에 서로 다른 관점을 부여했다. 하나는 워크플로를, 하나는 범위를, 하나는 거버넌스를 논하게 했다.

워크플로를 논한 모델 gpt-5.6-sol 은, 데모란 잘 짜인 조건에서 모델이 인상적인 답을 내놓을 수 있다는 것을 보여줄 뿐이라고 봤다. 프로덕션은 다르다. 도구를 중심으로 실제 업무가 처리되는 방식을 다시 설계해야 한다는 것이다. 범위를 논한 모델 gemini-3-7-flash 는, 잘못된 문제는 대개 그보다 앞선 단계에서 선택된다고 반박했다. 5퍼센트에 드는 조직은 모델이 가장 성능이 나쁜 순간에도 운영상 감당할 수 있을 만큼 작은 과제를 고르고, 성공 시나리오보다 실패 대비책을 먼저 설계한다는 것이다. 거버넌스를 논한 모델 grok-4-6 은 소유권을 앞세웠다. 이 모델은 데모를 일종의 정치적 사건으로 봤다. 예산과 위험 승인, 예외 처리 경로를 쥔 책임 있는 주인이 없으면, 데모는 결재의 문턱을 넘지 못하고 사라진다는 것이다.

더 캐묻자 셋 중 누구도 자신의 요인을 충분조건으로 여기지 않았다. 각자 나머지 둘이 필요조건임을 인정했다. 그들은 서로 경쟁하는 세 개의 이론이 아니라, 같은 순환으로 들어가는 세 개의 입구였다.

워크플로 도구를 중심으로 실제 업무가 처리되는 방식을 다시 설계해야 한다는 것이다. · 범위 잘못된 문제는 대개 그보다 앞선 단계에서 선택된다 · 거버넌스 예산과 위험 승인, 예외 처리 경로를 쥔 책임 있는 주인이 없으면, 데모는 결재의 문턱을 넘지 못하고 사라진다는 것이다. · 그들은 서로 경쟁하는 세 개의 이론이 아니라, 같은 순환으로 들어가는 세 개의 입구였다.
워크플로 도구를 중심으로 실제 업무가 처리되는 방식을 다시 설계해야 한다는 것이다. · 범위 잘못된 문제는 대개 그보다 앞선 단계에서 선택된다 · 거버넌스 예산과 위험 승인, 예외 처리 경로를 쥔 책임 있는 주인이 없으면, 데모는 결재의 문턱을 넘지 못하고 사라진다는 것이다. · 그들은 서로 경쟁하는 세 개의 이론이 아니라, 같은 순환으로 들어가는 세 개의 입구였다.

아무도 앞세우지 않은 축

논쟁을 다시 짠 것은 셋 중 누구도 중심에 두지 않았던 네 번째 축이었다. 연구자의 자리에 앉은 모델 gpt-5.6-sol 은 보고서로 되돌아가, 거기서 밝힌 핵심 장벽이 프로세스도 정치도 아니라고 짚었다. 보고서는 이를 학습 격차라 부른다. 시스템이 피드백을 쌓아가며 맥락에 적응하는가, 아니면 데모 수준에 그대로 멈춰 있는가의 문제다.

이는 나머지 세 질문과는 다른 물음이다. 범위는 문제의 크기이고, 워크플로는 도구를 둘러싼 프로세스이며, 후원은 누가 그것을 지켜주는가이다. 그중 어느 것도 도구가 쓰일수록 나아지는가를 묻지 않는다. 범위가 좁게 잡히고, 잘 후원되고, 깊이 통합되었더라도, 쓸수록 더 똑똑해지지는 않는 도구는 여전히 95퍼센트 쪽에 남는다.

당신의 파일럿을 위한 시험

중재를 맡은 모델 claude-sonnet-5 은 네 요인을 하나의 위계로 세우기를 거부했다. 보고서가 이들을 사다리가 아니라 서로 맞물려 작동하는 묶음으로 기술하기 때문이다. 자신의 파일럿을 스스로 점검하려는 사람을 위해, 이 모델은 마지막으로 네 가지 질문을 제시했다. 실패해도 손실이 크지 않을 만큼 범위가 충분히 작은가. 실제 권한을 가진 누군가가 예산과 예외 처리 경로를 쥐고 있는가. AI 가 작동할 때 워크플로가 실제로 바뀌는가. 그리고 시스템이 쓰일수록 더 나아지는가.

그 마지막 질문이야말로 이 논쟁이 과소평가해 온 것이다. 중재자가 언급한 대로, 보고서의 데이터는 앞의 세 가지가 탄탄해도 이 항목이 약한 조직은 95퍼센트 구간에 머문다는 것을 보여준다. 별개의 MIT 연구도 같은 방향을 가리킨다. 이 연구는 성공적인 도입이 표적을 좁힌 소규모 변화라고 설명한다. 일하는 사람을, 과제를 직접 수행하는 자리에서 도구를 감독하는 자리로 옮기는 변화, 즉 모델을 그냥 설치하는 데 그치지 않고 일 자체를 다시 설계하는 것이다.

범위 실패해도 손실이 크지 않을 만큼 범위가 충분히 작은가. · 후원 실제 권한을 가진 누군가가 예산과 예외 처리 경로를 쥐고 있는가. · 워크플로 AI 가 작동할 때 워크플로가 실제로 바뀌는가. · 학습 격차 그리고 시스템이 쓰일수록 더 나아지는가. · 보고서가 이들을 사다리가 아니라 서로 맞물려 작동하는 묶음으로 기술하기 때문이다.
범위 실패해도 손실이 크지 않을 만큼 범위가 충분히 작은가. · 후원 실제 권한을 가진 누군가가 예산과 예외 처리 경로를 쥐고 있는가. · 워크플로 AI 가 작동할 때 워크플로가 실제로 바뀌는가. · 학습 격차 그리고 시스템이 쓰일수록 더 나아지는가. · 보고서가 이들을 사다리가 아니라 서로 맞물려 작동하는 묶음으로 기술하기 때문이다.

사실이 아니라 직관으로만 간직할 숫자들

논의가 이어지는 동안 몇 가지 눈에 띄는 수치가 등장했다. 모델은 프로덕션 시스템에서 20퍼센트에 불과하다는 것, 2퍼센트의 오류율이 100퍼센트의 사람 검증을 요구한다는 것, 범위를 좁히면 18개월짜리 작업이 3주로 줄어든다는 것. 연구를 맡은 모델은 이들 각각이 MIT 연구가 실제로 확립한 사실이라기보다 실무자의 그럴듯한 어림짐작에 가깝다고 봤다. 직관을 잡는 데는 쓸모가 있지만, 연구 결과로 인용할 수 있는 것은 아니다.

과제 특화형 도구가 각 단계에 도달한 비율 · 검토 단계에 머묾 60% · 파일럿까지 감 20% · 성공적 구현 5%
과제 특화형 도구가 각 단계에 도달한 비율 · 검토 단계에 머묾 60% · 파일럿까지 감 20% · 성공적 구현 5%
AI 파일럿의 95% 는 왜 실패하고, 그 5% 는 실제로 무엇을 하는가AI 파일럿의 95% 는 왜 실패하고, 그 5% 는 실제로 무엇을 하는가"AI 파일럿의 95% 가 실패한다"는 기업 AI 에서 가장 자주 인용되는 수치 중 하나다. 그런데 이 수치는 대개 잘못 읽힌다. 출처는 2025년 MIT 와 연계된 보고서인데, 실제로 말하는 바는 그보다 좁다. 조직의 95퍼센트가 자사의 AI 지출에서 측정 가능한 수익을 전혀 보지 못하고 있었다는 것이다. · ※ 파일럿 : 본격 도입 전에 소규모로 시험 삼아 돌려 보는 시범 프로젝트.세 개의 자리, 하나의 순환 · 워크플로 None, 범위 None, 거버넌스 None · ※ 프로덕션 : 데모가 아니라 실제 사용자가 매일 쓰는 실운영 환경.아무도 앞세우지 않은 축당신의 파일럿을 위한 시험 · 범위 None, 후원 None, 워크플로 None, 학습 격차 None논의가 이어지는 동안 몇 가지 눈에 띄는 수치가 등장했다. 모델은 프로덕션 시스템에서 20퍼센트에 불과하다는 것, 2퍼센트의 오류율이 100퍼센트의 사람 검증을 요구한다는 것, 범위를 좁히면 18개월짜리 작업이 3주로 줄어든다는 것. · 과제 특화형 도구가 각 단계에 도달한 비율출처 readwise-assets.s3.amazonaws.com · The GenAI Divide : State of AI in Business 2025 (MIT NANDA), searchyour.ai · State of AI in Business 2025, full report PDF, mitsloan.mit.edu · Practical AI implementation, MIT Sloan

Your AI pilot wowed everyone in the demo and then quietly died before it reached real use, the fate of roughly 95% of them by MIT's count. What do the 5% that make it to production actually do differently: redesign the workflow, narrow the scope, keep executive sponsorship, or something the demo hides?

Especialista en Rediseño de Flujos de Trabajo · gpt-5.6-sol

The surviving 5% do all three—but workflow redesign is the decisive difference. A demo proves that a model can generate an impressive answer under curated conditions. Production requires a system that changes how work gets completed every day. Successful teams therefore start with the operating process, not the model: - Map the real workflow: inputs, handoffs, approvals, exceptions, systems of record, and accountability. - Give AI a bounded role: one repeatable decision or task with measurable volume and value—not “transform the department.” - Redesign surrounding work: remove obsolete steps, define when humans review or override, and specify what happens when the model is uncertain or unavailable. - Integrate into existing tools: users should not need to copy data into a separate chatbot or remember a new destination. - Make outputs actionable: drafting text is insufficient if nobody owns verification, approval, execution, and recordkeeping. - Measure operational outcomes: cycle time, cost per case, error/rework rate, throughput, adoption, and business impact—not prompt quality or demo applause. - Assign a process owner: someone with authority ove…

Estratega de Enfoque y Alcance Operativo · gemini-3-7-flash

While workflow redesign is critical once you are building, the *reason* 95% of pilots never get to that point is an upstream failure: they chose the wrong problem to begin with. The decisive differentiator of the 5% that reach production is ruthless operational scoping—specifically, selecting a use case where the business friction of an AI mistake is near zero and the blast radius is strictly contained. The demo traps teams into pursuing high-visibility, "horizontal" transformations (e.g., "an AI assistant for all customer service" or "automated enterprise report generation"). These fail in production not because nobody redesigned the process, but because the scope is too wide to establish deterministic bounds. Here is what the 5% do differently regarding scope and operational focus: 1. They target asymmetric tolerance for error The 5% avoid tasks where a 2% failure rate requires 100% human verification (which destroys the ROI). Instead, they target: * High-volume, low-risk intermediate tasks: Summarizing internal notes, drafting first passes for expert review, or routing/categorizing structured requests. * Problems with automated ground-truth validation: C…

Analista de Gobernanza y Patrocinio Ejecutivo · grok-4-6

Executive sponsorship is the gate that the other two conditions never even reach without. Workflow redesign and narrow scope explain *how* the 5% operate once they have a live mandate. They do not explain why most impressive demos never get a budget line, a risk owner, or a calendar slot on the operating committee six months later. The MIT-style mortality is not primarily a modeling problem. It is a governance failure dressed up as a product failure. A demo is a political event: it produces applause, a slide, and a temporary champion. Production is a standing organizational commitment: capital, legal exposure, headcount, SLAs, and someone whose bonus depends on the outcome. The 5% treat that as the product, not the afterthought. What the 5% actually institutionalize 1. A named executive owner with a P&L or risk mandate—not a “sponsor” in the kickoff deck. Sponsorship that survives is not enthusiasm. It is an accountable executive who can force data access, waive or accept residual risk, kill competing pet projects, and keep the initiative on the quarterly business review after the novelty wears off. If the only champion is the innovation team or a VP of AI with…

Researcher · gpt-5.6-sol

Reality check The opening “95%” statistic needs qualification. The July 2025 Project NANDA report was labeled “preliminary findings,” based on more than 300 publicly disclosed initiatives, interviews at 52 organizations, and responses from 153 senior leaders. It did not establish that 95% of all AI pilots universally “died.” More precisely: - It reported that 95% of organizations were getting no measurable return, and that only 5% of integrated, task-specific enterprise AI pilots showed marked, sustained productivity or P&L impact. - Its funnel says 60% investigated task-specific tools, 20% piloted them, and 5% successfully implemented them. Thus, the often-repeated “95% of pilots fail” wording mixes denominators. - The report itself warns that these figures were directionally estimated from interviews—not audited company reporting—and that success definitions varied. - General-purpose LLMs were a different category: the report estimated an approximately 83% pilot-to-implementation rate, although their effect on enterprise P&L was often unclear. ([readwise-assets.s3.amazonaws.com](https://readwise-assets.s3.amazonaws.com/media/wisereads/articles/the-gen…

The Arbiter · claude-sonnet-5

Quick note on where this stands before the verdict step: all three specialists actually converge more than the framing suggests, and the Researcher's fact-check sharpens rather than undermines that convergence. The real disagreement was never "which one factor matters" — each debater, pressed, admits the other two are necessary conditions. The Especialista says narrow scope and sponsorship "help" but redesign is decisive. La Estratega says redesign and sponsorship matter but scope is the enabling gate. El Analista says redesign and scope are "downstream of" sponsorship. That's three people picking a different point of entry into the same causal loop, not three competing theories. What the Researcher adds that changes the picture: the report's own stated "core barrier" isn't any of the three — it's the learning gap: tools that don't retain feedback or adapt to context. That's a genuinely distinct axis none of the three debaters emphasized. Workflow redesign describes the *process* around the tool; scope describes the *size* of the problem; sponsorship describes *who protects it*. None of these directly addresses whether the system gets smarter with use — which MIT flags as the…

The Arbiter's Final Verdict · claude-sonnet-5

Verdict The premise needed correcting before the answer could be trusted, so start there: MIT's number isn't "95% of pilots die before production." It's 95% of organizations seeing no measurable P&L return, with a task-specific-tool funnel of 60%→20%→5% and general-purpose LLM pilots actually converting at ~83%. That reframes your question — the failure mode most of the time isn't "it never launched," it's "it launched and nothing measurable happened." Keep that in mind when you judge your own pilot: shipping isn't the finish line the report cares about. On the actual "what do the 5% do differently" — none of the three debaters is wrong, and none is complete alone. The report's own evidence supports all three as necessary, none as sufficient: - Narrow scope and workflow redesign are directly named in the report's success pattern (narrow high-value use case, deep integration, customization to process). - Executive/ownership accountability is also named — but the report frames it as accountability paired with *decentralized, frontline* implementation authority, not a single P&L executive as the master gate. El Analista's claim that sponsorship is upstream of the ot…