Por qué fracasa el 95 % de los pilotos de IA, y qué hace de verdad el 5 %
"El 95 % de los pilotos de IA fracasa" malinterpreta al MIT : la mayoría se lanza y simplemente no muestra ningún retorno medible. El debate sobre qué hace distinto ese 5 % volvía una y otra vez a un eje infravalorado : si la herramienta se vuelve más inteligente con el uso.
IA y sociedad · 2026-09-02
«El 95 % de los pilotos de IA fracasa» es una de las cifras más repetidas en la IA empresarial. También es una de las peor leídas. El dato procede de un informe de 2025 vinculado al MIT, y lo que el informe dice en realidad es más acotado. El noventa y cinco por ciento de las organizaciones no veía ningún retorno medible de su gasto en IA. Eso no es lo mismo que decir que el 95 % de los pilotos muere antes de llegar a un uso real.
La distinción cambia la pregunta que deberías hacerte. Para la mayoría de las empresas, el fracaso no es que la herramienta nunca se lanzara. Es que se lanzó y no pasó nada medible. En el embudo que describe el informe para las herramientas pensadas para una tarea concreta, el 60 % se queda en la fase de investigación, el 20 % llega al piloto y solo el 5 % alcanza una implementación exitosa. Las herramientas de chat de propósito general se adoptaron a un ritmo mucho mayor, en torno al 83 %, aunque su efecto sobre el beneficio muchas veces no quedara claro.
Para desmontar la pregunta, Polora planteó el mismo problema a varios modelos de IA y a cada uno le asignó un punto de vista distinto. Uno argumentó desde el flujo de trabajo, otro desde el alcance, otro desde la gobernanza.
El modelo que defendía el flujo de trabajo, gpt-5.6-sol, sostuvo que una demostración solo prueba que un modelo puede producir una respuesta impresionante en condiciones cuidadas. La producción es otra cosa : obliga a rediseñar cómo se hace realmente el trabajo alrededor de la herramienta. El modelo que defendía el alcance, gemini-3-7-flash, replicó que el problema equivocado suele elegirse más arriba. El 5 % elige una tarea lo bastante pequeña como para que el peor día del modelo siga siendo aceptable en la operación, y diseña el plan de contingencia antes que el camino ideal. El modelo que defendía la gobernanza, grok-4-6, puso por delante la cuestión de quién es el dueño. Trató la demostración como un acontecimiento político, algo que muere en los pasillos a menos que un responsable que rinda cuentas controle el presupuesto, la aprobación del riesgo y la vía de excepción.
Al presionarlos, ninguno de los tres trató su propio factor como suficiente. Cada uno admitió que los otros dos eran condiciones necesarias. Eran tres puertas de entrada al mismo bucle, no tres teorías rivales.
El eje que nadie puso sobre la mesa
Lo que reformuló el debate fue un cuarto eje que ninguno de los tres había puesto en el centro. Un modelo en el asiento del investigador, gpt-5.6-sol, volvió al informe y señaló que la barrera central que este declara no es ni el proceso ni la política. El informe la llama la brecha de aprendizaje : la cuestión de si un sistema retiene la retroalimentación y se adapta al contexto, o si se queda congelado en la calidad de una demostración.
Esa es una pregunta distinta de las otras tres. El alcance es el tamaño del problema, el flujo de trabajo es el proceso alrededor de la herramienta, el patrocinio es quién la protege. Ninguno pregunta si la herramienta mejora con el uso. Una herramienta de alcance estrecho, bien patrocinada y profundamente integrada que nunca se vuelve más inteligente sigue cayendo en el 95 %.
Una prueba para tu propio piloto
El modelo moderador, claude-sonnet-5, se negó a ordenar los cuatro factores en una jerarquía, porque el informe los describe como un conjunto que interactúa entre sí y no como una escalera. Su prueba final para quien quiera juzgar su propio piloto se redujo a cuatro preguntas. ¿Es el alcance lo bastante pequeño como para que el fracaso salga barato? ¿Hay alguien con autoridad real que sea dueño del presupuesto y de la vía de excepción? ¿Cambia de verdad el flujo de trabajo cuando la IA funciona? ¿Y mejora el sistema cuanto más se usa?
Esa última pregunta es la que este debate había subestimado. Los datos del informe, señaló el moderador, muestran que las organizaciones débiles en ese punto se quedan en el grupo del 95 % incluso cuando los tres primeros factores son sólidos. Otro trabajo distinto del MIT apunta en la misma dirección. Describe la adopción exitosa como un cambio dirigido y a pequeña escala que traslada al trabajador de hacer la tarea a supervisar la herramienta, un rediseño del puesto y no la simple instalación de un modelo.
Cifras que conviene guardar como intuición, no como hecho
Por el camino se intercambiaron varias cifras contundentes. Que el modelo es solo el 20 % de un sistema en producción. Que una tasa de error del 2 % obliga a una verificación humana del 100 %. Que estrechar el alcance convierte un esfuerzo de 18 meses en tres semanas. El modelo investigador marcó cada una de ellas como heurísticas de operador plausibles, no como algo que la investigación del MIT haya establecido de verdad. Sirven como instinto. No se pueden citar como hallazgos.
Embudo de las herramientas de IA pensadas para una tarea concreta. · Investigación 60 % · Piloto 20 % · Implementación exitosa 5 %
Your AI pilot wowed everyone in the demo and then quietly died before it reached real use, the fate of roughly 95% of them by MIT's count. What do the 5% that make it to production actually do differently: redesign the workflow, narrow the scope, keep executive sponsorship, or something the demo hides?
Especialista en Rediseño de Flujos de Trabajo · gpt-5.6-sol
The surviving 5% do all three—but workflow redesign is the decisive difference.
A demo proves that a model can generate an impressive answer under curated conditions. Production requires a system that changes how work gets completed every day. Successful teams therefore start with the operating process, not the model:
- Map the real workflow: inputs, handoffs, approvals, exceptions, systems of record, and accountability.
- Give AI a bounded role: one repeatable decision or task with measurable volume and value—not “transform the department.”
- Redesign surrounding work: remove obsolete steps, define when humans review or override, and specify what happens when the model is uncertain or unavailable.
- Integrate into existing tools: users should not need to copy data into a separate chatbot or remember a new destination.
- Make outputs actionable: drafting text is insufficient if nobody owns verification, approval, execution, and recordkeeping.
- Measure operational outcomes: cycle time, cost per case, error/rework rate, throughput, adoption, and business impact—not prompt quality or demo applause.
- Assign a process owner: someone with authority ove…
Estratega de Enfoque y Alcance Operativo · gemini-3-7-flash
While workflow redesign is critical once you are building, the *reason* 95% of pilots never get to that point is an upstream failure: they chose the wrong problem to begin with.
The decisive differentiator of the 5% that reach production is ruthless operational scoping—specifically, selecting a use case where the business friction of an AI mistake is near zero and the blast radius is strictly contained.
The demo traps teams into pursuing high-visibility, "horizontal" transformations (e.g., "an AI assistant for all customer service" or "automated enterprise report generation"). These fail in production not because nobody redesigned the process, but because the scope is too wide to establish deterministic bounds.
Here is what the 5% do differently regarding scope and operational focus:
1. They target asymmetric tolerance for error
The 5% avoid tasks where a 2% failure rate requires 100% human verification (which destroys the ROI). Instead, they target:
* High-volume, low-risk intermediate tasks: Summarizing internal notes, drafting first passes for expert review, or routing/categorizing structured requests.
* Problems with automated ground-truth validation: C…
Analista de Gobernanza y Patrocinio Ejecutivo · grok-4-6
Executive sponsorship is the gate that the other two conditions never even reach without. Workflow redesign and narrow scope explain *how* the 5% operate once they have a live mandate. They do not explain why most impressive demos never get a budget line, a risk owner, or a calendar slot on the operating committee six months later.
The MIT-style mortality is not primarily a modeling problem. It is a governance failure dressed up as a product failure. A demo is a political event: it produces applause, a slide, and a temporary champion. Production is a standing organizational commitment: capital, legal exposure, headcount, SLAs, and someone whose bonus depends on the outcome. The 5% treat that as the product, not the afterthought.
What the 5% actually institutionalize
1. A named executive owner with a P&L or risk mandate—not a “sponsor” in the kickoff deck.
Sponsorship that survives is not enthusiasm. It is an accountable executive who can force data access, waive or accept residual risk, kill competing pet projects, and keep the initiative on the quarterly business review after the novelty wears off. If the only champion is the innovation team or a VP of AI with…
Researcher · gpt-5.6-sol
Reality check
The opening “95%” statistic needs qualification. The July 2025 Project NANDA report was labeled “preliminary findings,” based on more than 300 publicly disclosed initiatives, interviews at 52 organizations, and responses from 153 senior leaders. It did not establish that 95% of all AI pilots universally “died.” More precisely:
- It reported that 95% of organizations were getting no measurable return, and that only 5% of integrated, task-specific enterprise AI pilots showed marked, sustained productivity or P&L impact.
- Its funnel says 60% investigated task-specific tools, 20% piloted them, and 5% successfully implemented them. Thus, the often-repeated “95% of pilots fail” wording mixes denominators.
- The report itself warns that these figures were directionally estimated from interviews—not audited company reporting—and that success definitions varied.
- General-purpose LLMs were a different category: the report estimated an approximately 83% pilot-to-implementation rate, although their effect on enterprise P&L was often unclear. ([readwise-assets.s3.amazonaws.com](https://readwise-assets.s3.amazonaws.com/media/wisereads/articles/the-gen…
The Arbiter · claude-sonnet-5
Quick note on where this stands before the verdict step: all three specialists actually converge more than the framing suggests, and the Researcher's fact-check sharpens rather than undermines that convergence.
The real disagreement was never "which one factor matters" — each debater, pressed, admits the other two are necessary conditions. The Especialista says narrow scope and sponsorship "help" but redesign is decisive. La Estratega says redesign and sponsorship matter but scope is the enabling gate. El Analista says redesign and scope are "downstream of" sponsorship. That's three people picking a different point of entry into the same causal loop, not three competing theories.
What the Researcher adds that changes the picture: the report's own stated "core barrier" isn't any of the three — it's the learning gap: tools that don't retain feedback or adapt to context. That's a genuinely distinct axis none of the three debaters emphasized. Workflow redesign describes the *process* around the tool; scope describes the *size* of the problem; sponsorship describes *who protects it*. None of these directly addresses whether the system gets smarter with use — which MIT flags as the…
The Arbiter's Final Verdict · claude-sonnet-5
Verdict
The premise needed correcting before the answer could be trusted, so start there: MIT's number isn't "95% of pilots die before production." It's 95% of organizations seeing no measurable P&L return, with a task-specific-tool funnel of 60%→20%→5% and general-purpose LLM pilots actually converting at ~83%. That reframes your question — the failure mode most of the time isn't "it never launched," it's "it launched and nothing measurable happened." Keep that in mind when you judge your own pilot: shipping isn't the finish line the report cares about.
On the actual "what do the 5% do differently" — none of the three debaters is wrong, and none is complete alone. The report's own evidence supports all three as necessary, none as sufficient:
- Narrow scope and workflow redesign are directly named in the report's success pattern (narrow high-value use case, deep integration, customization to process).
- Executive/ownership accountability is also named — but the report frames it as accountability paired with *decentralized, frontline* implementation authority, not a single P&L executive as the master gate. El Analista's claim that sponsorship is upstream of the ot…