Skip to content
HIVO

Enterprise AI

Amit Amir · Co-founder & CEO

Why enterprise AI pilots stall

Most enterprise AI pilots stall because the pilot borrows something production cannot. A person in the room supplies the organizational context the model does not have. McKinsey found in August 2026 that 80 percent of respondents say AI has improved their individual productivity, while only 37 percent can attribute any EBIT impact to it at all.

Layered panes of frosted glass receding into soft focus, lit from behind in warm cream light.

If you have sat through a board discussion about artificial intelligence in the last year, you have heard the number. Ninety five percent of AI pilots fail. It arrives without a source, it settles the room, and it is repeated by people arguing for more investment and by people arguing for less.

It is not a reliable number. It is worth understanding exactly how it is unreliable, because the error points at the thing that actually goes wrong between a pilot and a rollout.

Is the 95 percent AI failure statistic real?

The figure comes from a working paper called "The GenAI Divide: State of AI in Business 2025", published by four authors associated with Project NANDA at the MIT Media Lab. The document is cover dated July 2025, began circulating in August 2025, and is labelled "Preliminary Findings". The copy that circulated most widely carried the filename v0.1_State_of_AI_in_Business_2025_Report.pdf; the document itself carries no version number.

Two things are worth saying plainly before anything else. The paper has not been retracted, and nothing here alleges bad faith. It is a preliminary working paper that was read as though it were a finished study, and most of the damage was done in transit rather than at the source.

The paper's own headline finding is that 95 percent of organizations are getting zero return, which is a claim about a different population from the one everybody quotes. The paper then makes the same slip itself, on the next line and again four pages later, calling its own chart a 95 percent failure rate for enterprise AI solutions.

That distinction is the whole problem, and the paper contains the arithmetic that shows it.

What the paper actually counted

The adoption exhibit reports a funnel for task specific generative AI tools. Sixty percent of organizations evaluated such tools. Twenty percent reached pilot stage. Five percent reached successful implementation.

Read those three numbers in order. The five percent is a share of all organizations surveyed, not a share of organizations that ran a pilot. Among the organizations that actually piloted something, five in twenty reached production. That is twenty five percent, not five percent.

25%
of organizations that ran a pilot reached successful implementation, on the paper's own funnel of 60 percent evaluated, 20 percent piloted, 5 percent implemented.
MIT NANDA, The GenAI Divide, July 2025

Roughly four fifths of the famous ninety five percent is made up of organizations that never attempted a pilot at all. Counting a company that never started as a company that failed is not an interpretation you can argue with. It is a division error, and anyone repeating the headline is publishing it.

What the paper counted as success

The second problem is the definition. The exhibit note states it directly: successfully implemented tools are ones that "users or executives have remarked as causing a marked and sustained productivity and/or P&L impact".

Success is a remark in an interview. The same exhibit is footnoted with the warning that the figures are "directionally accurate based on individual interviews rather than official company reporting". The methodology section states that impact was measured six months after the pilot, and lists as an explicit limitation that six months "may be insufficient", "potentially understating success rates".

No measurable P&L impact does not mean anyone measured P&L. It means nobody volunteered a strong enough remark inside six months.

The paper states its basis as "a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders collected across four major industry conferences". The 300 are public announcements rather than companies studied, and the 153 are people recruited at conferences rather than a random sample of enterprises.

How the number got worse on the way to you

Most people did not meet this statistic in the paper. They met it in Fortune, on 18 August 2025, which described the study as "150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments".

As reportedAs stated in the paper
Interviews150 interviews with leadersrepresentatives from 52 organizations
Surveya survey of 350 employees153 senior leaders, at four conferences
Deployments300 public AI deploymentsover 300 publicly disclosed initiatives
Fortune, 18 August 2025, against the paper's own front matter. Fortune revised the piece on 16 December 2025 and, as of September 2026, no correction has been appended.

The interview base was overstated roughly threefold and the survey roughly twofold, and "employees" became the wrong word for a group of senior leaders recruited at industry conferences. The numbers 150 and 350 appear nowhere in the paper. Every article that repeats them is repeating Fortune, not the research.

The paper was also hard to read at the moment it went viral. On the 80,000 Hours podcast episode "The story behind the bad AI stat that moved markets and misled millions", published 28 April 2026, Rob Wiblin noted that "the article wasn't even available to journalists around the world trying to cover the story".

The methodological objection has been made by named academics. LeadDev, in November 2025, put Wharton professor Kevin Werbach's objection this way: "the 95% figure is presented in one sentence, but the authors offer no detail on how they came up with it".

The paper carries no conflict of interest statement beyond a line stating that the views are the authors' own. It also names NANDA, the authors' own project, among the infrastructure it describes as the way forward. That is worth knowing when a paper concludes that the cure is the category its authors work in.

Where can I find a figure I can actually use?

There is a properly sampled measurement of the same gap, and it is not hiding. McKinsey's "The state of AI in 2026: On the road to ROI" was published on 25 August 2026, surveyed between 4 May and 8 June 2026, and drew 1,719 responses across 97 nations.

80%
of respondents say AI has improved their individual productivity.
McKinsey, The state of AI in 2026: On the road to ROI, 25 August 2026, n=1,719
37%
attribute any EBIT impact to AI, essentially unchanged from a year earlier.
McKinsey, The state of AI in 2026: On the road to ROI, 25 August 2026, n=1,719

A documented field window, a published sample size, a named institution that stands behind the work. The distance between those two numbers is the pilot to production gap, measured properly. The individual gain is close to universal. Running the business on it is rare, and over a year it did not become less rare.

McKinsey reports respondents, not organizations, and its own verb for enterprise reach is progressive: 44 percent say AI "is scaling" across their enterprise, up from 38 percent a year ago. Scaling is not the same claim as scaled, and secondary coverage swaps them constantly.

So what actually goes wrong between the pilot and the rollout?

A pilot runs on a narrow slice of the business, with somebody in the room who knows the slice. That person chose the example. They know which of the four systems holds the true version of the record. They know which project the invoice belongs to, which approval it needs, what state the job is in, and who owns it. If the answer had been wrong, they would have caught it.

None of that was written down. It was supplied, live, by a human being who understood the organization.

A pilot succeeds because a person supplies the organizational context the model does not have. Production fails because at scale, nobody can.

At ten users the supply is invisible. At a thousand it is the constraint. This is why the failure looks like a technology failure and is not one: the model performed the same in both cases, and what changed was whether anyone was still filling the gap.

What is missing is a model of the organization, not a better model

The systems already hold the data. The ERP holds the transactions, the document system holds the files, and the spreadsheets hold everything that never found a home. What none of them holds is the structure: what the central operating object is, and what the organization has agreed is true about it - what it cost, who approved it, what governs it, and where it stands today.

Every department sees that object from its own angle and calls it something slightly different. Finance sees a cost centre, operations sees a job, and legal sees an engagement. The object is the same. The disagreement is not about data quality. It is about the absence of an agreed structure for the thing the company earns from.

That structure is what the model is. It is also why a better language model does not close the gap: the missing information was never in the text.

What to ask before funding the next pilot

  1. What is the central operating object, and does every system agree on its identity? If four systems hold four identifiers for the same job, that is the first problem, and it is not a model problem.
  2. Who supplies the context in this pilot that nobody will supply at scale? Name the person. If you cannot, the pilot has not started yet.
  3. What is the definition of success, written down before the pilot starts? A remark from a satisfied executive is not a measurement.
  4. Over what window is it measured, and who agreed to that window in advance? Six months is a choice, not a fact, and it is short for anything that touches an operating routine.
  5. When the system produces an answer, what does it change in the systems you already run? If the answer is nothing, you have bought a report.
  6. If it acts, what is the record of what it did? An action with no receipt is not something you can run a business on.

None of these questions require a vendor to answer. They are answerable by the people who already run the work, and they are the questions that separate a pilot that can scale from one that cannot.

Q&A

Is it true that 95% of AI pilots fail?

No. The figure comes from a preliminary working paper by Project NANDA at the MIT Media Lab, cover dated July 2025, whose actual claim is that 95 percent of organizations are getting zero return. On the paper's own funnel of 60 percent evaluated, 20 percent piloted and 5 percent implemented, about 25 percent of organizations that ran a pilot reached production.

Did MIT publish the 95% AI failure study?

It was published by four authors associated with Project NANDA at the MIT Media Lab, and it is labelled "Preliminary Findings" and cover dated July 2025. It is a working paper rather than peer reviewed research, and describing it as an MIT study overstates its status.

How far has AI actually got inside enterprises?

Further for individuals than for the business. In McKinsey's "The state of AI in 2026: On the road to ROI", published 25 August 2026 with 1,719 respondents across 97 nations, nearly nine in ten report regular AI use in at least one business function and 44 percent say AI is scaling across their enterprise, but 80 percent report a gain in their own productivity against 37 percent who can attribute any EBIT impact to it. That last figure is essentially unchanged from a year earlier.

Why does an AI pilot work but the rollout fail?

Because the pilot borrows organizational context from the people in the room. Somebody knows which system holds the true record, which project a document belongs to, and what state the job is in. That knowledge is supplied live and never written down, so it does not scale past the pilot group.

Is the problem data quality?

Usually not in the way it is described. The data exists across the ERP, the document system and the spreadsheets. What is missing is an agreed structure: one identity for the central operating object, and agreement on what has to be true about it - what it cost, who approved it, what governs it, and where it stands.

Sources

  1. The state of AI in 2026: On the road to ROI

    McKinsey and Company. 25 August 2026, surveyed 4 May to 8 June 2026, n=1,719 across 97 nations.

  2. The GenAI Divide: State of AI in Business 2025 (preliminary findings)

    Project NANDA, MIT Media Lab. July 2025.

  3. MIT report: 95% of generative AI pilots at companies are failing

    Fortune. 18 August 2025.

  4. If 95% of generative AI pilots fail, what is going wrong?

    LeadDev. 24 November 2025.

  5. The story behind the bad AI stat that moved markets and misled millions

    80,000 Hours. 28 April 2026.

Where does this break in your organization?

Tell us about one process you actually run. We answer with what we would look at first, not with a deck.

You may unsubscribe at any time.

The information you provide is voluntary and will be used by HIVO IO Technologies Ltd. (517268264) to respond to your inquiry and follow up regarding relevant HIVO services.

How your information is handled

If you do not provide the required information, we may be unable to respond. Your information may be processed by service providers that support our website and business operations, including outside Israel, for those purposes. You may request access to or correction of your personal information by contacting office@hi-vo.io.

See our Privacy Policy for more information.

office@hi-vo.io+972 73-348-8855

Ness Ziona, Israel

Related reading

Operational systems

What an AI agent needs before it can act

Credentials decide what an agent may do. They do not decide whether the action is the right one. Agent safety is a semantics problem before it is a permissions problem.

After the demo

What happens after the demo

A demo works because a person in the room supplies what the model is missing. Rollout is the moment that supply runs out. What has to be built in between is not a better model.