Free executive brief · Agent Playbooks

The "95% of AI pilots fail" number in your board deck is wrong.

It comes from a preliminary, non-peer-reviewed paper built on 52 interviews. The methodology quoted in the press contradicts the methodology printed in the PDF. And the figure everyone repeats is not the figure the paper reports.

Published July 2026 by Anodize Labs. Every source below was checked against a live primary document on 26 July 2026, and the date is printed next to each one. Nothing on this page is gated, there is no popup, and you do not need to give us an email address to read any of it.


The three corrections, in one place

  1. The 95% figure. It traces to a July 2025 preliminary paper from Project NANDA at the MIT Media Lab, based on 52 structured interviews, 153 survey responses and a review of 300-plus public initiatives. The widely circulated description of its method (150 interviews, a 350-person employee survey, 300 deployments) does not match the document. The 95% is also a production rate for one narrow category of tool, not a failure rate for AI.
  2. Self-reported time savings are not evidence. In a randomised trial, experienced developers were 19% slower when allowed to use AI while estimating afterwards that they had been 20% faster. That single result retires the hours-saved survey, which is the instrument most boards are currently shown.
  3. Your governance timeline moved. The EU AI Act's high-risk obligations were pushed to 2 December 2027 and 2 August 2028 by the adopted Digital Omnibus, and Colorado's AI Act was repealed and replaced before it ever took effect. Most published guidance still carries the old dates.

1. The claim you have probably repeated

Some version of this sentence has been in front of nearly every executive in the last year: ninety-five percent of AI pilots fail. It appears in board decks, vendor pitches, consultancy proposals and internal memos arguing both for and against spending more money. It is the closest thing enterprise AI has to a shared fact.

It is worth knowing exactly what produced it, because the answer changes what you should do with it.

2. Where the number actually came from

The source is The GenAI Divide: State of AI in Business 2025, by Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, published under Project NANDA at the MIT Media Lab in July 2025 and picked up widely through Fortune in mid-August 2025. Its own notes page states the method:

"This report is based on a multi-method research design that includes a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders collected across four major industry conferences."

Now compare that with the description that circulated. Secondhand write-ups characterised the study as 150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments. The paper says 52 and 153. The methodology most people can recite for this statistic is not the methodology of the paper the statistic comes from.

What the 95% actually measures

The report describes a funnel for custom or vendor-sold, task-specific generative AI tools: 60% of organisations evaluated one, 20% piloted, and 5% reached production. That is where 95% comes from. It is a production rate for a narrow category of build, measured six months out.

The same report measures general-purpose tools such as ChatGPT and Copilot separately, and the numbers are almost inverted: 80% investigated, 50% piloted, 40% implemented. That is a pilot-to-implementation rate of roughly 83%.

The denominator swap

"95% fail" and "83% of general-purpose pilots reach implementation" are both from the same paper, measured on different populations. Whenever someone quotes a failure rate at you, the question that resolves almost everything is: failure of what, counted against what?

Five problems, four of them conceded by the authors

  1. Not peer reviewed, and self-labelled preliminary. The cover page reads "Preliminary Findings from AI Implementation Research from Project NANDA." The listed reviewer is also a listed author. There is no journal and no replication package.
  2. Small and non-random. 52 interviews and 153 conference attendees. The report concedes "Selection bias possible in organizations willing to participate in AI research."
  3. A subjective and short success bar. Success is partly coded from whether "users or executives have remarked" on impact, measured six months after the pilot. The report itself warns that the "Six-month observation period may be insufficient... potentially understating success rates."
  4. A structural conflict of interest. NANDA builds agentic and memory interoperability infrastructure, and the report's prescribed remedy is agentic systems plus those protocols. The diagnosis deserves more weight than the prescription.
  5. The circulating methodology contradicts the document. This one is not conceded, because the authors did not write it. It was added in transmission.

The report also states its own limitation plainly: "These figures are directionally accurate based on individual interviews rather than official company reporting. Sample sizes vary by category, and success definitions may differ across organizations."

Lead author Challapally has since publicly softened the framing, saying that "Some large companies' pilots and younger startups are really excelling with generative AI."

3. What the evidence actually supports

Here is the defensible version of the sentence, and it is still a useful finding:

In a small, non-random, non-peer-reviewed 2025 study, roughly 5% of custom enterprise generative AI tools reached production with measurable impact within six months.

Cross-check that against the larger surveys and a range appears rather than a number.

SourceWhat it measuresFigure
MIT NANDA, July 2025 (n=52 organisations) Custom or vendor-built task-specific tools reaching production 5%
MIT NANDA, July 2025 General-purpose LLM pilots reaching implementation About 83%
IBM Institute for Business Value, 2025 CEO Study (n=2,000 CEOs) AI initiatives delivering expected ROI over three years 25%
McKinsey, State of AI, November 2025 (n=1,993) Organisations attributing any level of EBIT impact to AI 39%, and most of those say under 5% of EBIT
McKinsey, State of AI, November 2025 "AI high performers": 5%+ of EBIT attributable plus significant reported value About 6%
BCG, The Widening AI Value Gap, September 2025 Firms reaping "hardly any material value" 60%, with 35% scaling and 5% "future-built"
Menlo Ventures, December 2025 (n=495 decision-makers) AI deals converting to production 47%, versus 25% for traditional SaaS

The honest range for "pilots that produce measurable value" is roughly 5% to 40%, and where you land inside it is decided entirely by what you count. None of these are independent of each other in the way they are usually presented, either. Stanford HAI's widely quoted adoption percentages are imported from the McKinsey survey, so citing both for adoption is citing one survey twice.

Who is selling the remedy

McKinsey, BCG and Deloitte publish the numbers most often quoted about AI programme failure, and all three sell AI transformation work. IBM sells watsonx and the consulting around it. Menlo Ventures is an investor in the category it surveys. None of that makes the numbers false. It does mean the sampling frame and the definition of success deserve a look before the figure goes on a slide.

4. The finding that retires your current ROI case

Almost every AI business case in circulation rests on time saved, and almost every estimate of time saved comes from asking people how much time they saved. There is now a randomised controlled trial that measures both at once.

METR, a nonprofit evaluation organisation with no vendor funding, ran 16 experienced open-source developers across 246 randomised tasks in mature repositories they had worked in for years. From the paper:

"developers forecast that allowing AI will reduce completion time by 24%... developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%."

Expected 24% faster. Believed 20% faster afterwards. Measured 19% slower. The gap between perception and measurement is the finding, and it is roughly 39 percentage points wide.

State the caveats, because they matter

n=16 is small and the confidence interval is wide. The study deliberately chose the least favourable case for AI: experts working in codebases they know deeply. It is not evidence that AI slows everyone down, and anyone presenting it that way is making the same mistake in the opposite direction. What it does establish is narrower and more damaging to your reporting: practitioners cannot introspect their own productivity change, and they get the sign wrong.

The rest of the causal evidence

Surveys measure opinions. These measure output. Taken together they point somewhere quite specific.

StudySample and designResult
Dell'Acqua et al., Navigating the Jagged Technological Frontier, HBS WP 24-013 (2023), Organization Science 758 BCG consultants; randomised, no AI versus GPT-4 versus GPT-4 with prompt training Inside the capability frontier: about 12.2% more tasks, 25% faster, 40% higher quality. Outside it: about 19 percentage points lower correctness with AI. Bottom-half performers gained about 43% on quality versus about 17% for the top half
Brynjolfsson, Li & Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2), 2025 About 5,172 support agents at a Fortune 500 software firm; staggered rollout 14 to 15% more issues resolved per hour, about 34% for novice and low-skilled agents, minimal effect on the most experienced
Cui, Demirer, Jaffe, Musolff, Peng & Salz, Management Science, 2025 4,867 developers across Microsoft, Accenture and an anonymous Fortune 100; three pre-registered RCTs "a 26.08% increase (SE: 10.3%) in completed tasks... less experienced developers had higher adoption rates and greater productivity gains"
Noy & Zhang, Science 381:187-192, 2023 453 college-educated professionals; randomised, mid-level writing tasks Time on task down about 40%, quality up about 18%, gains concentrated among lower-ability writers
METR, arXiv:2507.09089, July 2025 16 experienced developers, 246 randomised tasks in mature repositories 19% slower with AI, while believing they were 20% faster
Peng, Kalliamvakou, Cihon & Demirer, arXiv:2302.06590, 2023 About 95 developers, one synthetic greenfield task 55.8% faster. Three of four authors were at Microsoft or GitHub, the task was synthetic, and no review or maintenance burden was measured

The pattern is consistent and it is the opposite of how most rollouts are targeted. Every rigorous study finds gains concentrated in less-experienced staff on well-specified tasks, and neutral to negative effects for experts doing complex work in systems they already know. Most enterprises pilot with their senior people, because that is where the political capital sits. The evidence says pilot where the skill floor is.

5. The governance dates in your policy document are probably stale

Two changes landed recently enough that most published guidance, including several popular compliance trackers, has not caught up.

The EU AI Act's high-risk deadlines moved

Regulation (EU) 2024/1689 entered into force on 1 August 2024 with a staged timeline. The Digital Omnibus on AI, procedure 2025/0359(COD), was proposed on 19 November 2025 and moved the two high-risk dates. It is adopted, not merely proposed: Parliament adopted it in plenary on 16 June 2026 (423 for, 57 against, 174 abstentions), the Council adopted it on 29 June 2026, and the final act was signed on 8 July 2026.

DateWhat appliesStatus
2 February 2025 Article 5 prohibited practices and Article 4 AI literacy In force now
2 August 2025 General-purpose AI obligations, governance structures, most penalties In force now
2 August 2026 Article 50 transparency, Article 27 fundamental rights impact assessment, Article 86 right to explanation Unchanged
2 December 2027 Stand-alone Annex III high-risk obligations, including the employment and recruitment category Moved by the omnibus. Previously 2 August 2026
2 August 2028 Annex I product-embedded high-risk systems Moved by the omnibus. Previously 2 August 2027

One honest caveat on this one. The Commission's regulatory framework page describes the omnibus as having entered into force in July 2026, while the Legislative Observatory still showed the file as completed and awaiting publication in the Official Journal when we checked on 26 July 2026. The substance is settled; the final Official Journal citation and the precise entry-into-force date are not yet pinned down, so do not put a citation string in a policy document until they are. We have also seen no confirmation of whether the Article 4 AI literacy duty was altered, so treat it as binding.

Two things are worth separating from the delay, because they are in force today and they catch ordinary employers with no AI product of their own:

  • Article 5(1)(f) prohibits AI used to infer emotions of a natural person in the workplace, except for medical or safety reasons. Emotion or sentiment analytics pointed at your own staff is not a 2027 problem. It has been prohibited since February 2025.
  • Article 4 requires providers and deployers to ensure "a sufficient level of AI literacy" among staff and others operating AI systems on their behalf, in force since February 2025.
  • Article 2 reaches a company with no EU establishment where "the output produced by the AI system is used in the Union." A US-only headcount does not settle the question.
  • Article 99(6) applies the lower of the percentage or the fixed amount for small and medium enterprises and startups. This clause is directly relevant to a company of 50 to 5,000 people and is almost never mentioned in the popular coverage.

Colorado's AI Act was repealed before it took effect

Colorado SB 24-205, signed 17 May 2024, was the first comprehensive US state AI law and became the template a great deal of governance advice was written against. It never applied to anyone. Its effective date was pushed from 1 February 2026 to 30 June 2026 by SB 25B-004 in August 2025, and it was then repealed and reenacted as SB 26-189, the Automated Decision-Making Technology Act, signed 14 May 2026.

Survived into the new act

  • Developer documentation duties
  • Pre-use notice to consumers
  • A plain-language explanation within 30 days of an adverse outcome
  • Rights to access and correct data, and to request meaningful human review
  • Three-year recordkeeping

Deleted

  • The duty of reasonable care against algorithmic discrimination
  • Annual impact assessments
  • The NIST and ISO aligned risk management programme requirement

The new act's statutory effective date is 12 August 2026, with substantive developer and deployer obligations running to 1 January 2027. Enforcement is by the Attorney General only, with no private right of action, penalties up to $20,000 per violation and a 60-day cure period until 1 January 2030.

If your AI governance programme was built to satisfy the original Colorado act, roughly half of what it obliges you to do is now voluntary in Colorado. That may still be the right thing to do. It is no longer the thing the statute requires, and knowing the difference is the point.

There is no federal preemption

The proposed 10-year moratorium on state AI laws was stripped from the One Big Beautiful Bill Act by a Senate vote of 99 to 1 on 1 July 2025, and was omitted again from the FY2026 NDAA conference text released 8 December 2025. Executive Order 14365 of 11 December 2025 instead directs the Attorney General to establish a task force to challenge state AI laws. For a board, the practical answer is that state AI law is fully operative, and preemption is a litigation and funding risk rather than a legal fact.

6. What to do differently on Monday

Five changes, each of which follows directly from something above and none of which requires a budget.

  1. Ask for the denominator, out loud, every time. When a failure rate or a success rate is presented, the question is "counted against what population, over what window, with success defined how?" On the 95% figure alone, that question moves the answer between 5% and 83% using the same source document.
  2. Stop accepting self-reported hours saved as evidence of anything. Treat it as a measure of enthusiasm, which is worth knowing and is not worth booking. Where a business case rests on time saved, ask what independent measurement would falsify it, and whether that measurement is being taken.
  3. Establish a baseline before the pilot, not after it. Capture how the work happens today: time, cost and error rate, across several cycles rather than the first attempt. A pilot without a pre-agreed baseline cannot produce a result, only an impression.
  4. Pilot where the skill floor is. Every rigorous study puts the gains in less-experienced staff on well-specified work. Starting with your most senior people is the politically obvious move and the empirically worst one.
  5. Date-stamp your governance document and re-check the two live obligations. Put a verification date on every regulatory claim in it. Then check two things that apply today regardless of the delayed high-risk deadlines: whether anything in your stack infers employee emotion or sentiment, and whether you have anything that would satisfy an AI literacy duty.

One further note on the projections you will be shown. The most quoted enterprise AI success story of 2024, Klarna's customer service automation, was published as "projected to deliver $40 million in profit improvement." Projected, not realised, and the company later walked back parts of the automation. Never book a projection in a business case without the measurement plan that would falsify it.

Sources

Everything above traces to a document you can open. Each was checked against a live primary source on 26 July 2026. Where a citation could not be resolved to a primary document, that is said at the point of use rather than hidden here.

The 95% claim

  • The GenAI Divide: State of AI in Business 2025, Challapally, Pease, Raskar and Chari, Project NANDA, MIT Media Lab, July 2025. Working PDF mirror. Verified 26 July 2026.
  • Paul Roetzer, Marketing AI Institute, on what the 95% definition excludes. marketingaiinstitute.com. Verified 26 July 2026.

Base rates and market surveys

  • McKinsey, The state of AI in 2025: Agents, innovation, and transformation, November 2025, n=1,993 in 105 nations. mckinsey.com. Verified 26 July 2026.
  • IBM Institute for Business Value, 2025 CEO Study, 6 May 2025, n=2,000 CEOs. ibm.com. Verified 26 July 2026.
  • BCG, The Widening AI Value Gap, 30 September 2025. BCG does not publish the respondent count or sampling frame on the public page. bcg.com. Verified 26 July 2026.
  • Deloitte, State of AI in the Enterprise, 2026 edition, n=3,235 senior leaders in 24 countries. deloitte.com. Verified 26 July 2026.
  • Menlo Ventures, 2025: The State of Generative AI in the Enterprise, 9 December 2025, n=495. menlovc.com. Verified 26 July 2026.
  • Stanford HAI, 2026 AI Index Report, April 2026. Its adoption percentages are imported from the McKinsey survey and are not independent of it. hai.stanford.edu. Verified 26 July 2026.

The causal evidence

  • METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025. arXiv:2507.09089 and metr.org. Verified 26 July 2026.
  • Dell'Acqua et al., Navigating the Jagged Technological Frontier, HBS Working Paper 24-013 (2023), published in Organization Science. doi 10.1287/orsc.2025.21838. Verified 26 July 2026.
  • Brynjolfsson, Li & Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2):889, 2025. academic.oup.com. Verified 26 July 2026.
  • Cui, Demirer, Jaffe, Musolff, Peng & Salz, The Effects of Generative AI on High-Skilled Work, Management Science, June 2025, pre-registered AEARCTR-0014530. Author copy. Verified 26 July 2026.
  • Noy & Zhang, Science 381:187-192, July 2023. doi 10.1126/science.adh2586. Verified 26 July 2026.
  • Peng, Kalliamvakou, Cihon & Demirer, February 2023. arXiv:2302.06590. Verified 26 July 2026.

Regulation

  • Regulation (EU) 2024/1689 of 13 June 2024, published in the Official Journal on 12 July 2024, in force from 1 August 2024. Article and Annex references above are to the regulation text. Verified 26 July 2026.
  • Digital Omnibus on AI, Commission proposal COM(2025) 836 of 19 November 2025, procedure 2025/0359(COD). Parliament plenary adoption 16 June 2026 (T10-0198/2026), Council adoption 29 June 2026, final act signed 8 July 2026. Tracked through the European Parliament Legislative Observatory and Legislative Train. Final Official Journal citation still unresolved as of 26 July 2026.
  • Colorado SB 24-205 (signed 17 May 2024), SB 25B-004 (28 August 2025), and SB 26-189, the Automated Decision-Making Technology Act (signed 14 May 2026). Verified 26 July 2026.
  • Executive Order 14365, Ensuring a National Policy Framework for Artificial Intelligence, 11 December 2025, 90 FR 58499. Verified 26 July 2026.

The Klarna example

  • OpenAI, AI in the Enterprise: Lessons from seven frontier companies. cdn.openai.com. The $40 million figure appears there as a projection. Verified 26 July 2026.

If this was useful

This brief is three corrections. It is not a sales page with a chapter attached, and if you close the tab here you have the part that changes what you do on Monday.

We also publish The Agent Playbook for Executive Leaders, which is built the same way: named sources with sample sizes, dates and sponsors, and a plain statement wherever a widely repeated number is weak. It covers where value actually shows up by function, build versus buy versus adopt, org design, budget and unit economics with dated vendor prices, measurement, and rollout. It costs $199 once, it is a PDF, and there is a 60-day refund with no questions asked. It is genuinely not for you if you want a certification, a cohort, or a hands-on prompting course.

The executive guide

Everything above, applied to a decision you actually have to make.

This brief corrects three figures. The guide is 114 pages of the same work: where value shows up by function, build versus buy versus adopt, org design, budget and unit economics with dated vendor prices, how to measure it without relying on self-reported hours, and the questions to ask your teams and your vendors.

$199 once. A PDF, delivered immediately. A 60-day refund, no questions asked. Team and department licences are on the product page.

See the full contents first

114-page PDF · one-time payment, not a subscription · instant download plus an emailed link · 60-day refund, no questions asked. We use your email only to deliver the file and its updates. See the privacy policy.

Updates to this brief

Two of the claims above have expiry dates.

The EU Official Journal citation is unresolved, and the Colorado obligations phase in through January 2027. If you want the corrections when they land, leave an email address. We send updates to this brief and nothing else: no drip sequence, no course launch, and you can leave at any time. See the privacy policy.

Done. Check your inbox for a confirmation, which also carries the unsubscribe link.

That did not work. Check the address and try again, or email us and we will add you by hand.

You are unsubscribed. You will not hear from us about this again.

See the executive guide

Corrections to this brief only. If the form does not work, email support@anodizelabs.com with the word "brief" and we will add you by hand.

Agent Playbooks is published by Anodize Labs, an independent shop in Spokane, Washington. We are not affiliated with, endorsed by or sponsored by MIT, METR, Anthropic, OpenAI, or any organisation named on this page, and no one paid for a mention. Contact details and business information.