Writing / 2026

The Statistic Nobody Can Reconstruct

I tried to reconstruct '95% of AI pilots fail.' It isn't in the report it cites. Card the numbers that steer your decisions: licensed, expired, or retired.

Sometime this quarter, in a budget meeting near you, someone will say “well, 95% of AI pilots fail,” and a decision will move. I went looking for where that number comes from, expecting to find a stale-but-real measurement. What I found is better, in the way autopsies are better: the statistic cannot be reconstructed at all.

The trail leads to “The GenAI Divide: State of AI in Business 2025,” from MIT’s Project NANDA, published July 2025. It’s a non-peer-reviewed report built on 52 executive interviews and 153 conference-collected survey responses, measuring outcomes about six months post-pilot. The report contains a narrow finding, that only around 5% of custom enterprise AI tools reached production, and a separate executive-summary assertion that roughly 95% of organizations saw no measurable return. “95% of pilots fail” is neither claim; it’s the two welded together by headline writers. It gets worse. Kevin Werbach, a Wharton professor, read the report repeatedly, publicly said he could not locate where the 95% figure is derived, and called on the authors to release the supporting data or retract. Secondary write-ups can’t even agree on the sample sizes. And as of when I checked, the original PDF’s URL redirects to the project’s overview page. The most-quoted AI statistic of the past year points at a document that no longer answers for it.

Nothing about that stopped the number, and here is the discipline worth extracting from the wreck. Any statistic that recurs in your decisions gets a provenance card, six lines, kept where the decision documents live. Watch it work on this one. Source: Fortune headline, August 2025, over the NANDA report; original document currently unavailable at its published URL. Observed: January to June 2025, outcomes read at six months. Population: 52 interviews plus 153 survey responses, recruited partly from organizations willing to discuss implementation challenges. Definition: contested; “fail” conflates tools-not-reaching-production with organizations-reporting-no-return. Sponsor interest: the publishing project builds agentic-AI infrastructure, and the report’s prescription points at its products. Status: retired; underlying claim not reconstructible, and a named academic has requested retraction. Ten minutes of an analyst’s time, and the number that was about to steer your portfolio is revealed as unusable for any decision at all. That is the strongest possible advertisement for the card, and it’s why the card leads with Source. A provenance card without one is process satire.

The objection writes itself: a CFO under time pressure cannot card every slide number, and heuristics exist because they’re rational. Granted entirely, and the trigger is the fix. Card only the statistics that recur or that materially move a fund, kill, or vendor decision; let the rest stay heuristics, labeled as such. The problem was never shorthand. It’s shorthand dressed as measurement, steering capital .

Two rules keep the cards usable. First, every card has an owner and a review date, and past the date the statistic can inform history but not premise a current decision. Three states, so the meeting can use them as vocabulary: licensed for this decision, expired, or retired. Second, symmetry, meaning the same acceptance test for numbers you like. That October, a Wharton-affiliated survey reported roughly three in four enterprise leaders seeing positive AI returns. Same year as the 95%, opposite headline, both self-reported, and the entire gap is definitional: “positive return” versus “measurable P&L impact.” Cite either without the other and a hostile reader will catch you. Card both and the contradiction dissolves into two defensible, narrow claims. The adoption numbers deserve the same treatment: a16z’s April 2026 finding that 29% of the Fortune 500 are live paying customers of a leading AI startup is a real observation, by sellers’ investors, about purchases, not about success. A leadership team that cards only the scary numbers hasn’t built an evidence practice; it’s built a permission structure.

Your organization is currently building freshness checks so agents won’t act on stale context . Extend the courtesy to the executives. The 95% will be in a boardroom near you for years; numbers this useful don’t die of natural causes. When it arrives, don’t argue with it. Ask for its card.