Writing / 2026

The Diffusion Gap: Why Pilots Don't Become Capability

Most pilots succeed and die. The gap from demo to daily operation is ownership and cadence—write exit criteria as operating changes, not model scores.

The money in an AI pilot rarely disappears in a crash. It goes quietly, in licenses paid and habits unchanged.

A claims team pilots an assistant that drafts the first-pass adjustment letter. The trial is clean: on a held-out sample it produces a usable draft for 78% of cases, and the two adjusters running it cut letter-writing time roughly in half. The deck circulates. Everyone agrees it works. Eight months later, fourteen adjusters hold a license and nine have opened the tool twice. The default is still the old template button in the case system, because that is the button that has always been there.

The model did not fail. What failed is the work nobody scoped: moving a proven capability from “passes a trial” to “is how the job is done.” That is the distance between a model that performs and an organization that operates differently because of it, and it is where the money went.

Most companies running pilots have a graveyard of these. The audit is cheap: list every pilot that cleared its technical bar, then mark which daily workflow each one actually changed. The blanks are your real result, and they cluster for one reason. A pilot gets scored on whether the model works, which answers a question nobody important was asking. The question that pays is whether work will route through it next quarter.

Three things have to move for that to happen, and they fail as a set: ownership, cadence, and incentives.

The first two are org-chart edits a director can make before lunch. Ownership leaves the pilot team for a named operator whose performance review includes the adoption number. Cadence puts that number into a standing review, so it is read on a fixed clock rather than admired once. Do only these and you get a well-governed tool nobody uses. That is the claims pilot’s eight months exactly: a sponsor, a slide in the monthly deck, and usage in the low twenties, because the cheapest action on screen was still the old one.

Incentives is the lever everyone hand-waves as “drive adoption,” and it is the only hard one, because “path of least resistance” is not a mood. It is three concrete edits to the screen and the scorecard.

The first is to delete the old path for the cases the tool handles. The highest-leverage move is removing the legacy button: take the manual “insert template” action out of the case screen, so drafting through the assistant is the one-click route and the old way costs four extra clicks. That is not the same as deleting the template for the 22% the assistant can’t draft; those cases still route to it. Adoption is a UI default before it is a culture.

The second is to move the measured metric. Stop scoring adjusters on raw letters closed per day and start scoring first-pass acceptance, a number the assistant moves and the manual path does not. Now the comp-relevant metric rewards the new behavior instead of being neutral to it.

The third is to make the old path visible when chosen. The two holdouts still writing by hand show worse handle time on that same scorecard, which is a line in the weekly review, not a policy memo.

Incentives gets skipped because it is the expensive lever. It means touching a scorecard people are measured on, deleting a tool they are comfortable with, and eating the complaints when the familiar button is gone.

So write the pilot’s exit criteria as operating changes before it starts, not model targets after it ends:

  • Which human task stops, and who stops doing it?
  • Who owns the capability in production, with their name on the adoption metric?
  • Which existing meeting reviews it, on what clock?
  • What is the old path, and what is the plan to remove it for the cases the tool covers?

None of those mention accuracy. A pilot that cannot answer them is funding a build and assuming diffusion is free; it never is. The work that closes the gap is moving a capability from impressive to load-bearing : name the owner, put the adoption number in a standing review, and remove the old step for the cases the new one covers, until the assistant is the route people actually take.