Writing / 2026

The Board's AI Oversight Problem Is Operational

AI oversight is metrics, owners, halt authority, and an incident rule—not a literacy seminar.

The number that matters is rarely accuracy; it is how long a misbehaving system runs before someone with authority stops it. That interval is a P&L line, and slow oversight is its own decision latency that drags on margin and risk . A board that knows its error rate to two decimals but cannot name who halts the system, how fast, and at what cost has measured everything except the thing that bounds the loss.

So make those three answerable in calm, before the next bad week instead of during it. Who is a named owner per production system with an explicit halt path, not a committee that assembles after the damage. How fast is the elapsed time from the first bad signal to the system actually pulled, and that only shrinks if the halt is rehearsed rather than improvised. At what cost is the incident threshold agreed in advance: which incidents reach the room, and what the firm is willing to forgo to pull the cord. Wire that into the cadence you already run, a short standing report and an owner who shows up, and you have governance without bureaucracy . It is the lever a board can actually pull.

The full apparatus has four parts: a named owner per production system, an explicit halt path, a rule for which incidents reach the room, and a failure rate measured the same way every quarter. The first three are easy to recite. The failure rate is genuinely hard, and it is exactly where boards get fooled.

The fooling arrives on a slide: the AI system’s error rate fell from 2.1% to 1.7%. Heads nod, the program looks healthy, the meeting moves on. Both figures are closer to fiction than measurement, and a board that cannot say why has not overseen anything. It has been shown a chart.

Why “the same way every quarter” is the hard part

A non-deterministic system has no error rate to read off a gauge. The same input yields different outputs. The model version changes under you, the traffic mix shifts, and the humans grading “wrong” quietly revise what wrong means. So 1.7% is the product of a dozen choices: which outputs were sampled, how many, against what rubric, judged by whom. A defensible number pins all of them down.

  • A frozen definition of failure, written as a rubric a grader applies, with inter-rater agreement tracked so the bar does not drift loose over time.
  • A fixed sampling rule: N items drawn fresh from live traffic each quarter, never a static test set. The moment engineers know which prompts are scored, they teach to the test and the rate detaches from reality.
  • A confidence interval next to the point estimate. On a few hundred samples, 2.1% to 1.7% sits inside the noise band, and presenting it as progress is the lie the slide tells.
  • A re-baseline rule: when the rubric or the sample changes, run old and new on one overlap quarter and report both, so the trend is not an artifact of moving the goalposts.

One trap catches the diligent. A frozen golden set goes stale; as the world drifts it stops resembling live traffic and starts flattering you. The defense is to watch the gap between your scored sample and a fresh live sample, and treat divergence as a signal the metric is lying before the system is. A board does not run this machinery; that is the IC’s standard to meet. The board asks the questions that prove someone did, and refuses a number that cannot answer them.

Which is why the usual oversight advice underwhelms. Get an AI expert on the board, run a literacy seminar, issue a principles memo: none of that tells a director whether 1.7% is real, and none of it shortens the halt. Education is a fine supplement and a dangerous substitute. The failure it breeds is specific, a roomful of directors who can define the risk fluently and still cannot intervene when it walks in.