Writing / 2026

Power Belongs on the AI Roadmap

Power, not models or GPUs, is the binding constraint. Treat energy and capacity as roadmap dependencies with real lead times.

A large new grid connection is a multi-year request. The queue to energize big new load is measured in years, not quarters, and the transformers a substation needs are ordered on lead times of a year or more. You will never file one of those orders; the cloud you buy inference from does, and the megawatts stay someone else’s problem right up until that schedule becomes yours.

The lead time does not disappear when you outsource it; it reaches you renamed. A quota request that takes weeks and comes back denied. A reserved-capacity contract your provider will only sign for the region where the power already exists. A new accelerator generation that is available in three regions and waitlisted in the rest. Scarce capacity lands first where the grid is already built and trickles outward from there, so the map of where you can actually run is the map of where the power already sits. You did not secure that power; you inherited its schedule, and that schedule is unforgiving in a way software schedules are not. Assume the GPUs will be there and you have assumed away a constraint nobody upstream can wave away either.

Here is the shape of the miss. A team commits a 5x ramp to the board: volumes, latency targets, a confident curve. The model works in the demo, the math closes in the spreadsheet. Then the data-residency commitment pins the workload to one region, the provider’s on-demand quota in that region is capped, the reserved-capacity deal that would lift the cap is a negotiation measured in months and multi-year terms , and the ramp you already sold slips two quarters. That is not a forecasting miss. It is a dependency you mislabeled as an assumption.

So price it like a dependency. On the roadmap, every committed ramp carries, next to its volume and latency targets:

  • the region it actually lands in, and the residency rule that pins it there;
  • whether that capacity is on-demand (no guarantee) or reserved (contracted), and the lead time to convert one to the other;
  • the named owner of the provider relationship who can escalate a quota or open a region;
  • the fallback region, with the latency and residency cost of failing over to it.

The reserved-capacity line is the nearest thing you hold to a power-procurement order: contracted capacity you claim well ahead of need, not when the quota is already capped. That is how a roadmap survives contact with reality : it names the capacity constraint before the board discovers it.

Now stop spending the scarce capacity on work that does not need it. The premium path, frontier reasoning on the newest accelerators in the few regions that have them, should be rationed by design. Most production volume is high-frequency, low-stakes work that runs fine on a smaller model, a batched job, a cached result, or last-generation hardware that is abundant precisely because everyone else moved on. Routing that volume down buys back ramp without buying more capacity, and it is the one lever here you pull directly rather than petition a vendor for.

What remains is geography and timing. Workloads can move to where capacity is cheap and available; the commitments you made to customers usually cannot move as fast. So choose regions deliberately, hold a qualified fallback for each committed workload, and treat residency rules as the hard constraint they are rather than a detail to reconcile later. The failure mode is always the same: one assumed region, one assumed quota, one assumed timeline, and nothing staged for when any of the three moves. Stage the move before the cap forces it, because by then the cap is not negotiable and your customers already have the date.