Writing / 2026
The Review Queue Is Your Real Agent Limit
Plan agent rollouts like capacity: risk-weighted review demand against effective reviewer-hours. Past the constraint, seats buy inventory, not throughput.
The trend of the season, visible in the newsletters and in any engineering org you ask, is rising review load. I’ve argued the general mechanism already : speedups upstream of a fixed-capacity stage convert to queue, not throughput. This post is the specific machinery for the case that matters most to engineering leaders, code review under agent-generated change volume, because “the handoff tax exists” doesn’t answer the question a VP actually faces in Q4 planning: how many agent seats can we absorb, and what do we change to absorb more? That’s a capacity-planning question, and it deserves a capacity plan, not a sentiment.
The plan needs two quantities nobody currently measures. The first is offered review demand, and the unit is review-minutes, not pull requests. Counting PRs is like counting “some vehicles” at a bridge; a dependency bump and an authorization rewrite are different loads. Estimate minutes per change from size, surface touched, and blast radius, crudely at first, calibrated later from actuals. Note what this replaces: seat count. Seats are not demand. A seat can sit idle, can replace typing without adding submissions, or can produce changes that arrive with evidence and review faster. What hits the queue is minutes, so measure minutes. The second quantity is effective reviewer capacity: not headcount times forty hours, but the review-minutes your qualified reviewers can actually give, after meetings, on-call, and their own delivery work, measured from review activity rather than assumed.
Demand against capacity gives you a utilization ratio, and here queueing theory earns its keep as a warning model, no more. For a stage with high variability in job sizes, which review is, extremely, waiting times grow nonlinearly as the ratio approaches one, and the growth steepens with the variability. Don’t quote a magic threshold; the practical rule is humbler. Track the ratio and the queue-age tail together, per risk class, and treat a rising tail as the constraint announcing itself. The classes matter because the single most effective lever isn’t more reviewers. It’s refusing to price all changes identically.
That lever is risk-tiered review, and it should come before any cap. Low-risk classes, the reversible, well-tested, non-privileged surfaces, get automated checks plus sampled human review at a stated rate, with the sample watched for escapes; the tier earns its lighter process by evidence, and loses it the same way. High-risk classes get full human review from a capacity budget reserved for them, so an authorization change never queues behind forty dependency bumps. Tiering can multiply effective capacity outright. What it can’t do is repeal the arithmetic on the residual: the human-review tiers still have an arrival rate and a capacity, and when the first approaches the second, you cap work-in-progress. Not because Little’s Law commands it, but because the relationship it states means bounded queues are the only way to bound cycle time. When the cap binds, the rule for agents is specific, because “agents work on the queue” invites a category error: an author rechecking itself is not verification. Capped agents split oversized changes, rebase stale ones, answer review comments, repair failing checks, and attach evidence (tests, eval results, stated blast radius) that shrinks a reviewer’s minutes-per-change. Independent verification stays with an independent party, human or a separately-tasked checker , and new submissions wait.
Two boundaries on this program. The metrics need definitions before dashboards: attribute changes to agents at commit metadata, split queued-minutes from active-review-minutes, track p50/p95 queue age per tier, and count escapes per hundred merged changes in a fixed attribution window. Loose definitions here produce confident nonsense. And the apprenticeship pipeline is a capacity investment on a months-long crossover, not a Q4 lever; budget it like one, expecting senior minutes to rise before they fall.
The one-sentence version for the seat-count conversation: your agent capacity is not your seat count, it’s your verified-throughput rate, and the ratio of offered review-minutes to reviewer-minutes tells you which side of the constraint each team is on. Buy seats up to the constraint. Past it, spend the same money on the constraint itself, on tiering, evidence, and reviewers, because that’s where throughput now lives.