<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AI Strategy | Law Zava</title><link>https://lawzava.com/topics/strategy/</link><description>AI strategy for CTOs and technical leaders: capital allocation, roadmap discipline, throughput, and operating constraints.</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 13 Aug 2026 07:51:57 +0000</lastBuildDate><atom:link href="https://lawzava.com/topics/strategy/index.xml" rel="self" type="application/rss+xml"/><item><title>Unsupervised</title><link>https://lawzava.com/blog/2026-08-12-unsupervised/</link><pubDate>Wed, 12 Aug 2026 09:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-08-12-unsupervised/</guid><description>Producing the work can be handed to a machine. Answering for it cannot. What an organization values when machines do most of the making, what people are for, and four things to try this week.</description><content:encoded><![CDATA[<p>Producing plausible work has become close to free. Knowing whether it is any good
has not. Undoing it has not. Answering for it has not. Deciding what deserves to
exist has not, and never was. The scarcity moved, and practices built for the old
one keep running with nothing behind them.</p>
<p>What the machines give back is not room to ship more. It is the chance to attempt
work that used to need a company, to serve people the old cost floor priced out,
and to spend a working life on the part of the work that is worth a life. A
company that spends the surplus on volume has been handed a decade and bought a
quarter with it.</p>
<p>Everything here turns on one thing. The work is how people get made. People are
not how the work gets made.</p>
<h2 id="what-i-value">What I value</h2>
<p><strong>Verifying</strong> over producing.</p>
<p><strong>Earned autonomy</strong> over granted autonomy.</p>
<p><strong>Enforced limits</strong> over declared intent.</p>
<p><strong>Doing the work</strong> over approving it.</p>
<p><strong>Judgment that compounds</strong> over output that repeats.</p>
<p><strong>Apprenticeship</strong> over throughput.</p>
<p><strong>Problems surfaced</strong> over progress reported.</p>
<p><strong>Independent sources that can disagree</strong> over one that cannot be wrong.</p>
<p>I want the right column too. What moved is where the scarcity sits.</p>
<h2 id="how-i-work">How I work</h2>
<p>A process nobody can check is not automated. It is unsupervised. The check comes
before the thing it protects: the test before the pipeline, the alarm before the
deploy, the dry run before the announcement.</p>
<p>Verification rises with the consequence of being wrong. Produce enough to learn
from someone real, verify enough to trust an act that repeats or cannot be taken
back, and where the work decides nothing, do neither.</p>
<p>Autonomy is earned with evidence and taken back the same way. For machines it
attaches to a step. For people it widens across a domain as their judgment earns
it, and narrows at the smallest place the evidence indicts. Nobody should live
permanently on probation. And evidence expires: what earned trust last year
describes last year&rsquo;s system, and nothing announces the change.</p>
<p>Nothing is done when it works. It is done when something checks it, an owner
carries the outcome, and the way back has been walked at least once. A rollback
nobody has rehearsed is a rumour.</p>
<p>Some work has no decisive check at a price anyone will pay. Mark it, keep people
on it, and never let a number stand in for what cannot be measured.</p>
<p>Nobody is the only source of evidence that they are doing the thing, and that
includes the machine that did it. You cannot feel whether the machine helped:
work goes faster in the hand and slower on the clock, and nobody has trained that
out of anyone.</p>
<p>Producing the work can be handed to a machine. Answering for it cannot. The
machine wrote it in a second and somebody carries it for a year.</p>
<p>Deciding to kill something is not killing it. It is alive for exactly as long as
its permissions still work.</p>
<p>Measure the system in public. Measure the person in private, and publish who did
the measuring, on what evidence, and who can overturn it. Pay and promote for
judgment under uncertainty, for spreading what works, and for carrying what does
not.</p>
<h2 id="what-people-are-for">What people are for</h2>
<p>Every time a machine takes the making from someone, they are owed a harder
question and not a smaller job. If the harder question does not come with more
standing, it was a smaller job with better words.</p>
<p>Beginner&rsquo;s work is what made the expert, and it is disappearing. Give that work
to beginners anyway, because a trade that admits nobody has stopped being a
trade. A company that ships more this year and can make fewer experts than last
year has eaten capital and called it margin. No dashboard you already have will
show it, so count the people you took from beginner to independent and publish
that beside what you shipped.</p>
<p>Refuse work whose value depends on weakening the agency of the people it touches.
Such work can be profitable and perfectly controlled. Whether its success is
worth causing is the question, and it is asked before the controls and not
after. Between two things
you could build and defend, build the one that leaves people more able to act
without you.</p>
<p>Anyone a decision lands on can reach a person who can explain it and change it.
Not because the machine is often wrong, but because being processed by something
that cannot be asked is a harm at any error rate.</p>
<p>If the work is handed to you rather than by you, all of that is what somebody
owes you: work that can go wrong, a harder question when a machine takes the
making, a name and a route when a decision goes against you. Ask for them by
name. Asking is not insubordination, and if it is treated as insubordination
where you are, you have your answer on your first day instead of in your third
year.</p>
<h2 id="where-to-start">Where to start</h2>
<p>None of this needs agreeing with first. Four things are doable this week, and
each hands back its own evidence.</p>
<p>Ship one change only after a check you did not write.</p>
<p>Before you let anything run on its own, walk the way back once, by hand. You will
find a rollback or you will find a rumour.</p>
<p>Find something you declared dead last year and try its permissions. Whatever
still answers was never killed.</p>
<p>Put a name and a route on one automated decision that lands on people, and ask
the ones it went against whether they were told anything they could use.</p>
<p>Then stop doing them as an audit, because an audit is a thing you pass. Each one
is fastened to a moment in the work and comes back when the moment does: a check
you did not write, every time something ships whose failure would cost you; the
way back walked, every time something is trusted to run alone; the permissions
tried, every time something is retired; a name and a route, every time a decision
lands on a person.</p>
<h2 id="how-you-would-know-i-am-wrong">How you would know I am wrong</h2>
<p>Every line above should be a bet, and a bet that cannot lose is not a bet. A test
has to be losable by the line being wrong rather than by the world moving, so a
clause that fires when the machine improves is a timer and not a bet. These five
are mine. Where a line here has no test yet, that is a debt I owe and not a
privilege I am claiming.</p>
<p>Verifying is misplaced when the work you checked went wrong as often as the work
you let through.</p>
<p>Earning is theatre when what you trusted on sight held up as well as what you
made climb for it.</p>
<p>Apprenticeship is sentiment when people who were never given work that could go
wrong turn out able to make the same calls unaided as the people who were.</p>
<p>Refusing is posturing when the work you came closest to turning down, and built
anyway, left the people it touched more able to act than your safer work did.</p>
<p>The reachable person is a courtesy when the people who had a decision go against
them used it and said the wait was not worth it.</p>
<p>When one of these arrives, the failure counts against the line and not against
you for having believed it. Drop it and say so.</p>
<h2 id="how-this-dies">How this dies</h2>
<p>The last manifesto did not die. It won, and was hollowed out, which is worse. It
spread as vocabulary while behaviour stayed home, and its heirs made it immune,
so no failure could count against it. A method that explains every failure as
incomplete adoption has arranged never to learn.</p>
<p>So if these words arrive and no decision changes, they failed there. And if most
of this does not survive contact with your organization, it was wrong about your
organization. A manifesto that cannot be wrong is not a manifesto. It is a
liturgy.</p>
<p>Correct me rather than quote me.</p>
]]></content:encoded></item><item><title>Power Belongs on the AI Roadmap</title><link>https://lawzava.com/blog/2026-08-06-power-belongs-on-the-ai-roadmap/</link><pubDate>Thu, 06 Aug 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-08-06-power-belongs-on-the-ai-roadmap/</guid><description>Power, not models or GPUs, is the binding constraint. Treat energy and capacity as roadmap dependencies with real lead times.</description><content:encoded><![CDATA[<p>A large new grid connection is a multi-year request. The queue to energize big new load is measured in years, not quarters, and the transformers a substation needs are ordered on lead times of a year or more. You will never file one of those orders; the cloud you buy inference from does, and the megawatts stay someone else&rsquo;s problem right up until that schedule becomes yours.</p>
<p>The lead time does not disappear when you outsource it; it reaches you renamed. A quota request that takes weeks and comes back denied. A reserved-capacity contract your provider will only sign for the region where the power already exists. A new accelerator generation that is available in three regions and waitlisted in the rest. Scarce capacity lands first where the grid is already built and trickles outward from there, so the map of where you can actually run is the map of where the power already sits. You did not secure that power; you inherited its schedule, and that schedule is unforgiving in a way software schedules are not. Assume the GPUs will be there and you have assumed away a constraint nobody upstream can wave away either.</p>
<p>Here is the shape of the miss. A team commits a 5x ramp to the board: volumes, latency targets, a confident curve. The model works in the demo, the math closes in the spreadsheet. Then the data-residency commitment pins the workload to one region, the provider&rsquo;s on-demand quota in that region is capped, the reserved-capacity deal that would lift the cap is a  <a href="/blog/2026-06-09-ai-vendor-negotiation-playbook/"
   
   >negotiation measured in months and multi-year terms</a>
, and the ramp you already sold slips two quarters. That is not a forecasting miss. It is a dependency you mislabeled as an assumption.</p>
<p>So price it like a dependency. On the roadmap, every committed ramp carries, next to its volume and latency targets:</p>
<ul>
<li>the region it actually lands in, and the residency rule that pins it there;</li>
<li>whether that capacity is on-demand (no guarantee) or reserved (contracted), and the lead time to convert one to the other;</li>
<li>the named owner of the provider relationship who can escalate a quota or open a region;</li>
<li>the fallback region, with the latency and residency cost of failing over to it.</li>
</ul>
<p>The reserved-capacity line is the nearest thing you hold to a power-procurement order: contracted capacity you claim well ahead of need, not when the quota is already capped. That is how a  <a href="/blog/2026-05-28-ai-roadmaps-survive-reality/"
   
   >roadmap survives contact with reality</a>
: it names the capacity constraint before the board discovers it.</p>
<p>Now stop spending the scarce capacity on work that does not need it. The premium path, frontier reasoning on the newest accelerators in the few regions that have them, should be rationed by design. Most production volume is high-frequency, low-stakes work that runs fine on a smaller model, a batched job, a cached result, or last-generation hardware that is abundant precisely because everyone else moved on.  <a href="/blog/2024-03-18-multi-model-strategies/"
   
   >Routing that volume down</a>
 buys back ramp without buying more capacity, and it is the one lever here you pull directly rather than petition a vendor for.</p>
<p>What remains is geography and timing. Workloads can move to where capacity is cheap and available; the commitments you made to customers usually cannot move as fast. So choose regions deliberately, hold a qualified fallback for each committed workload, and treat residency rules as the hard constraint they are rather than a detail to reconcile later. The failure mode is always the same: one assumed region, one assumed quota, one assumed timeline, and nothing staged for when any of the three moves. Stage the move before the cap forces it, because by then the cap is not negotiable and your customers already have the date.</p>
]]></content:encoded></item><item><title>Your Vendor's Balance Sheet Is Your Risk</title><link>https://lawzava.com/blog/2026-07-28-your-vendors-balance-sheet-is-your-risk/</link><pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-07-28-your-vendors-balance-sheet-is-your-risk/</guid><description>Integrate a model vendor and you inherit their balance sheet. Subsidized pricing is a repricing waiting to become your outage—price solvency as a dependency.</description><content:encoded><![CDATA[<p>Single-vendor exposure on a critical path is the risk that ends up unowned. Reliability carries the pager for uptime, not for the vendor&rsquo;s solvency; the teams shipping conversion keep deepening the dependency, because that is where the roadmap pays. So the warm fallback that would carry you through a vendor&rsquo;s bad quarter sits in no one&rsquo;s budget, and loses to conversion work until someone forces the question. Force it with diligence you can actually run.</p>
<p>Price the same million output tokens three ways before you wire a workload to a model vendor: their API, a competitor&rsquo;s API, and the same job self-hosted on an open-weight model of comparable quality, billed by the GPU-hour. You will never see the vendor&rsquo;s balance sheet. You do not need it. You need a floor.</p>
<p>Scale buys real efficiency, so a price below your own self-host cost proves nothing; a large operator running batched, high-utilization fleets genuinely serves tokens cheaper than you can. Set the floor lower, then: the cost a well-run fleet at full utilization would carry, and compare against that. When the offered price sits below even that number, no efficiency story closes the gap. The discount is being financed by someone, and financing gets withdrawn. That spread between their price and the floor is the subsidy you will repay later, on their schedule.</p>
<p>This is the diligence that works, because the diligence usually prescribed does not. &ldquo;Ask what funds the price, durable margin or capital that must be raised again&rdquo; is a question a private vendor will never answer honestly and you cannot audit. The price-versus-floor spread you can compute yourself, repeat each quarter, and act on.</p>
<h2 id="repricing-arrives-as-a-deprecation-notice">Repricing arrives as a deprecation notice</h2>
<p>It rarely shows up as &ldquo;we raised prices.&rdquo; It shows up as a sunset date. The cheap model you built on is scheduled for retirement, and the supported replacement is a reasoning model that emits three to five times the tokens to answer the same prompt. The headline per-token price holds or even falls, and your cost per call still triples, because the only model you are now allowed to call thinks out loud before it answers.  <a href="/blog/2026-07-21-token-prices-fell-ai-bills-did-not/"
   
   >Token prices fell; bills did not</a>
 is the same mechanism viewed from the invoice. The hit lands on your  <a href="/blog/2026-06-25-ai-profit-engines-unit-economics/"
   
   >unit economics</a>
 and your roadmap in the same week, when you have the least slack to absorb it.</p>
<h2 id="two-clauses-written-at-signing">Two clauses, written at signing</h2>
<p>Two terms cost nothing at signing and cannot be retrofitted during an outage: a price lock paired with a minimum notice window on any model deprecation, and an explicit right to export your fine-tunes, prompt logic, and request logs in a portable form. You are not negotiating to be reassured. You are negotiating for terms that hold through the vendor&rsquo;s bad quarter.</p>
<h2 id="the-control-is-a-warm-fallback-with-an-owner">The control is a warm fallback with an owner</h2>
<p>Comparisons and clauses narrow the risk; neither removes single-vendor exposure on a critical path. One thing does: a warm fallback, a second provider or that self-hosted model, kept able to carry the workload at degraded quality on short notice. Treat it as a  <a href="/blog/2026-04-30-ai-build-vs-buy/"
   
   >build-versus-buy portfolio rather than a binary</a>
, optionality you fund on purpose, not a hedge you improvise mid-incident.</p>
<p>Warm means it runs. Route a thin slice of real production traffic through the alternate on a fixed cadence, weekly rather than when someone remembers, so the path stays exercised and you know its true error rate before the day you need it. The question it answers is blunt: the morning your largest model vendor posts a shutdown notice, how many days until your product runs on something else? That redundancy costs real money and competes for budget, which is the point: redundancy you can measure is a line a board can fund.</p>
<p>So give the line an owner. On-call carries the failover; conversion carries the dependency that makes it necessary, and will keep deepening it until someone with budget authority puts the fallback above the next feature. Until that owner exists, the vendor sets the answer to the shutdown-notice question, on the quarter that suits them.</p>
]]></content:encoded></item><item><title>The Benchmark You Didn't Build</title><link>https://lawzava.com/blog/2026-07-16-your-benchmarks-are-lying-to-you/</link><pubDate>Thu, 16 Jul 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-07-16-your-benchmarks-are-lying-to-you/</guid><description>Public benchmarks are contaminated and gamed. The only eval that matters runs on your traffic, your failure modes, your bar—and you own it.</description><content:encoded><![CDATA[<p>Public benchmarks are useful, and the genre that tells you to ignore them is wrong. A model sitting twenty points behind the field on a task that resembles yours will not clear your bar either, so you can drop it without spending a cent. A leaderboard is a cheap negative filter and a rough ceiling on capability. What it cannot do is separate two strong candidates on the work that pays you: the gap between first and third place is usually inside the noise, and the test set has likely leaked into someone&rsquo;s training run. Shortlist on the public number. Decide on your own. A benchmark you did not build is a benchmark someone else optimized for.</p>
<p>None of that is new, which is the actual problem. &ldquo;Build your own evals&rdquo; is the most recycled advice in this field, and plenty of teams take it, stand up a set of real cases, and quietly stop trusting it within a quarter. The failure is never collecting the cases. It is the part everyone skips: who scores them, on every change, without going broke or going blind.</p>
<h2 id="the-cost-trap">The cost trap</h2>
<p>Pull thirty to a hundred real cases from your logs and weight them toward the unhappy paths. That part is easy. Then someone has to grade the outputs, and if that someone is a human reviewer working every model, prompt, and routing change, the set runs once a quarter and becomes a ritual instead of an  <a href="/blog/2026-05-05-measure-ai-progress-without-theater/"
   
   >instrument</a>
. Re-running a hand-graded set on every change is the cost objection that kills most owned-eval programs. It is real, and &ldquo;unglamorous and cheap&rdquo; does not answer it.</p>
<p>Tiering does. Most cases have a checkable answer: a required field, a JSON shape that must parse, a string that must or must not appear, a number within tolerance. Encode those as programmatic assertions. They are deterministic, free, never drift, and run in CI on every commit. In practice sixty to eighty percent of a real case set collapses to assertions. Only the residue needs judgment: was the summary faithful, was the refusal correct.</p>
<p>For that residue, use an LLM as judge, but treat the judge as versioned code, not a vibe. Pin the judge model, freeze the rubric, anchor it with a few labeled examples in the prompt. Now the expensive layer costs cents per case and runs on every change. Humans grade the grader, not the traffic: keep a small gold set you labeled by hand, and on a schedule (and whenever the judge&rsquo;s model version moves) measure judge-versus-human agreement. When agreement slips, fix the rubric before you trust another run. Auditing the judge instead of the model is what makes the cheap layer safe to lean on.</p>
<h2 id="the-noise-trap">The noise trap</h2>
<p>A set of thirty cannot tell a four-point regression from sampling noise. At an eighty percent pass rate over thirty cases, the standard error on the aggregate is about seven points, so any single-digit move is meaningless. You will chase ghosts or ship regressions while admiring a stable number. The aggregate is the wrong statistic. Run the old and new model on identical cases and read the per-case flips: which cases passed before and fail now. Three clean regressions out of thirty is a signal even when the headline barely moves, because each case is compared to itself rather than to an average. Run each case a few times so model nondeterminism does not masquerade as a change. Pairing, not volume, gives a small set teeth.</p>
<h2 id="what-it-buys">What it buys</h2>
<p>The payoff is narrow and concrete. A new model tops a public set, clears your shortlist, and your eval shows two regressions on the exact request class that quietly generates support tickets. You hold it back, with receipts, and you can tell the board why the number did not move on your terms instead of the vendor&rsquo;s. That is the whole return: not a better score, an argument you can defend. Separate it from  <a href="/blog/2024-10-14-ai-cost-benchmarking/"
   
   >cost per correct answer</a>
, which is a different decision, and the eval set becomes the one measurement a vendor cannot tune.</p>
]]></content:encoded></item><item><title>Sovereignty-by-Design for AI: How to Win Regulated Enterprise Deals</title><link>https://lawzava.com/blog/2026-07-09-sovereignty-design-ai-enterprise-deals/</link><pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-07-09-sovereignty-design-ai-enterprise-deals/</guid><description>Sovereignty is an architecture you can demonstrate, not a checklist you assert. Trust boundaries decide revenue boundaries.</description><content:encoded><![CDATA[<p>Bring-your-own-key is the strongest claim most vendors make to a regulated buyer: we hold only ciphertext, so we cannot read your data. A good auditor breaks it with a single question. How does your retrieval pipeline assemble a context window? You decrypt the record, in-region, on every request. Plaintext exists, transiently, inside your own infrastructure. The auditor has just found the seam, and you will spend the rest of the call rebuilding trust you spent for nothing.</p>
<p>Make the honest claim instead, because it is the one that survives the room. Customer-held keys do not make you subpoena-proof; a service that decrypts on request can be compelled to capture what it decrypts. What BYOK actually buys is narrower and more defensible: the customer holds the kill switch, every key use is logged on their side, and every copy at rest stays opaque without their KMS. That is revocation leverage and an audit trail, not cryptographic immunity. Sell the first as the second and the auditor stops believing your other answers too.</p>
<p>The fix for the questionnaire is not a longer policy. It is an architecture where the answers are forced by construction. Trace one record through it.</p>
<p>A customer record hits ingestion and is classified before it is stored: tagged with a data class at the boundary, the tag traveling as metadata through every downstream system. Untagged data does not pass. It is encrypted under the customer&rsquo;s key, held in their KMS. When the application needs plaintext, to build a retrieval context, it requests a short-lived scoped token, decrypts in the record&rsquo;s own region (data tagged EU-resident never reaches non-EU hardware), uses the plaintext for that one request, and drops it. A no-train flag rides along, enforced in the pipeline so the record is excluded from fine-tuning and eval sets by code, not by a reviewer&rsquo;s good intentions.</p>
<p>Deletion is where most sovereignty pitches quietly contradict themselves, promising both instant crypto-shredding and a multi-store delete fan-out as if they were one mechanism. They are two, and they cover two different populations.</p>
<p>Everything encrypted under the customer key (primary store, snapshots, the immutable backups you cannot selectively edit) is handled by revocation. You do not run a delete job against an append-only backup; you can&rsquo;t, that is what append-only means. Revoke the key version and the ciphertext is unrecoverable in place, backups included, precisely because you never needed write access to them.</p>
<p>Revocation does nothing for plaintext that escaped the envelope: lines in application logs, embeddings derived from decrypted text, a copy a subprocessor holds under its own key. None were encrypted under the customer key, so none shred with it. Those need an actual delete, per store, each returning a signed receipt with its timestamp and key version.</p>
<p>So the engineering goal is to shrink that second population until it nearly vanishes: log token references instead of plaintext, encrypt the vector index under the same customer key so embeddings die on revocation, push subprocessors onto the same custody so their copies do too. The more your deletion story is &ldquo;revoke one key,&rdquo; the less of it is &ldquo;trust that six jobs ran.&rdquo;</p>
<p>This is also why sovereignty is cheaper built early than retrofitted. Key custody, in-region routing, and a small escaped-plaintext surface are load-bearing assumptions. Bolt them onto a system that already copies data freely and you are rewriting your data plane under a customer&rsquo;s deadline. Treat  <a href="/blog/2026-04-06-sovereign-systems-privacy-non-optional/"
   
   >residency and custody as non-optional from the first commit</a>
, and  <a href="/blog/2026-05-07-ai-governance-without-bureaucracy/"
   
   >governance becomes something you exhibit</a>
 rather than a process you defend.</p>
<p>The deletion story is the part an auditor can actually test, and the test is one record&rsquo;s deletion evidence. If you cannot say which copies revocation kills and which copies it doesn&rsquo;t, neither can your buyer&rsquo;s auditor, and that gap, not the questionnaire, is why the deal stalls. &ldquo;We will delete your data on request&rdquo; is a sentence in an MSA. The answer that clears the room has the evidence attached to it: these two stores still hold plaintext derivatives, here is each delete receipt, and everything else went dark when we revoked key version 47.</p>
]]></content:encoded></item><item><title>The AI Strategy Stack: What Boards Mistake for Moats</title><link>https://lawzava.com/blog/2026-06-30-ai-strategy-stack-boards-mistake-moats/</link><pubDate>Tue, 30 Jun 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-06-30-ai-strategy-stack-boards-mistake-moats/</guid><description>Most AI moat claims are distribution theater; durable moats come from routing economics, proprietary workflow data, and operational reliability.</description><content:encoded><![CDATA[<p>The strongest argument against this whole essay is short: foundation models keep getting better, so whatever gap your proprietary data closes this quarter, the next base model closes for free. If that were fully true, no data moat in AI would be worth funding. It is half true. The half it gets wrong is the half boards keep paying for.</p>
<p>Start with what the strategy deck stacks up as defensible. Three layers, usually. The model itself, rented, and your competitor can rent the same one. The scaffolding on top of it: prompt library, routing logic, eval harness, all shaped around one vendor&rsquo;s behavior and all of it breaks the morning you change providers. <strong>If the moat disappears when the vendor changes, it was never a moat. It was a dependency.</strong> That clears two of the three layers off the slide. Swap the provider in your head; whatever still works the next morning is the only candidate worth the word.</p>
<p>What survives is the third layer, and it is the one boards cannot tell apart from its imitation: data your own operation produces by running. Here is the mechanism, and the exact place it breaks.</p>
<p>Take support automation. The model drafts a resolution; a human approves before it ships. Every rejection or rewrite captures a labeled triple: the input, the output the model produced, the output the human accepted. Not a log line. A graded example of where your model was wrong and what right looked like, on your tickets, in your domain.</p>
<p>Now the load-bearing step, the one most decks wave through. How does that triple make a cheaper model tier handle a class it used to escalate? Two mechanisms, two different bills.</p>
<p>Retrieval, the few-shot route: index the accepted exemplars and inject the nearest ones into the prompt at inference. Cheap to stand up, live the moment you index a correction, but it taxes every call in tokens and latency, and the lift is capped because you are renting the base model&rsquo;s in-context learning.</p>
<p>Distillation, the fine-tune route: train the small model on the correction set. Latency stays flat, the behavior is baked in, per-call cost drops, but you pay a training-and-eval cycle up front and re-pay it on every base-model upgrade.  <a href="/blog/2024-03-18-multi-model-strategies/"
   
   >Which tier absorbs which class</a>
 is a cost decision, not a model-quality one: retrieval for the long tail, distillation for the high-volume classes once they stop drifting.</p>
<p>Either way, one number tells you which you have: escalation rate on a single request class, quarter over quarter. Falling and sustained while quality holds is the loop compounding. Flat is a warehouse with a dashboard bolted to it. A logging pipeline and a compounding loop look identical in the architecture diagram and behave nothing alike in the P&amp;L.</p>
<p>I cannot hand you a rival&rsquo;s P&amp;L to prove the good case, and any deck that shows you a clean before-and-after percentage is selling the illustration as the evidence. The honest test is one you run on your own numbers: name the class, name the two quarters it improved, name why a competitor on the same vendor cannot reproduce it. The answer to the last one is never the model. It is the  <a href="/blog/2026-05-14-build-the-system-the-model-cannot-break/"
   
   >system around it</a>
 that turns each rejection into an exemplar only you hold.</p>
<p>Then the part the optimistic version omits: this asset depreciates. When the next base model ships, it absorbs your easy classes for free, everyone&rsquo;s, not only yours, and that compresses the set of failures where your corrections still move the number. Your edge is only ever the residual: corrections illegible outside your context, your product&rsquo;s quirks, your contractual edge cases, your policy language. The vendor will productize the capture loop; they already sell feedback buttons and fine-tuning APIs. What they cannot aggregate is a residual that means nothing without your business wrapped around it. The loop compounds only while you generate domain-specific corrections faster than a better base model erases the generic ones.</p>
]]></content:encoded></item><item><title>From Model Demos to Profit Engines: The CTO Playbook for AI Unit Economics</title><link>https://lawzava.com/blog/2026-06-25-ai-profit-engines-unit-economics/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-06-25-ai-profit-engines-unit-economics/</guid><description>AI value is won in routing and failure-cost control, not in picking a single “best” model.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>A beautiful demo is not a business model. It only proves the model can look useful before the business pays for edge cases. The bill arrives when the system hits real users, real load, and real failure conditions. At that point AI stops being a model-selection problem and becomes a  <a href="/blog/2024-03-18-multi-model-strategies/"
   
   >routing problem</a>
, a fallback problem, and a repair problem. Good CTOs do not buy &ldquo;smart.&rdquo; They buy systems that stay cheap enough, predictable enough, and reliable enough to survive the week.</p>
<h2 id="unit-economics-start-with-routing">Unit economics start with routing</h2>
<p>The wrong AI architecture sends every request to the most expensive path. That feels elegant until the invoice arrives. Mature systems route by value and by risk.</p>
<p>A practical routing model usually splits work into classes:</p>
<ul>
<li>trivial tasks that should  <a href="/blog/2026-03-09-the-end-of-fat-cloud-agentic-economy/"
   
   >stay cheap and local</a>
</li>
<li>medium-value tasks that deserve a balanced model tier</li>
<li>high-stakes tasks that justify expensive reasoning and stronger checks</li>
</ul>
<p>This is not model worship. It is cost discipline.</p>
<h2 id="the-hidden-cost-is-rarely-the-model-line-item">The hidden cost is rarely the model line item</h2>
<p>Teams fixate on tokens because tokens are visible. The real bill sits around the model: retries, context assembly, human correction, support escalation, and the work of proving the output is acceptable.</p>
<p>If a system saves one minute for a customer and creates two minutes of cleanup, it is destroying margin.</p>
<p>A finance-aware CTO should be able to answer these questions without hand-waving:</p>
<ul>
<li> <a href="/blog/2024-10-14-ai-cost-benchmarking/"
   
   >what each class of request costs to serve</a>
</li>
<li>where the rework happens</li>
<li>what failure costs when the model is wrong</li>
<li>which parts of the workflow justify premium inference</li>
</ul>
<h2 id="the-real-decision-is-not-model-choice-it-is-failure-cost">The real decision is not model choice, it is failure cost</h2>
<p>&ldquo;Best model&rdquo; is usually the wrong conversation. The useful conversation is about failure cost.</p>
<p>A cheaper model that fails gracefully can beat a more expensive model that fails silently. A  <a href="/blog/2026-05-14-build-the-system-the-model-cannot-break/"
   
   >local fallback</a>
 that keeps the system alive during a rate-limit event can matter more than a small quality lift in the happy path.</p>
<p>The CTO playbook is simple: optimize the whole system, not the benchmark screenshot.</p>
<h2 id="measure-margin-at-the-workflow-level">Measure margin at the workflow level</h2>
<p>The right unit of measure is the workflow, not the model call.</p>
<p>Ask:</p>
<ul>
<li>how much does this workflow cost end to end?</li>
<li>how often does it need human repair?</li>
<li>how long does it take to reach a trustworthy answer?</li>
<li>what is the revenue or labor value of the result?</li>
</ul>
<p>That is where the business truth lives. A model that looks slightly less accurate in isolation may create better margin if it is cheaper, faster, and easier to trust.</p>
<h2 id="a-practical-threshold">A practical threshold</h2>
<p>If the system does not improve margin, then it needs to improve risk or speed. If it improves neither, it is a demo that escaped the lab.</p>
<p>AI work that  <a href="/blog/2026-04-16-ai-capital-allocation-what-to-stop-funding/"
   
   >survives budget review</a>
 answers one of four questions:</p>
<ul>
<li>does it lower cost per task?</li>
<li>does it reduce human labor?</li>
<li>does it increase throughput?</li>
<li>does it unlock new revenue with acceptable risk?</li>
</ul>
<p>If not, the demo should stay in the demo lane.</p>
<h2 id="key-takeaways">Key Takeaways</h2>
<ul>
<li>Route cheap work cheaply.</li>
<li>Model cost is only part of the bill.</li>
<li>Measure workflow margin, not call cost.</li>
<li>If it does not improve  <a href="/blog/2026-04-28-margin-risk-speed-ai-strategy-metrics/"
   
   >margin, risk, or speed</a>
, it does not belong in production.</li>
</ul>
]]></content:encoded></item><item><title>The New Talent Stack: Product, Platform, and Applied AI Must Work as One System</title><link>https://lawzava.com/blog/2026-06-18-new-talent-stack-for-ai-organizations/</link><pubDate>Thu, 18 Jun 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-06-18-new-talent-stack-for-ai-organizations/</guid><description>AI organizations create leverage when product, platform, and applied AI are designed as one operating system instead of three kingdoms.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Most  <a href="/blog/2024-12-02-building-ai-teams/"
   
   >AI hiring plans</a>
 are trying to fix an interface problem with resumes.</p>
<p>If product, platform, and applied AI are not built as one operating system, new headcount adds motion but not leverage. The constraint is usually not talent scarcity. It is system design.</p>
<h2 id="recruiting-alone-cannot-fix-a-broken-stack">Recruiting Alone Cannot Fix a Broken Stack</h2>
<p>AI organizations often describe their issue as “we need stronger talent.” In many cases, they already have capable people. What they lack is a clear operating contract across teams.</p>
<p>The pattern is familiar:</p>
<ul>
<li>product optimizes for release velocity</li>
<li>platform optimizes for reliability and control</li>
<li>applied AI optimizes for model behavior and evaluation quality</li>
</ul>
<p>Each goal is rational. The breakdown happens at the handoffs.</p>
<p>When interfaces are unclear, every launch becomes a negotiation. When interfaces are explicit, the same teams produce compounding output.</p>
<h2 id="the-three-layer-talent-stack">The Three-Layer Talent Stack</h2>
<p>A healthy stack has three interlocking layers with distinct responsibilities:</p>
<ol>
<li><strong>Product</strong> — owns user outcomes and business success metrics.</li>
<li><strong>Platform</strong> — owns safe defaults, deployment paths, and observability.</li>
<li><strong>Applied AI</strong> — owns workflow behavior, retrieval/prompting/routing choices, and evaluation quality.</li>
</ol>
<p>These are not departments in competition. They are system components with different jobs.</p>
<p>If product outruns platform, quality debt accumulates.
If platform outruns product, infrastructure becomes generic overhead.
If applied AI outruns both, you get technically impressive demos that never operationalize.</p>
<h2 id="where-organizations-usually-break">Where Organizations Usually Break</h2>
<p>Most failures are boundary failures, not individual failures.</p>
<p>Common symptoms:</p>
<ul>
<li>no explicit owner for the model-to-product handoff</li>
<li>platform operating as a  <a href="/blog/2026-05-14-why-ai-platform-teams-become-bottlenecks/"
   
   >ticket queue instead of an enablement layer</a>
</li>
<li>applied AI measured by demo novelty instead of  <a href="/blog/2026-05-19-stop-building-internal-ai-tools-no-one-uses/"
   
   >adoption in live workflows</a>
</li>
<li>product committing features that infra cannot support safely</li>
</ul>
<p>A concise diagnosis: <strong>org debt is usually interface debt with better branding.</strong></p>
<h2 id="design-the-stack-intentionally">Design the Stack Intentionally</h2>
<p>The fix is not “more syncs.” The fix is explicit decision rights.</p>
<ul>
<li>product owns problem selection and business tradeoffs</li>
<li>platform owns reliability guardrails and release safety</li>
<li>applied AI owns workflow performance and  <a href="/blog/2026-04-23-ai-evaluation-maturity/"
   
   >evaluation integrity</a>
</li>
<li>leadership owns  <a href="/blog/2026-06-10-ai-leadership-bench-roles-interfaces/"
   
   >escalation rules</a>
 when tradeoffs conflict</li>
</ul>
<p>Once this is explicit, hiring quality improves. You stop searching for mythical generalists and start  <a href="/blog/2026-05-26-hiring-operators-for-ai-teams/"
   
   >hiring operators</a>
 who can perform inside a coherent system.</p>
<h2 id="what-to-evaluate-before-adding-headcount">What to Evaluate Before Adding Headcount</h2>
<p>Before opening new roles, run this short check:</p>
<ol>
<li>Are cross-team handoffs documented and current?</li>
<li>Does each layer have clear success metrics it actually controls?</li>
<li>Are escalation paths clear when speed, reliability, and quality disagree?</li>
<li>Are teams rewarded for system outcomes rather than local optimization?</li>
</ol>
<p>If those answers are weak, fix interfaces first. New hires will scale the current operating model, good or bad.</p>
<h2 id="key-takeaways">Key Takeaways</h2>
<ul>
<li>Strong AI organizations are designed as a system, not staffed as silos.</li>
<li>Product, platform, and applied AI need explicit interfaces and decision rights.</li>
<li>Boundary clarity is a bigger lever than raw headcount.</li>
<li>Hiring works best after the operating contract is clear.</li>
</ul>
]]></content:encoded></item><item><title>The Post-Prototype AI Org: Operating Models That Survive Year Two</title><link>https://lawzava.com/blog/2026-06-10-post-prototype-ai-org/</link><pubDate>Wed, 10 Jun 2026 12:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-06-10-post-prototype-ai-org/</guid><description>Year-two AI failure usually comes from org-design mismatch, not model-quality mismatch. The handoffs are where the system slows down.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>A lot of AI orgs look healthy in month three and brittle by year two. The model usually did not fail. The operating model did. Prototype energy is easy to create; durable coordination is not.</p>
<p>The question is not whether the team can ship something exciting. The question is whether the company can keep shipping after the novelty fades.</p>
<h2 id="why-the-prototype-phase-hides-the-real-problem">Why the prototype phase hides the real problem</h2>
<p>In the early phase, AI teams often succeed because everyone is close to the work. Decisions are informal, context is shared, and the whole system fits in a few people’s heads. That stops scaling almost immediately.</p>
<p>As soon as the team grows, the same strengths turn into liabilities:</p>
<ul>
<li>knowledge becomes hidden</li>
<li>approvals multiply</li>
<li>handoffs slow down</li>
<li>nobody owns the  <a href="/blog/2026-06-10-ai-leadership-bench-roles-interfaces/"
   
   >interface boundaries</a>
</li>
</ul>
<p>What worked when the team was small no longer works when the company needs predictability.</p>
<h2 id="the-operating-model-should-be-explicit">The operating model should be explicit</h2>
<p>A post-prototype AI org needs to define how work moves.</p>
<p>The model should answer:</p>
<ul>
<li>who owns the user problem?</li>
<li>who owns the runtime?</li>
<li>who owns the quality signal?</li>
<li>who owns the  <a href="/blog/2026-05-07-ai-governance-without-bureaucracy/"
   
   >risk boundary</a>
?</li>
<li>who can stop the release?</li>
</ul>
<p>Without those answers, the team is improvising around gaps that will eventually become incidents or delays.</p>
<h2 id="handoffs-are-the-hidden-bottleneck">Handoffs are the hidden bottleneck</h2>
<p>Most  <a href="/blog/2026-05-28-ai-roadmaps-survive-reality/"
   
   >AI roadmaps</a>
 do not fail because the team lacks ideas. They fail because each handoff adds ambiguity.</p>
<p>The problem shows up in predictable places:</p>
<ul>
<li>product asks for speed, platform asks for safety</li>
<li>applied AI wants more freedom, compliance wants more proof</li>
<li>leadership wants output, the system wants more control</li>
</ul>
<p>That tension is normal. What is not normal is leaving it unresolved.</p>
<p>A good operating model turns tension into a documented interface, not a recurring crisis.</p>
<h2 id="scale-requires-less-heroics-not-more">Scale requires less heroics, not more</h2>
<p>The post-prototype org has to depend less on heroic behavior and more on repeatable behavior.</p>
<p>That usually means:</p>
<ul>
<li>clearer ownership</li>
<li> <a href="/blog/2026-06-10-decision-latency-p-and-l-variable/"
   
   >smaller decision surfaces</a>
</li>
<li>stronger  <a href="/blog/2026-04-23-ai-evaluation-maturity/"
   
   >eval gates</a>
</li>
<li>visible  <a href="/blog/2026-05-14-build-the-system-the-model-cannot-break/"
   
   >rollback paths</a>
</li>
<li>fewer ambiguous exceptions</li>
</ul>
<p>This can feel slower at first, but it is the only way the org gets faster at scale.</p>
<h2 id="a-simple-test">A simple test</h2>
<p>Ask whether the AI system can survive a senior person going on vacation for two weeks.</p>
<p>If the answer is “not really,” the organization is still running on hidden tribal knowledge.</p>
<p>If the answer is “yes, with documented ownership and a stable operating model,” the company is moving from prototype to production.</p>
<p>That is the real year-two test.</p>
<h2 id="key-takeaways">Key Takeaways</h2>
<ul>
<li>Prototype energy does not scale on its own.</li>
<li>The year-two problem is usually organizational, not model-related.</li>
<li>Ownership, interfaces, and escalation paths matter more than the demo itself.</li>
<li>A durable AI org is designed for scale before the prototype succeeds.</li>
</ul>
]]></content:encoded></item><item><title>Decision Latency as a P&amp;L Variable: The Leadership Metric Nobody Owns</title><link>https://lawzava.com/blog/2026-06-10-decision-latency-p-and-l-variable/</link><pubDate>Wed, 10 Jun 2026 09:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-06-10-decision-latency-p-and-l-variable/</guid><description>Decision latency is measurable and should be treated as a direct cost driver.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Slow decisions look like caution. In practice, they are hidden expense.</p>
<p>Decision latency belongs on the P&amp;L. Every day a real decision sits unresolved, the business pays in delay, rework, and attention.</p>
<h2 id="why-decision-latency-matters">Why Decision Latency Matters</h2>
<p>A team can  <a href="/blog/2026-05-05-measure-ai-progress-without-theater/"
   
   >look productive</a>
 and still be dragging the business down if every meaningful decision takes too long.</p>
<p>Decision latency shows up as:</p>
<ul>
<li>stalled launches</li>
<li>expired opportunities</li>
<li>duplicated work</li>
<li>growing frustration in the teams closest to the customer</li>
</ul>
<p>When leaders do not measure this, they blame execution when the real problem is delay. The work may be moving. The organization is not.</p>
<h2 id="what-decision-latency-looks-like-in-practice">What Decision Latency Looks Like in Practice</h2>
<p>You can usually find it by asking a few questions:</p>
<ul>
<li>How long does a high-signal issue sit before someone decides?</li>
<li>How many people need to weigh in before the first answer exists?</li>
<li>How often do decisions get reopened because no one owned the original call?</li>
<li>How much work is blocked waiting for alignment that never arrives?</li>
</ul>
<p>Those are not soft questions. They are economic questions.</p>
<p>If a release,  <a href="/blog/2026-05-26-hiring-operators-for-ai-teams/"
   
   >hiring decision</a>
,  <a href="/blog/2026-06-09-ai-vendor-negotiation-playbook/"
   
   >vendor decision</a>
, or  <a href="/blog/2026-06-02-ai-incident-review-changes-architecture/"
   
   >architecture decision</a>
 sits for weeks, the business is paying rent on uncertainty.</p>
<p>A useful line: <strong>ambiguous ownership is the most expensive architecture in your company.</strong></p>
<h2 id="make-it-visible">Make It Visible</h2>
<p>If you want leaders to care, make the metric visible.</p>
<p>Track:</p>
<ul>
<li>time from issue raised to decision made</li>
<li>time from decision made to action taken</li>
<li>number of escalations per decision class</li>
<li>number of decisions reopened after approval</li>
</ul>
<p>Once those numbers are in the open, patterns become hard to deny. You can see which teams move fast, which questions keep getting rerouted, and where the organization is burning time on decisions that should have been routine.</p>
<h2 id="how-to-reduce-it">How to Reduce It</h2>
<p>Decision latency drops when teams do four things well:</p>
<ol>
<li>Define  <a href="/blog/2026-06-10-ai-leadership-bench-roles-interfaces/"
   
   >who owns each decision class</a>
.</li>
<li>Set  <a href="/blog/2026-05-07-ai-governance-without-bureaucracy/"
   
   >decision boundaries</a>
 before the crisis.</li>
<li>Reduce the number of people required for routine calls.</li>
<li>Make escalation fast when the decision is truly material.</li>
</ol>
<p>This is not about making every decision unilateral. It is about making routine decisions quick and risky decisions explicit.</p>
<p>If the call is small, the system should move. If the call is material, the system should know exactly who has to weigh in.</p>
<h2 id="key-takeaways">Key Takeaways</h2>
<ul>
<li>Decision latency is a real cost driver.</li>
<li>Measure the time from issue to decision and from decision to action.</li>
<li>Ownership clarity reduces hidden opex.</li>
<li>The best organizations make routine decisions quickly and unusual decisions deliberately.</li>
</ul>
]]></content:encoded></item><item><title>The AI Vendor Negotiation Playbook for CTOs</title><link>https://lawzava.com/blog/2026-06-09-ai-vendor-negotiation-playbook/</link><pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-06-09-ai-vendor-negotiation-playbook/</guid><description>Vendor leverage in AI comes from architecture readiness, eval data, and exit credibility — not procurement theater.</description><content:encoded><![CDATA[<p>Use this before any AI vendor contract renewal, initial procurement, or pricing negotiation. Most CTOs walk in under-prepared — the vendor knows your dependency footprint better than you do. This worksheet closes that gap. Work through it the day before the meeting.</p>
<hr>
<h2 id="1-workload-facts-you-must-have">1. Workload Facts You Must Have</h2>
<p>The vendor’s first move is to define your usage for you. Don’t let them.</p>
<ul>
<li><input disabled="" type="checkbox"> Total request volume per month, broken out by use case
<em>A single aggregate number is not enough. Know which workflows drive cost.</em></li>
<li><input disabled="" type="checkbox"> Cost per task class (e.g., generation vs. classification vs. retrieval)
<em>If you cannot name your top three cost drivers, you cannot challenge the invoice.</em></li>
<li><input disabled="" type="checkbox"> Latency p50/p95 by workflow, measured from  <a href="/blog/2025-03-31-ai-observability-deep/"
   
   >your own instrumentation</a>

<em>Vendor SLAs are measured at their edge, not yours.</em></li>
<li><input disabled="" type="checkbox"> Percentage of spend attributable to this vendor vs.  <a href="/blog/2026-04-16-ai-capital-allocation-what-to-stop-funding/"
   
   >total AI budget</a>

<em>Concentration creates leverage — for them. Know the number.</em></li>
<li><input disabled="" type="checkbox"> Named owner of the vendor relationship on your side
<em>If no one owns it, no one negotiates it.</em></li>
</ul>
<h2 id="2-architecture-leverage-check">2. Architecture Leverage Check</h2>
<p>Leverage is an architecture property. Answer these before you sit down.</p>
<ul>
<li><input disabled="" type="checkbox"> Is the vendor’s API called directly from product code, or through an  <a href="/blog/2024-03-18-multi-model-strategies/"
   
   >abstraction layer</a>
?
<em>Direct calls = switching costs measured in months. Abstraction = measured in days.</em></li>
<li><input disabled="" type="checkbox"> How many distinct integration points does this vendor touch?
<em>Write the number. Fewer than five is manageable. More than ten is a dependency.</em></li>
<li><input disabled="" type="checkbox"> What is the estimated engineering cost to swap this vendor?
<em>Get a real estimate, even a rough one. &ldquo;Unknown&rdquo; is not an answer.</em></li>
<li><input disabled="" type="checkbox"> Do you have a secondary provider you have already integrated, even partially?
<em>Yes/No. If no, you have no credible threat.</em></li>
<li><input disabled="" type="checkbox"> Does your data pipeline depend on vendor-specific formats or  <a href="/blog/2023-07-10-embedding-models-deep-dive/"
   
   >embeddings</a>
?
<em> <a href="/blog/2026-05-14-build-the-system-the-model-cannot-break/"
   
   >Format lock-in</a>
 is often more expensive than API lock-in.</em></li>
</ul>
<h2 id="3-evaluation-evidence">3. Evaluation Evidence</h2>
<p>Vendors sell on benchmark claims. Counter with your data.</p>
<ul>
<li><input disabled="" type="checkbox"> Do you have  <a href="/blog/2026-04-23-ai-evaluation-maturity/"
   
   >evals that measure model performance on your actual workload</a>
?
<em>Yes/No. If no, you are buying on their terms by default.</em></li>
<li><input disabled="" type="checkbox"> Which models have you tested against your task suite in the last 90 days?
<em>List them. If the answer is only theirs, you have no comparison point.</em></li>
<li><input disabled="" type="checkbox"> What is your acceptable quality threshold, defined numerically?
<em>&ldquo;Good enough&rdquo; is not a threshold. A number is.</em></li>
<li><input disabled="" type="checkbox"> Have you run a  <a href="/blog/2024-10-14-ai-cost-benchmarking/"
   
   >cost-per-correct-output comparison</a>
 across providers?
<em>Price per token is a distraction. Price per correct result is the metric.</em></li>
<li><input disabled="" type="checkbox"> Who owns your eval framework and can demo it in the meeting if needed?
<em>Named person, not a team.</em></li>
</ul>
<h2 id="4-exit-credibility">4. Exit Credibility</h2>
<p>A vendor that believes you cannot leave does not negotiate. Make them uncertain.</p>
<ul>
<li><input disabled="" type="checkbox"> Do you have a documented migration plan, even a sketch?
<em>It does not need to be final. It needs to exist.</em></li>
<li><input disabled="" type="checkbox"> What is your contractual notice period to exit?
<em>Know this before they remind you of it.</em></li>
<li><input disabled="" type="checkbox"> Have you identified which vendor you would move to first if pricing increased 40%?
<em>Name them. Vague alternatives are not alternatives.</em></li>
<li><input disabled="" type="checkbox"> Is there a sunset timeline for any features that are vendor-exclusive today?
<em>If yes, the vendor knows your dependency has an expiration date.</em></li>
<li><input disabled="" type="checkbox"> Can your team absorb a two-week migration sprint without derailing the roadmap?
<em>Yes/No. Honest answer only.</em></li>
</ul>
<hr>
<p>If you cannot fill in the workload numbers, you are not done preparing — you are about to negotiate against someone who has already modeled your spend. If you have no eval data, you will accept their performance claims by default. If there is no exit plan, any number they name is essentially a take-it-or-leave-it offer. The meeting itself is the wrong place to discover these gaps. Thirty minutes with this worksheet before you walk in is worth more than any negotiation tactic once you are in the room.</p>
]]></content:encoded></item><item><title>How Great CTOs Design AI Roadmaps That Survive Contact With Reality</title><link>https://lawzava.com/blog/2026-05-28-ai-roadmaps-survive-reality/</link><pubDate>Thu, 28 May 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-05-28-ai-roadmaps-survive-reality/</guid><description>AI roadmaps fail when they are sequenced around ambition instead of dependency, verification, and rollback cost.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>AI roadmaps fail when ambition is treated as sequencing. Dependencies slip, rollback gets expensive, and the team discovers the missing work only after the launch date is already spoken for.</p>
<p>A survivable roadmap is not a prettier Gantt chart. It is a dependency-aware budget for uncertainty.</p>
<h2 id="roadmaps-fail-at-the-edges">Roadmaps Fail at the Edges</h2>
<p>The core mistake is treating the roadmap like a statement of intent instead of a statement of sequencing.</p>
<p>AI work fails at the edges:</p>
<ul>
<li>data access is slower than expected</li>
<li>model behavior is less stable than expected</li>
<li>review cycles take longer than expected</li>
<li>vendor changes arrive earlier than expected</li>
</ul>
<p>If your roadmap does not account for those edges, it is not a plan. It is a confidence exercise.</p>
<p>Most teams only find out those edges are missing after the launch date is already public.</p>
<p>The fix is to move the hidden work into the plan before the promise is made.</p>
<h2 id="budget-the-dependency-chain">Budget the Dependency Chain</h2>
<p>Every AI feature has a dependency chain:</p>
<ul>
<li>data availability</li>
<li> <a href="/blog/2024-07-22-context-window-strategies/"
   
   >context assembly</a>
</li>
<li> <a href="/blog/2024-03-18-multi-model-strategies/"
   
   >model routing</a>
</li>
<li>evaluation</li>
<li>deployment</li>
<li>fallback</li>
</ul>
<p>If any one of those links is not ready, the feature will not survive real use.</p>
<p>If the chain is incomplete, the roadmap is lying by omission.</p>
<p>The most honest roadmap is the one that writes the chain down first. That slows the conversation, but it also keeps the team from selling a feature that depends on work nobody has budgeted.</p>
<p>Slower conversations are cheaper than broken launches.</p>
<h2 id="make-rollback-a-first-class-requirement">Make Rollback a First-Class Requirement</h2>
<p>Good roadmaps assume the first version will be wrong.</p>
<p>That means every AI initiative should answer four questions:</p>
<ul>
<li>How do we turn this off?</li>
<li> <a href="/blog/2025-03-31-ai-observability-deep/"
   
   >How do we know it is hurting us?</a>
</li>
<li>How fast can we revert?</li>
<li>What manual path exists if the model degrades?</li>
</ul>
<p>If those answers are fuzzy, the roadmap is overconfident.</p>
<p>If you cannot turn it off quickly, you have  <a href="/blog/2026-05-14-build-the-system-the-model-cannot-break/"
   
   >shipped a liability with a product label</a>
.</p>
<p>Roadmaps should not only describe the happy path. They should budget for the probability that the first version is wrong, the vendor changes terms, or the model regresses under load.</p>
<p>That is not pessimism. It is operational seriousness.</p>
<h2 id="wip-limits-matter-more-than-hope">WIP Limits Matter More Than Hope</h2>
<p>A roadmap that promises too many parallel AI experiments is usually a roadmap that does not respect WIP.</p>
<p>The more novel the work, the lower the WIP should be.</p>
<p>Concurrency feels productive until it multiplies rework.</p>
<p>Strong teams set rules like:</p>
<ul>
<li>no more than one high-risk AI launch per squad at a time</li>
<li>no feature ships without  <a href="/blog/2026-04-23-ai-evaluation-maturity/"
   
   >evaluation coverage</a>
</li>
<li>no vendor migration without a fallback path</li>
<li>no roadmap item enters “done” until the operational notes exist</li>
</ul>
<p>That may sound strict. It is. Novel work punishes loose concurrency.</p>
<h2 id="what-a-survivable-roadmap-looks-like">What a Survivable Roadmap Looks Like</h2>
<p>Survivable roadmaps are dependency-explicit, rollback-aware, and honest about capacity.</p>
<p>A roadmap is not a promise. It is a bet with visible failure modes.</p>
<p>If the failure modes are invisible, the roadmap is pretending.</p>
<p>You do not need a roadmap that impresses the room. You need one the organization can execute without pretending the hard parts are somebody else&rsquo;s problem.</p>
<h2 id="key-takeaways">Key Takeaways</h2>
<ul>
<li>AI roadmaps fail at dependency and rollback boundaries.</li>
<li>Treat the roadmap as a budget for uncertainty, not a wish list.</li>
<li>Limit WIP, make rollback explicit, and require evaluation coverage before launch.</li>
<li>The best roadmap is the one the organization can survive.</li>
</ul>
]]></content:encoded></item><item><title>Technical Leadership in the AI Era (It’s About Throughput, Not Trends)</title><link>https://lawzava.com/blog/2026-05-21-ai-technical-leadership/</link><pubDate>Thu, 21 May 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-05-21-ai-technical-leadership/</guid><description>Technical leadership in mid-2026: anchor decisions in throughput, verification, and operability instead of chasing the latest agent framework.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>AI does not change the core job of technical leadership. It changes the cost of being vague. In 2026, the best leaders still do the same three things: set direction, remove friction, and keep production systems measurable. The difference is that AI makes every weak assumption show up faster.</p>
<p>The real mandate is throughput. Not more noise. Not more  <a href="/blog/2026-05-05-measure-ai-progress-without-theater/"
   
   >experimentation theater</a>
. Throughput.</p>
<h2 id="the-leadership-pivot-focus-on-throughput">The Leadership Pivot: Focus on Throughput</h2>
<p>Organizations do not pay technical leaders to keep up with model releases. They pay them to improve  <a href="/blog/2026-03-30-throughput-engineer-headcount-lagging-metric/"
   
   >organizational throughput</a>
.</p>
<p>That means reducing cognitive overhead, tightening verification, and making deployment paths boring enough that teams can move without drama. If you cannot measure what an AI workflow produced, or what it cost to produce it, you do not have an operating system yet. You have a prototype with invoices.</p>
<p>The leadership question is simple: are we removing blockers faster than we are adding complexity?</p>
<h2 id="decision-making-in-practice">Decision-Making in Practice</h2>
<p>AI work gets messy when teams debate tools before they define the outcome.</p>
<p>Good leaders force the conversation back to first principles:</p>
<ul>
<li>What business metric should change if we ship this?</li>
<li>What latency budget do we actually have?</li>
<li> <a href="/blog/2026-05-14-build-the-system-the-model-cannot-break/"
   
   >What happens when the model is wrong?</a>
</li>
</ul>
<p>Those questions cut through a lot of noise. They keep the team from turning architecture meetings into opinion contests about  <a href="/blog/2023-04-03-vector-databases-explained/"
   
   >vector databases</a>
, prompt styles, or the latest agent framework.</p>
<p>If the answer to any of those questions is fuzzy, the work is not ready for serious implementation.</p>
<h2 id="define-good-enough-and-measure-it">Define “Good Enough” and Measure It</h2>
<p>Reliability is not just accuracy. It is consistency, cost, and the ability to  <a href="/blog/2025-03-31-ai-observability-deep/"
   
   >catch degradation before customers do</a>
.</p>
<p>Sometimes  <a href="/blog/2024-08-05-small-models-big-impact/"
   
   >a smaller, cheaper model</a>
 is the right answer. Sometimes the frontier model is worth the price. The point is not to be religious about either option. The point is to define the bar, test against it, and choose the system that meets it with the least operational pain.</p>
<p>Your job is not to build a perfect AI system. It is to build one where failure is bounded, expected, and visible.</p>
<h2 id="the-cultural-shift">The Cultural Shift</h2>
<p>Technical leadership still has a change-management problem. Engineers will worry about ownership, safety, and the volatility of the ecosystem. Those concerns are real.</p>
<p>The right response is not debate for its own sake. It is instrumentation.</p>
<p>Stop arguing in design docs about whether a model will work. Build the telemetry that shows whether it works. Stop treating every new framework like a strategy reset. Run small, contained experiments that either produce evidence or die cheaply.</p>
<p>The strongest teams are not the ones that sprint toward the newest beta API. They are the ones that can absorb change without losing control.</p>
<h2 id="final-take">Final Take</h2>
<p>AI rewards leaders who are disciplined about outcomes and ruthless about verification. If the team can move quickly, measure clearly, and recover cleanly, AI becomes leverage. If not, it becomes another source of drag.</p>
]]></content:encoded></item><item><title>Stop Building Internal AI Tools No One Uses</title><link>https://lawzava.com/blog/2026-05-19-stop-building-internal-ai-tools-no-one-uses/</link><pubDate>Tue, 19 May 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-05-19-stop-building-internal-ai-tools-no-one-uses/</guid><description>Internal AI tools fail when teams optimize for launch instead of habit formation, trust, and workflow fit.</description><content:encoded><![CDATA[<p>The demo went well. A mid-size logistics company — roughly 800 people, enough procurement complexity to  <a href="/blog/2026-04-16-ai-capital-allocation-what-to-stop-funding/"
   
   >justify the investment</a>
 — had spent three months building an internal AI tool to surface contract terms during vendor negotiations. The launch Slack channel hit 40 reactions in the first hour. A VP called it the kind of thing that changes how the team operates.</p>
<p>Six weeks later, the channel had five messages in it, four of them automated. The procurement leads were still pulling PDFs manually and copying terms into a shared spreadsheet. One support engineer, who had quietly championed the project from the beginning, had reverted to her old database query because &ldquo;the tool doesn&rsquo;t know about the amendments.&rdquo; The tool was still running. Nobody had officially abandoned it. It had simply become invisible.</p>
<p>This pattern is not unusual. It is almost the default.</p>
<h2 id="what-actually-failed">What Actually Failed</h2>
<p>The postmortem conversation usually centers on the wrong things — model choice, interface design, rollout timing. Those are symptoms. The root causes are structural.</p>
<p>The contract tool was built around a narrow slice of the negotiation workflow: surfacing base terms. But procurement work is not base terms. It is base terms plus amendments plus prior history plus the relationship context the lead carries in her head. The tool knew one layer of a five-layer problem. It looked complete in a demo because demos are controlled. Real work is not controlled.</p>
<p>The output trust problem arrived fast. In week two, the tool surfaced an incorrect payment term — technically correct in the original contract, superseded by a signed amendment it had not been given access to. The lead caught it before it caused damage, but she stopped relying on it after that. <em>One unexplained wrong answer is enough to demote a tool from co-worker to footnote.</em> The team had not  <a href="/blog/2026-04-23-ai-evaluation-maturity/"
   
   >built evaluation into the system</a>
, so there was no way to know how often this happened, which made the uncertainty worse, not better.</p>
<p>Nobody owned adoption after the launch. The engineer who built it moved to a different priority. The VP who celebrated it never checked  <a href="/blog/2025-07-07-ai-product-metrics/"
   
   >sustained usage</a>
. When procurement leads developed workarounds, there was no one watching the signal and no one with a mandate to respond. The tool drifted.</p>
<h2 id="when-it-works">When It Works</h2>
<p>A different team at a professional services firm built something structurally simpler: a tool that drafted the engagement summary section of a client report, pulling from structured notes the consultant had already entered into their project management system. Narrow scope. No novel context required. One predictable output format, reviewed every time before it went anywhere.</p>
<p>The tool stuck. Not because it was more technically impressive — it was considerably less so. It stuck because it removed a specific, recurring task that consultants genuinely disliked, it used context they were already maintaining anyway, and the output was always human-reviewed before it mattered. The failure mode was visible and safe. The value was obvious the first time you used it and every time after.</p>
<p>The team lead reviewed usage weekly for the first two months and made three small adjustments based on what she saw. That ownership — unglamorous, persistent, post-launch — is what made the difference.</p>
<h2 id="the-structural-difference">The Structural Difference</h2>
<p>Both companies built  <a href="/blog/2025-08-04-ai-workflow-automation/"
   
   >AI tools for internal workflows</a>
. One failed quietly, one became a habit. The gap was not the model. It was not the interface. It was whether the tool was designed around how work actually moves or around  <a href="/blog/2026-05-05-measure-ai-progress-without-theater/"
   
   >what would look good in a demo</a>
.</p>
<p>Tools that survive are ones that fit a narrow, complete slice of a workflow, produce output that is either verifiable or bounded enough to trust, require no context the user does not already have, and have someone whose job it is to watch whether people are actually using them.</p>
<p>That last part is the one most teams skip. Usage is not a launch outcome. It is an operating responsibility.</p>
]]></content:encoded></item><item><title>Build the System the Model Cannot Break</title><link>https://lawzava.com/blog/2026-05-14-build-the-system-the-model-cannot-break/</link><pubDate>Thu, 14 May 2026 09:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-05-14-build-the-system-the-model-cannot-break/</guid><description>A manifesto for building AI-native organizations. Twelve tenets across strategy, architecture, economics, and people — and the only test that matters in year two.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>An AI-native company is not a company that uses AI. It is a company whose operating model — decisions, ownership, interfaces, capital, and failure boundaries — has been built so AI compounds inside it instead of evaporating around it.</p>
<p>The model will change. The system around it should not.</p>
<p>This is a manifesto. It is opinionated, deliberately. Twelve tenets, four movements, one test. Borrow what works. Argue with the rest.</p>
<hr>
<h1 id="movement-i--strategy">Movement I — Strategy</h1>
<h2 id="1-the-operating-model-is-the-strategy">1. The operating model is the strategy</h2>
<p>The model is the most expensive dependency in your stack. It is not the brain. The brain is everything you build around it: context assembly, retrieval, validation, retries, telemetry, fallback, escalation.</p>
<p>Two companies buy the same frontier model on the same Tuesday. One ships in six weeks with a deterministic fallback, a typed validator, and an eval gate on every PR. The other ships in six months with a notebook of &ldquo;good prompts&rdquo; and a Slack channel for incidents. Same model. Different company.</p>
<p>If your AI plan begins with &ldquo;which model should we buy,&rdquo; you are solving the easiest problem in the room. <strong>The moat is everything around the model.</strong></p>
<h2 id="2-capital-allocation-is-the-first-product-decision">2. Capital allocation is the first product decision</h2>
<p>Great AI teams do not start with a roadmap. They start with  <a href="/blog/2026-04-16-ai-capital-allocation-what-to-stop-funding/"
   
   >a kill list</a>
. Capital is finite. Attention is finite. Support burden is finite.</p>
<p>Three questions before any AI initiative gets funded:</p>
<ol>
<li>Does this increase <strong>margin</strong>, reduce <strong>risk</strong>, or improve <strong>speed</strong>?</li>
<li>Can we measure that effect within one to three quarters?</li>
<li>Do we own the <strong>fallback</strong> if the model or vendor changes?</li>
</ol>
<p>If the answer to all three is not yes, the default is no.</p>
<p>The most common pattern across Series B–D companies that quietly stalled in 2024–2025: somewhere between $1M and $3M of engineering and infra burned on internal copilots that never crossed adoption threshold, plus a duplicate prompt orchestration layer because two teams built one in parallel. Neither project had a measurable failure mode. Both had a sponsor.</p>
<p>A four-dimension scorecard makes the next budget meeting honest:</p>
<ul>
<li><strong>Adoption</strong> — are real users using it in a real workflow?</li>
<li><strong>Reliability</strong> — does it fail in bounded, observable ways?</li>
<li><strong>Margin</strong> — does it reduce cost or improve unit economics?</li>
<li><strong>Speed</strong> — does it shorten a real business cycle time?</li>
</ul>
<p><strong>If you cannot defend it with numbers, the project is not innovative. It is unpriced.</strong></p>
<h2 id="3-decision-latency-is-a-pl-variable">3. Decision latency is a P&amp;L variable</h2>
<p>Slow decisions look like caution. In practice, they are hidden expense. Every day a real decision sits unresolved, the business pays in delay, rework, and attention.</p>
<p>Headcount is an input.  <a href="/blog/2026-03-30-throughput-engineer-headcount-lagging-metric/"
   
   >Throughput is an outcome</a>
. Adding the tenth engineer to a system that takes nine days to approve a deploy adds nine more days of waiting, not 10% more output.</p>
<p>Track four numbers with the same seriousness as revenue:</p>
<ul>
<li>time from issue raised to decision made</li>
<li>time from decision made to action taken</li>
<li>escalations per decision class</li>
<li>decisions reopened after approval</li>
</ul>
<p><strong>Ambiguous ownership is the most expensive architecture in your company.</strong></p>
<hr>
<h1 id="movement-ii--architecture">Movement II — Architecture</h1>
<h2 id="4-build-firewalls-not-masterpieces">4. Build firewalls, not masterpieces</h2>
<p>A statistical engine cannot be expected to behave like deterministic infrastructure. If your architecture only works when the model is correct 100% of the time, it is not architecture. It is wishful thinking with a demo budget.</p>
<p>Three failure modes, three firewalls. They are not the same thing and they are not solved by the same code:</p>
<ul>
<li><strong>Inbound sanitization.</strong> What data is permitted into the prompt context. PII strippers, schema enforcers, retrieved-document trust scoring. This is also where indirect prompt injection — instructions hidden in a vendor PDF, a customer message, or a tool output — gets caught before it reaches the model.</li>
<li><strong>Outbound validation.</strong> A typed schema checker stands between the model and the operational database. Malformed JSON, out-of-range values, and policy-violating outputs are rejected at the boundary, not absorbed by downstream services.</li>
<li><strong>Operational fallback.</strong> Circuit breakers for vendor outages and rate limits. If the model returns invalid output three times in a row, the system degrades to a deterministic path — not a stack trace in front of the user.</li>
</ul>
<p>Each of these is a separate piece of code with a separate owner, a separate test surface, and a separate failure mode. A &ldquo;kill switch&rdquo; that catches all three is a slide, not a system.</p>
<p><strong>You cannot prompt your way out of entropy. You have to architect your way out of it.</strong></p>
<h2 id="5-evaluation-is-the-spine">5. Evaluation is the spine</h2>
<p>If you cannot define an eval suite before shipping a feature, you do not understand the system well enough to ship it.</p>
<p>A  <a href="/blog/2026-04-23-ai-evaluation-maturity/"
   
   >five-level maturity ladder</a>
:</p>
<ol>
<li><strong>Vibes-based.</strong> Someone eyeballs prompts before release.</li>
<li><strong>Spreadsheet.</strong> Suite exists, runs occasionally, blocks nothing.</li>
<li><strong>CI/CD-integrated.</strong> Evals run on every PR. A failed gate stays failed.</li>
<li><strong>Continuous telemetry.</strong> Production samples scored asynchronously. Incidents become regression tests.</li>
<li><strong>Governance as moat.</strong> Evaluation shapes architecture before code. Margin, latency, and sovereignty tradeoffs are quantified, not asserted.</li>
</ol>
<p>Below Level 3 is not a production system. It is a demo with a pager.</p>
<p>Level 4 is where most organizations get stuck, and the reason is rarely effort. Judge models drift, ground truth ages, sampling bias creeps in, and your asynchronous scoring quietly stops tracking the failure mode you cared about. Mature teams hold a small, hand-labeled golden set as the anchor, treat the judge model as a versioned dependency, and re-calibrate when either changes.</p>
<p>Eval portability is a year-two survival trait. If your eval suite is hand-tuned to one model&rsquo;s tokenizer and one vendor&rsquo;s output quirks, you have not built an eval suite. You have built a benchmark for the model you are about to be unable to leave.</p>
<h2 id="6-agentic-systems-run-on-a-reliability-contract">6. Agentic systems run on a reliability contract</h2>
<p>Agents are not magical workers. They are autonomous systems with more ways to fail. The reliability discipline gets stricter, not looser.</p>
<p>Every production agent answers five questions in one meeting, without hand-waving:</p>
<ul>
<li>what is it allowed to do?</li>
<li>what is it explicitly not allowed to do?</li>
<li>what metrics prove it is healthy?</li>
<li>what happens when the model degrades?</li>
<li>who can stop it, and how fast?</li>
</ul>
<p>But the five questions are a meeting checklist. The contract is a published artifact with <strong>SLOs, blast-radius caps in dollars or rows or API calls, rollback latency targets, and a named owner per failure mode.</strong> Blast radius is the real design variable: data scope, action scope, time scope, permission scope, fallback scope.</p>
<p>Kill switches are not weakness. They are governance that can move faster than the failure. A useful test of any AI control: <strong>could an engineer follow this rule at 2 a.m. without calling a committee?</strong></p>
<p>A roadmap that ships an agent without answers to these questions is a roadmap that has shipped a liability with a product label. Every initiative names how it turns off, how it knows it is hurting, how fast it reverts, and what manual path exists when the model degrades.</p>
<p><em>Companion:  <a href="/docs/agent-reliability-contract"
   
   >Agent Reliability Contract template</a>
.  <a href="/docs/rollback-template"
   
   >Rollback document template</a>
.</em></p>
<p><strong>Autonomy without a reliability contract is just an incident waiting for a timeline.</strong></p>
<hr>
<h1 id="movement-iii--economics--externals">Movement III — Economics &amp; Externals</h1>
<h2 id="7-unit-economics-live-at-the-workflow-not-the-model-call">7. Unit economics live at the workflow, not the model call</h2>
<p>Teams fixate on tokens because tokens are visible. The real bill sits around the model: retries, context assembly, human correction, support escalation, and the work of proving the output is acceptable.</p>
<p>Route by value and by risk. Trivial work stays cheap and local. High-stakes work earns expensive inference and stronger checks. A finance-aware leader can answer, without hand-waving:</p>
<ul>
<li>what each class of request costs to serve, end to end</li>
<li>where the rework happens</li>
<li>what failure costs when the model is wrong</li>
<li>which parts of the workflow justify premium inference</li>
</ul>
<p>The cost question nobody owns until it explodes: <strong>when product ships a feature that 10x&rsquo;s tokens, who pays?</strong> If the answer is &ldquo;we&rsquo;ll figure it out,&rdquo; you have not designed an operating model. You have deferred a fight.</p>
<p>Compute placement is part of this calculation, not a separate one. For high-frequency agentic workloads, a chain of round-trips across regions and vendors compounds into real latency tax and real egress cost. Local-first, hardware-aware patterns earn their place where the workload mix justifies them — and create a worse outcome where it does not. Measure first, place compute second.</p>
<p><strong>A cheaper model that fails gracefully beats an expensive model that fails silently.</strong></p>
<h2 id="8-sovereignty-is-an-architecture-constraint">8. Sovereignty is an architecture constraint</h2>
<p> <a href="/blog/2026-04-06-sovereign-systems-privacy-non-optional/"
   
   >Privacy is not a feature you bolt on</a>
 before an enterprise contract closes. It is the shape of the system.</p>
<p>A sovereign system controls the full lifecycle of every piece of data — where it lives, who can access it, how long it persists, and what happens when someone asks you to delete it. In practice, four concrete patterns:</p>
<ul>
<li><strong>Customer-managed keys.</strong> BYOK or hold-your-own-key. If your cloud provider holds the only copy of the encryption key, &ldquo;we cannot access your data&rdquo; is a policy promise, not a verifiable claim.</li>
<li><strong>Regional routing with storage isolation.</strong> EU data does not leave EU infrastructure. The application layer handles the routing. The deployment pipeline ships multi-region.</li>
<li><strong>Scoped, short-lived access.</strong> No ambient credentials. Service-to-service tokens with explicit grants and automatic expiry.</li>
<li><strong>Immutable audit trails.</strong> Append-only, tamper-evident logging of every access, transformation, and movement.</li>
</ul>
<p>&ldquo;We use AWS&rdquo; is not an answer to &ldquo;where does my data live.&rdquo; <strong>Sovereignty is about specificity.</strong></p>
<p>The compounding bill arrives when you try to add this later. The discount arrives when you build it in early and close enterprise contracts without an architectural retrofit.</p>
<h2 id="9-the-threat-model-is-the-manifesto">9. The threat model is the manifesto</h2>
<p>An AI manifesto without a threat model is marketing copy. Four risks every operator names explicitly:</p>
<ul>
<li><strong>Indirect prompt injection.</strong> Instructions hidden in retrieved documents, tool outputs, and user uploads — not just in the user&rsquo;s direct prompt. Treat every retrieved string as potentially adversarial. Validate before it reaches the model. Strip before it reaches the agent.</li>
<li><strong>Silent quality drift.</strong> The model returns <em>slightly</em> worse reasoning. The tone shifts. The retrieval starts ignoring critical documents. There is no stack trace. Only asynchronous production scoring, anchored to a golden set, catches this before customers do.</li>
<li><strong>Vendor and model lock-in by accident.</strong> Fine-tunes, preference data calibrated to one model family, and prompts hand-tuned to a specific tokenizer compound. By year two, your &ldquo;swappable&rdquo; model is a six-month migration. Discipline preserves optionality: prompt abstraction, eval portability, vendor-neutral preference data, and a quarterly review of what would break if the vendor changed terms tomorrow.</li>
<li><strong>Agent blast radius creep.</strong> Permissions accumulate. The agent that summarizes documents quietly gains write access to your billing API because someone needed it once. Audit scope quarterly. Treat agent permissions like database credentials, not like configuration.</li>
</ul>
<p>Threat modeling is not a one-time exercise. It is the bill of materials your system runs on.</p>
<hr>
<h1 id="movement-iv--people--failure">Movement IV — People &amp; Failure</h1>
<h2 id="10-interfaces-beat-titles">10. Interfaces beat titles</h2>
<p>Most AI hiring plans try to fix an interface problem with resumes. They rarely work.</p>
<p>A working leadership system is not a roster of senior titles. It is a decision map. Four owners with explicit decision rights and explicit escalation paths:</p>
<ul>
<li><strong>Product</strong> — user outcomes, adoption, business tradeoffs.</li>
<li><strong>Platform</strong> — safe defaults, deployment paths, observability, paved roads.</li>
<li><strong>Applied AI</strong> — workflow behavior, routing, prompting, retrieval, evaluation quality.</li>
<li><strong>Governance</strong> — risk boundaries, sovereignty controls, escalation thresholds.</li>
</ul>
<p>The titles can be anything. The interfaces cannot be ambiguous. If the answers depend on who is online that day, the system is not operational.</p>
<p>The same logic governs platform teams. A platform exists to make repeated decisions disappear into the default path — identity, routing, eval harnesses, logging, safe deployment, fallback behavior. The moment platform becomes a queue that has to bless every use case,  <a href="/blog/2026-05-14-why-ai-platform-teams-become-bottlenecks/"
   
   >the queue is the product</a>
 and waiting is the cost. <strong>A platform should remove waiting, not become a waiting room.</strong></p>
<p>Hiring works after the operating contract is clear, not before. New hires scale the current operating model, good or bad. <strong>Org debt is interface debt with better branding.</strong></p>
<h2 id="11-anti-fragility-requires-portability-discipline">11. Anti-fragility requires portability discipline</h2>
<p>Resilience is surviving the shock. Anti-fragility is using the shock to remove the next one.</p>
<p>Fragility hides in the org chart and in the stack. One engineer who knows the routing. One vendor whose terms changed last week. One fine-tune that took six months to train and would take six months to migrate. That is not an organization or a system. That is a single point of failure wearing a department badge or a model card.</p>
<p>Four design choices build strength:</p>
<ul>
<li><strong>Modular ownership.</strong> No critical function depends on one person&rsquo;s memory. Deputies are named.</li>
<li><strong>Resettable interfaces.</strong> A model, vendor, or workflow can be swapped without a rewrite. This is not free. It requires prompt abstraction, eval portability, vendor-neutral preference data, and a regular drill where the team actually proves a swap is possible.</li>
<li><strong>Fast learning loops.</strong> Every failure produces a tighter eval, a better fallback, or a clearer operating boundary.</li>
<li><strong>Cross-training on the boring parts.</strong> Alerts, evals, fallback logic, access boundaries. The unglamorous work is what keeps the organization elastic.</li>
</ul>
<p>A short anti-fragility check:</p>
<ul>
<li>Can you swap a model without rewriting the product?</li>
<li>Can you lose a key engineer without losing the system?</li>
<li>Can you absorb a vendor price increase without panic?</li>
<li>Can you turn a production incident into an improved control?</li>
</ul>
<p>If any answer is no, the organization is more brittle than it thinks. The most expensive lie an AI organization tells itself is that the model is swappable when nobody has tried.</p>
<h2 id="12-the-year-two-test">12. The year-two test</h2>
<p>A lot of AI organizations look healthy in month three and brittle by year two. The model did not fail. The operating model did. Prototype energy is easy to create. Durable coordination is not.</p>
<p>The single question that separates the two:</p>
<blockquote>
<p>Can the AI system survive a senior person going on vacation for two weeks?</p>
</blockquote>
<p>If the answer is &ldquo;not really,&rdquo; the organization is still running on hidden tribal knowledge.</p>
<p>If the answer is &ldquo;yes, with documented ownership, a published reliability contract, an eval suite that blocks releases, and a fallback path the on-call engineer can execute at 2 a.m.,&rdquo; the company is moving from prototype to production.</p>
<p>That is the only year-two test that matters. Everything else in this manifesto is in service of passing it.</p>
<hr>
<h2 id="what-this-manifesto-is-not">What this manifesto is not</h2>
<p>It is not a prediction about which model wins. It is not a framework for replacing engineers with agents. It is not a defense of any vendor, any cloud, or any stack.</p>
<p>It is a statement about how serious companies organize for AI when the easy money, the demo budgets, and the hype cycles are done — and only the operating model is left to do the work.</p>
<p>The model will change.</p>
<p>The system around it should not.</p>
<hr>
<p><em>Law Zava writes about the operating model behind serious AI execution. Companion artifacts:  <a href="/docs/agent-reliability-contract"
   
   >Agent Reliability Contract template</a>
 ·  <a href="/docs/rollback-template"
   
   >Rollback document template</a>
 ·  <a href="/docs/eval-starter-kit"
   
   >Eval Suite starter kit</a>
. The canonical reading path is at  <a href="/blog"
   
   >/blog</a>
.</em></p>
]]></content:encoded></item><item><title>The Board Deck Is Lying: How to Measure AI Progress Without Theater</title><link>https://lawzava.com/blog/2026-05-05-measure-ai-progress-without-theater/</link><pubDate>Tue, 05 May 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-05-05-measure-ai-progress-without-theater/</guid><description>Most AI progress reporting confuses activity with value. Executive measurement should collapse around adoption, reliability, margin, and delivery speed.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Most AI dashboards count motion, not progress. They record pilots, prompts, and meetings, then call that momentum. If the scorecard cannot show adoption, reliability, margin, or cycle-time improvement, it is a prop. A board should be able to read it and know whether the business is better off.</p>
<h2 id="the-theater-problem">The Theater Problem</h2>
<p>AI reporting drifts toward  <a href="/blog/2022-10-17-engineering-metrics-that-matter/"
   
   >vanity metrics</a>
 because vanity metrics are easy to collect and hard to argue with.</p>
<p>The usual suspects:</p>
<ul>
<li>number of pilots launched</li>
<li>number of prompts written</li>
<li>number of models tested</li>
<li>number of meetings held</li>
<li>number of slides in the board update</li>
</ul>
<p>None of those is useless on its own. The problem is that none of them answers the only question that matters: <strong>what improved because we shipped this?</strong></p>
<h2 id="a-better-executive-scorecard">A Better Executive Scorecard</h2>
<p>A serious AI scorecard should be small enough to remember and strong enough to force a decision.</p>
<p>Start with four dimensions:</p>
<ol>
<li><strong>Adoption</strong> — are real users using it in a real workflow?</li>
<li><strong>Reliability</strong> — does it fail in bounded, observable ways?</li>
<li><strong>Margin</strong> — does it reduce cost or improve unit economics?</li>
<li><strong>Speed</strong> — does it shorten a real business cycle time?</li>
</ol>
<p>If a project does not move at least one of those numbers, it is not strategic. It is a lab exercise with a budget.</p>
<p>The point is not to build a perfect dashboard. The point is to make it impossible to hide weak outcomes behind busy activity.</p>
<h2 id="what-to-report-weekly">What to Report Weekly</h2>
<p>A weekly AI review should be short, blunt, and decision-oriented.</p>
<p>Report:</p>
<ul>
<li>what shipped</li>
<li>what users actually did with it</li>
<li>what broke</li>
<li>what it cost</li>
<li>what decision changed because of the data</li>
</ul>
<p>That last bullet matters. Progress reporting without decisions is performance art.</p>
<p>A team can launch five experiments in a week and still have no strategy. Strategy shows up when the evidence sharpens the next choice.</p>
<h2 id="keep-the-dashboard-honest">Keep the Dashboard Honest</h2>
<p>There are two reliable ways AI dashboards lie.</p>
<p>First, they drift toward lagging metrics only. By the time the board sees the number, the product problem is already old.</p>
<p>Second, they reward volume instead of signal. A busy roadmap can still be a weak roadmap.</p>
<p>Keep the dashboard honest by requiring every metric on the top page to map to one of  <a href="/blog/2026-04-28-margin-risk-speed-ai-strategy-metrics/"
   
   >three board outcomes</a>
:</p>
<ul>
<li>margin expansion</li>
<li>risk compression</li>
<li>execution-speed advantage</li>
</ul>
<p>If a metric does not help the board understand at least one of those outcomes, it belongs lower in the stack or not at all.</p>
<p>A line worth keeping: <strong>if the scorecard cannot survive finance review, it is not strategy.</strong></p>
<h2 id="key-takeaways">Key Takeaways</h2>
<ul>
<li>Measure adoption, reliability, margin, and speed.</li>
<li>Weekly reviews should force decisions, not decorate slides.</li>
<li>Tie every visible metric to margin, risk, or execution speed.</li>
<li>If the dashboard cannot survive finance review, move it off the first page.</li>
</ul>
]]></content:encoded></item><item><title>The 2026 AI Build vs. Buy Calculus (It’s Just Operational Cost)</title><link>https://lawzava.com/blog/2026-04-30-ai-build-vs-buy/</link><pubDate>Thu, 30 Apr 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-04-30-ai-build-vs-buy/</guid><description>By mid-2026, AI build vs buy has nothing to do with novelty. It is a ruthless mathematical calculation of telemetry, context freshness, and infrastructure lock-in.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>In 2026, build vs. buy is not a taste question. It is an operational cost question. Are you prepared to own the telemetry, the fallback paths, and the failure modes that come with the stack? Buying gives you speed and leaves the analytics with someone else. Building gives you control and hands you the overhead.</p>
<h2 id="the-myth-of-the-headline-price">The Myth of the Headline Price</h2>
<p>Most teams compare API pricing to GPU rental and stop there. That is the wrong first-order model.</p>
<p>Token price is the easiest number to quote and the least useful number to trust. The real bill shows up in the work around the model:</p>
<ul>
<li><strong>Telemetry &amp; Evals:</strong> If you self-host, you must build the pipeline that captures, scores, and reviews output. Vendor APIs may bundle some of this, but then they also own the metadata.</li>
<li><strong>Graceful Degradation:</strong> When the provider throttles you at peak, do you have local fallback? Hybrid systems buy resilience, but they also add systems-engineering work.</li>
<li><strong> <a href="/blog/2026-04-06-sovereign-systems-privacy-non-optional/"
   
   >Data Sovereignty</a>
:</strong> Sometimes the reason to build is simple: the data cannot legally leave your VPC. Once that is true, the token price stops mattering.</li>
</ul>
<h2 id="when-to-buy-the-commodity-highway">When to Buy (The Commodity Highway)</h2>
<p>Buy when the AI capability is a feature, not the product.</p>
<p>If you are building an internal documentation chatbot, a support-ticket summarizer, or a semantic search overlay, buy the API. Do not spend engineering throughput standing up vLLM instances and chasing KV-cache optimizations for a problem that is not your moat.</p>
<p>The catch is lock-in at the integration layer. If your code imports vendor-specific classes directly, you will feel the squeeze when prices change or a model line is deprecated.  <a href="/blog/2024-03-18-multi-model-strategies/"
   
   >Keep the provider behind an internal interface</a>
.</p>
<h2 id="when-to-build-the-crucible-of-control">When to Build (The Crucible of Control)</h2>
<p>Build when AI sits inside unit economics or inside a hard trust boundary.</p>
<p>You must build if:</p>
<ol>
<li>Your margins depend on it. Billions of tokens a day can make  <a href="/blog/2026-03-09-the-end-of-fat-cloud-agentic-economy/"
   
   >the API tax</a>
 the difference between a healthy product and a broken one.</li>
<li>You operate under zero-trust or residency constraints. In healthcare, finance, or defense, the data cannot touch a multi-tenant cloud edge.</li>
<li>You need hardware-level optimization. Sub-150ms tail latency usually means quantization, attention fusion, and serious control over the runtime.</li>
</ol>
<p>That is the part teams underestimate. You are no longer building a prompt pipeline. You are operating a distributed, heavily constrained state machine. That takes engineers who understand memory bandwidth, not just prompting.</p>
<h2 id="the-hybrid-default">The Hybrid Default</h2>
<p>The mature pattern in 2026 is a barbell.</p>
<p>Buy frontier models for complex reasoning, planning, and high-context zero-shot tasks. Build or host  <a href="/blog/2024-08-05-small-models-big-impact/"
   
   >quantized, heavily tuned 8B models</a>
 for the large volume of routing, formatting, and classification work that sits underneath the product.</p>
<p>The CTO&rsquo;s job is not to choose a camp. It is to make the handoff between buy and build a config change, not a rewrite.</p>
]]></content:encoded></item><item><title>Margin, Risk, and Speed: The Three Numbers That Should Drive AI Strategy</title><link>https://lawzava.com/blog/2026-04-28-margin-risk-speed-ai-strategy-metrics/</link><pubDate>Tue, 28 Apr 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-04-28-margin-risk-speed-ai-strategy-metrics/</guid><description>Most AI strategy becomes clearer when leadership stops tracking novelty and starts forcing every decision through three numbers.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Most AI strategy decks are full of nouns and short on numbers. That is usually the tell. If a project cannot move margin, reduce risk, or shorten the path to an outcome, it is not strategy. It is activity with a steering committee.</p>
<h2 id="why-three-numbers-are-enough">Why Three Numbers Are Enough</h2>
<p>Leaders overcomplicate AI strategy because they do not want to choose.</p>
<p>But every AI decision eventually lands in one of three buckets:</p>
<ul>
<li><strong>Margin</strong> — does it improve unit economics?</li>
<li><strong>Risk</strong> — does it make the system safer or more controllable?</li>
<li><strong>Speed</strong> — does it shorten the path from decision to outcome?</li>
</ul>
<p>That is the executive frame. Everything else supports it.</p>
<p>If a project cannot clearly improve at least one of those numbers, it does not belong near the top of the roadmap.</p>
<h2 id="the-trap-of-novelty-metrics">The Trap of Novelty Metrics</h2>
<p>AI teams love the wrong metrics because the wrong metrics are easy to count.</p>
<p>Number of models tested. Number of pilots launched. Number of prompts written. Number of demos shown. Number of meetings held.</p>
<p>Those numbers can tell you whether work is happening. They do not tell you whether the company is getting more profitable, less exposed, or faster to act.</p>
<h2 id="build-a-scorecard-around-outcomes">Build a Scorecard Around Outcomes</h2>
<p>A serious AI scorecard is short.</p>
<ol>
<li>Did margin improve?</li>
<li>Did risk go down?</li>
<li>Did  <a href="/blog/2026-03-30-throughput-engineer-headcount-lagging-metric/"
   
   >cycle time</a>
 shorten?</li>
</ol>
<p>Everything else is instrumentation that helps answer those questions.</p>
<p>That does not mean you ignore adoption, reliability, or cost. It means you use them as inputs to the three executive numbers, not as substitutes for them.</p>
<p>The strongest boards and founders do not need twenty metrics. They need a few numbers that are hard to fake.</p>
<h2 id="make-the-three-numbers-operational">Make the Three Numbers Operational</h2>
<p>The framework only works if the numbers are real.</p>
<p>For each AI initiative, define:</p>
<ul>
<li>the baseline</li>
<li>the target</li>
<li>the measurement cadence</li>
<li>the owner</li>
<li>the rollback path if the numbers move the wrong way</li>
</ul>
<p>That keeps the conversation concrete and makes the project accountable.</p>
<p>A line worth keeping: <strong>if a strategy cannot change one of the three numbers, it is probably theater.</strong></p>
<h2 id="key-takeaways">Key Takeaways</h2>
<ul>
<li> <a href="/blog/2026-04-16-ai-capital-allocation-what-to-stop-funding/"
   
   >Margin, risk, and speed</a>
 are enough to evaluate AI strategy.</li>
<li>Stop reporting novelty metrics as if they were outcomes.</li>
<li>Give every project a baseline, target, owner, cadence, and rollback path.</li>
<li>If the work does not change the numbers, the work is not strategic.</li>
</ul>
]]></content:encoded></item><item><title>AI Capital Allocation: What Great CTOs Stop Funding First</title><link>https://lawzava.com/blog/2026-04-16-ai-capital-allocation-what-to-stop-funding/</link><pubDate>Thu, 16 Apr 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-04-16-ai-capital-allocation-what-to-stop-funding/</guid><description>Strong AI strategy starts with a kill list. If a project cannot defend margin, risk, or speed, it should not survive the next budget meeting.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Great AI teams do not start with a roadmap. They start with a kill list. If a project cannot defend margin, risk, or speed, it does not deserve the next budget cycle. Capital is finite. Attention is finite. Support burden is finite.</p>
<p>The real mistake most companies make is treating AI spend as a separate class of spend. It is not. It competes with product work, platform work, hiring, and operational debt. If you cannot explain why an AI initiative deserves scarce capital, you are not allocating capital. You are subsidizing hope.</p>
<h2 id="capital-allocation-is-the-first-product-decision">Capital Allocation Is the First Product Decision</h2>
<p>Capital allocation is not a finance problem that happens to engineering. It is a technical leadership problem with finance consequences.</p>
<p>Every AI project consumes three things:</p>
<ul>
<li>engineering time</li>
<li>infrastructure budget</li>
<li>organizational attention</li>
</ul>
<p>If the project does not improve one of three board-level outcomes — margin expansion, risk compression, or execution speed — it is likely a vanity project wearing a product costume.</p>
<p>That does not mean the project has to be immediately profitable. It does mean you should be able to state what gets better if the project works and what gets worse if it does not.</p>
<h2 id="what-should-die-first">What Should Die First</h2>
<p>The easiest place to make mistakes is the demo room. The second easiest is the budget meeting.</p>
<p>Stop funding these first:</p>
<ol>
<li>
<p><strong>Thin demos that do not survive workflow reality.</strong>
If the user needs three manual edits after every response, you have built a presentation layer, not a product.</p>
</li>
<li>
<p><strong>Duplicate platform work.</strong>
If two teams are building separate prompt orchestration, evaluation, or routing layers, one of them should stop. Duplication feels like speed until the maintenance bill lands.</p>
</li>
<li>
<p><strong>Ambiguous experiments with no owner.</strong>
“We should explore AI” is not a strategy. It is a permission slip for drift.</p>
</li>
<li>
<p><strong>Projects with no measurable failure mode.</strong>
If nobody can say what counts as bad output, bad latency, bad cost, or bad adoption, the project cannot be managed.</p>
</li>
</ol>
<p>There is a simple reason these projects linger: they are emotionally easy to defend. Nobody wants to kill a project that sounds innovative. But if you cannot defend it with numbers, the project is not innovative. It is unpriced.</p>
<h2 id="the-kill-list-rubric">The Kill-List Rubric</h2>
<p>A good kill list is not a spreadsheet of personal dislikes. It is a decision system.</p>
<p>Before funding a new AI initiative, ask three questions:</p>
<ul>
<li><strong>Does this increase margin, reduce risk, or improve speed?</strong></li>
<li><strong>Can we measure that effect within one quarter?</strong></li>
<li><strong>Do we own the fallback if the model or vendor changes?</strong></li>
</ul>
<p>If the answer to all three is not yes, the default should be no.</p>
<p>This is where a lot of teams get sentimental. They continue funding because the project has a sponsor, or because it already consumed sunk cost, or because it looks good in a board deck. Those are weak reasons to keep a system alive.</p>
<p>Strong reasons to keep funding an AI initiative usually look like this:</p>
<ul>
<li>it replaces high-volume manual work</li>
<li>it improves decision quality in a regulated workflow</li>
<li>it reduces customer wait time</li>
<li>it protects a revenue stream that depends on fast, accurate responses</li>
</ul>
<p>Notice that none of those reasons mention hype.</p>
<h2 id="what-to-keep-funding-instead">What to Keep Funding Instead</h2>
<p>The highest-return AI investments are boring in the best way.</p>
<p>Fund the parts that make the system measurable and durable:</p>
<ul>
<li>retrieval and context quality</li>
<li> <a href="/blog/2024-02-19-evaluating-llm-applications/"
   
   >evaluation harnesses</a>
</li>
<li>fallback logic</li>
<li>routing by task class</li>
<li>observability around bad outputs and retries</li>
<li>workflow-specific data collection</li>
</ul>
<p>The point is not to chase the smartest model. The point is to build a system that can absorb model churn without forcing a rewrite every six months.</p>
<p>A useful line to keep in mind: <strong>if a system cannot be measured under load, it is still a pilot.</strong> Pilots are fine. Pilots just should not keep consuming production budget forever.</p>
<h2 id="the-hard-part-is-saying-no">The Hard Part Is Saying No</h2>
<p>The best operators are not famous for being aggressive spenders. They are famous for being disciplined about what they do not fund.</p>
<p>That discipline becomes a reputation asset. The founder who sees you delete a weak AI project starts trusting your judgment. The board member who sees you cut duplicate work starts trusting your signal. The engineering team that sees you protect their time starts trusting your priorities.</p>
<p>Capital allocation is how you tell the truth about what matters. If a project cannot defend margin, risk, or speed, it should not survive by momentum alone. Fund the systems that make AI measurable, recoverable, and cheap to operate. Cut the rest.</p>
]]></content:encoded></item><item><title>AI Strategy: The CTO Perspective (It's Just Data Infrastructure)</title><link>https://lawzava.com/blog/2026-04-14-ai-cto-perspective/</link><pubDate>Tue, 14 Apr 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-04-14-ai-cto-perspective/</guid><description>A CTO&amp;amp;rsquo;s AI strategy is not about chasing models. It is about resilient data infrastructure, operational boundaries, and measured throughput.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>In 2026, a CTO&rsquo;s AI strategy is not a model shortlist. It is an operating model for data, latency, evaluation, and failure. The model will change. The system around it should not.</p>
<p>If your AI plan still starts with &ldquo;which model should we buy,&rdquo; you are solving the easiest problem in the room. The moat is the pipeline that feeds context, the eval loop that catches regressions, and the fallback path that keeps the product standing when the model misses.</p>
<h2 id="the-strategy-is-the-infrastructure">The Strategy Is the Infrastructure</h2>
<p>The single biggest mistake engineering organizations make is treating the model as the brain. It is not. It is the most expensive dependency in the stack.</p>
<p>The brain is everything you build around it: context assembly, retrieval, validation, retries, telemetry, and rollback.</p>
<p>A CTO must focus ruthlessly on three pillars:</p>
<h3 id="1-the-context-pipeline">1. The Context Pipeline</h3>
<p>The model is only as intelligent as the context you feed it. If Postgres, Cassandra, or Scylla takes five seconds to assemble structured context, encode it, and hand it to the orchestrator, your feature is already late before inference begins.</p>
<p>Strategy means architecting data replication, embedding generation, and caching so the latency budget stays intact for the inference layer. If your data infrastructure is not close to real time, your AI will not be either.</p>
<h3 id="2-the-evaluation-framework">2. The Evaluation Framework</h3>
<p>You cannot scale what you cannot measure. If your organization is still eyeballing model outputs before deployment, you are running a pilot, not a production system.</p>
<p>Leadership means demanding continuous evaluation. Every PR that touches an orchestration layer must be blocked by a CI pipeline that runs 500 deterministic evals against the new reasoning flow. Building that telemetry <em>is</em> the AI strategy.</p>
<h3 id="3-graceful-degradation-and-fallbacks">3. Graceful Degradation and Fallbacks</h3>
<p>LLMs fail. APIs throttle. Endpoints rotate. If a model hallucinates malformed JSON and your core application crashes, that is not an AI failure; that is an architectural failure.</p>
<p>A mature strategy wraps every AI interaction in circuit breakers. If the model fails three times, what is the deterministic fallback? If the cloud provider rate-limits you, where is the  <a href="/blog/2024-08-05-small-models-big-impact/"
   
   >local, quantized 8B-parameter fallback model</a>
 running in your own cluster?</p>
<h2 id="stop-chasing-the-frontier">Stop Chasing the Frontier</h2>
<p>The frontier-model conversation is a distraction. Unless you are OpenAI or Anthropic, you do not win by having the smartest model. You win by having the tightest feedback loop, the cleanest data access, and the lowest cost per transaction.</p>
<p>A strong CTO designs for  <a href="/blog/2024-03-18-multi-model-strategies/"
   
   >swapability</a>
: a single configuration commit, zero downtime, and telemetry that proves the new model performs 4% better on the exact workload that matters.</p>
<p>That is the strategy. Everything else is theater.</p>
]]></content:encoded></item><item><title>The Throughput Engineer: Why Headcount Is a Lagging Metric</title><link>https://lawzava.com/blog/2026-03-30-throughput-engineer-headcount-lagging-metric/</link><pubDate>Mon, 30 Mar 2026 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2026-03-30-throughput-engineer-headcount-lagging-metric/</guid><description>Headcount is a lagging metric. The best engineering organizations measure throughput: decision speed, defect containment, and constraint removal.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Headcount is an input. Throughput is an outcome. The best engineering organizations have stopped asking &ldquo;how many engineers do we need?&rdquo; and started asking &ldquo;what&rsquo;s blocking the engineers we have?&rdquo; Teams that optimize for decision speed, defect containment, and execution clarity outperform teams twice their size. Hiring more people into a broken system just makes the system break faster.</p>
<h2 id="the-metric-everyone-tracks-and-nobody-questions">The Metric Everyone Tracks and Nobody Questions</h2>
<p>Every quarterly planning cycle, the same conversation happens. The roadmap is too ambitious for the team. The proposed solution is more headcount. The exec team approves some fraction of the ask. Six months later, the team is bigger but the roadmap is still slipping.</p>
<p>This pattern persists because headcount is easy to measure and feels actionable. You can put a number on a slide. You can point to it in a board meeting and say &ldquo;we&rsquo;re investing in engineering.&rdquo;</p>
<p>But headcount measures capacity the way adding lanes measures highway throughput. It works up to a point, then coordination overhead offsets the capacity gain. The tenth engineer doesn&rsquo;t add 10% more output. They add 10% more communication paths, 10% more code review load, and another person who needs context on every architectural decision.</p>
<p>The organizations getting this right have shifted to outcome metrics. Not &ldquo;how many people do we have&rdquo; but &ldquo;how fast do decisions move from identification to resolution.&rdquo; Not &ldquo;how many PRs did we merge&rdquo; but &ldquo;what&rsquo;s our  <a href="/blog/2022-01-24-dora-metrics-implementation/"
   
   >change failure rate</a>
 and how quickly do we recover.&rdquo;</p>
<h2 id="staff-growth-versus-constraint-removal">Staff Growth Versus Constraint Removal</h2>
<p>Adding staff is an additive intervention. It puts more resources into the system. Constraint removal is a multiplicative intervention. It makes every existing resource more effective.</p>
<p>Consider a team of eight engineers where the average PR sits in review for 18 hours. Hiring two more engineers does nothing to fix the review bottleneck. It makes it worse because there are now more PRs competing for the same review bandwidth. But changing the review process, setting a 4-hour SLA, pairing reviewers with authors, and shrinking PR scope, can cut that 18 hours to 4 without adding a single person.</p>
<p>The same principle applies at every level. Slow deploys, unclear ownership, meetings that could be async documents, long approval chains. Each costs every engineer on the team hours per week. Multiply by team size and the waste is staggering.</p>
<p>If 20 engineers each lose 5 hours per week to process friction, that&rsquo;s 100 engineer-hours, equivalent to 2.5 full-time engineers doing nothing but waiting. Removing the friction is cheaper than hiring, faster to implement, and doesn&rsquo;t increase coordination costs.</p>
<p>AI tooling has made this dynamic sharper. A well-structured team with good tooling and clear ownership regularly outships teams twice its size. But a poorly structured team with AI tooling just generates more half-finished work faster. AI amplifies the system it operates in, good or bad.</p>
<h2 id="the-operating-system-of-a-high-throughput-team">The Operating System of a High-Throughput Team</h2>
<p>High-throughput teams share three operational patterns that have nothing to do with individual talent.</p>
<p><strong>Clear intent over detailed instructions.</strong> When an engineer picks up a task, they should know the outcome that matters, not the exact steps to get there. &ldquo;Reduce P95 latency on the search endpoint below 200ms&rdquo; is clear intent. &ldquo;Refactor the search query builder to use connection pooling&rdquo; is a solution masquerading as a task. The first lets the engineer use judgment. The second removes it.</p>
<p>Teams that operate on intent move faster because decisions happen at the point of most information, the engineer doing the work, rather than being routed through a manager who has less context. This requires trust, and trust requires that the intent is genuinely clear and that the engineer has the authority to make reasonable tradeoffs.</p>
<p><strong>Delegated authority with explicit boundaries.</strong> Every recurring decision type should have a documented owner and a decision boundary. &ldquo;The on-call engineer can roll back any deploy without approval&rdquo; is a delegation. &ldquo;Database schema changes require review from the data team&rdquo; is a boundary. When these are written down and understood, decisions happen in minutes instead of hours.</p>
<p>The failure mode is implicit authority. Nobody knows who can make the call, so everyone escalates. The escalation chain adds latency to every decision. In a team of 15, this can mean that a simple operational decision takes a day instead of an hour because it bounces between three people who each assume someone else owns it.</p>
<p><strong> <a href="/blog/2020-04-13-async-communication-practices/"
   
   >Async-first communication</a>
.</strong> Synchronous communication, meetings, Slack pings expecting immediate response, tap-on-the-shoulder interruptions, is the most expensive coordination mechanism. It requires everyone to be available simultaneously and context-switch away from focused work.</p>
<p>Async-first doesn&rsquo;t mean no meetings. It means meetings are for decisions that genuinely require real-time discussion. Everything else is a written document, a recorded decision in a ticket, or a code review comment.</p>
<h2 id="a-weekly-operating-cadence">A Weekly Operating Cadence</h2>
<p>Decision tempo separates high-throughput teams from slow ones. A lightweight weekly cadence keeps the system self-correcting without drowning in noise.</p>
<p><strong>Weekly: review leading metrics.</strong> Cycle time from commit to production, change failure rate, time to recover from incidents, review queue depth, and decision latency on open questions. Don&rsquo;t track vanity metrics like lines of code or number of PRs.</p>
<p><strong>Biweekly: connect signals to causes.</strong> Is cycle time creeping up? Is one team&rsquo;s change failure rate spiking? Are the same types of decisions getting stuck repeatedly? The goal is systemic diagnosis, not individual blame.</p>
<p><strong>Biweekly: pick one constraint to remove.</strong> &ldquo;This sprint, we&rsquo;re going to cut our deploy time from 45 minutes to under 10&rdquo; is a decision. &ldquo;We&rsquo;re going to improve developer experience&rdquo; is not. One thing, not five.</p>
<p><strong>Continuous: execute, measure, repeat.</strong> Act on the decision, measure the result, and feed it back into the next weekly review. If cutting deploy time didn&rsquo;t improve cycle time, the constraint was elsewhere. Move to the next one.</p>
<h2 id="incentives-that-reward-impact-over-activity">Incentives That Reward Impact Over Activity</h2>
<p>Most engineering organizations accidentally incentivize busyness. The engineer who closes the most tickets gets praised. The team that ships the most features gets the biggest headcount allocation. The manager who runs the most meetings looks the most engaged.</p>
<p>Throughput-oriented incentives look different.</p>
<p>Reward engineers who eliminate recurring work, not just complete it. The engineer who automates away a manual process that costs the team 10 hours per week has created more value than the engineer who ships a new feature used by 50 people.</p>
<p>Reward teams that improve their own throughput metrics, not just output volume. A team that cuts its change failure rate from 15% to 3% has freed up enormous capacity that was previously spent on rollbacks, hotfixes, and incident response. That&rsquo;s worth more than two new features.</p>
<p>Reward leaders who make themselves less necessary. The manager whose team operates smoothly when they&rsquo;re on vacation has built a better system than the manager who&rsquo;s cc&rsquo;d on every decision.</p>
<h2 id="a-12-week-operating-reset">A 12-Week Operating Reset</h2>
<p>For teams experiencing delivery drag, a structured reset works better than a reorg.</p>
<p><strong>Weeks 1-3: Measure.</strong> Instrument cycle time, change failure rate, review latency, and decision latency. Don&rsquo;t change anything yet. Establish a baseline that everyone agrees on.</p>
<p><strong>Weeks 4-6: Remove one constraint.</strong> Pick the biggest bottleneck revealed by the data. If review latency is the worst, fix the review process. If deploy time is the worst, fix the pipeline. One constraint at a time.</p>
<p><strong>Weeks 7-9: Delegate and document.</strong> Write down the top 10 recurring decision types and who owns each one. Set decision boundaries. Remove one layer of approval from the most common workflow.</p>
<p><strong>Weeks 10-12: Sustain.</strong> Establish the weekly review cadence. Compare throughput metrics to the week-1 baseline. Identify the next constraint. Make the cycle self-reinforcing.</p>
<p>Teams that complete this reset typically see 30-50% improvement in cycle time without adding staff. The improvement comes from removing friction that was invisible because everyone had adapted to it.</p>
<h2 id="board-facing-metrics-that-map-engineering-to-business-risk">Board-Facing Metrics That Map Engineering to Business Risk</h2>
<p>Boards understand risk and return. Translate engineering throughput into those terms.</p>
<p><strong>Cycle time</strong> maps to market responsiveness. &ldquo;We can respond to a competitor move in days, not months&rdquo; is a strategic capability that boards care about.</p>
<p><strong>Change failure rate</strong> maps to operational risk. &ldquo;5% of our changes cause incidents&rdquo; is a risk number a board can evaluate, especially when paired with the cost of those incidents.</p>
<p><strong>Recovery time</strong> maps to resilience. &ldquo;When something breaks, we fix it in under an hour&rdquo; is a durability statement that affects customer trust and revenue protection.</p>
<p><strong>Decision latency</strong> maps to organizational agility. &ldquo;Strategic decisions take 2 days to reach execution, not 2 weeks&rdquo; tells the board that the organization can adapt.</p>
<p>None of these metrics mention headcount. That&rsquo;s the point. Headcount funds capacity. These metrics measure whether that capacity produces results.</p>
<h2 id="key-takeaways">Key Takeaways</h2>
<p>Headcount tells you what you&rsquo;re spending. Throughput metrics, cycle time, change failure rate, recovery time, decision latency, tell you what you&rsquo;re getting.</p>
<p>The highest-leverage engineering work is constraint removal, not feature addition. Every hour of friction you eliminate pays dividends across every engineer on the team.</p>
<p>Stop asking &ldquo;how many engineers do we need?&rdquo; Start asking &ldquo;what&rsquo;s preventing the engineers we have from shipping?&rdquo;</p>
]]></content:encoded></item><item><title>Your AI Metrics Are Measuring the Wrong Thing</title><link>https://lawzava.com/blog/2025-07-07-ai-product-metrics/</link><pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2025-07-07-ai-product-metrics/</guid><description>Engagement metrics tell you people clicked. They tell you nothing about whether your AI feature actually helped anyone do anything.</description><content:encoded><![CDATA[<p>Every AI product review I sit in starts the same way: someone pulls up a dashboard showing adoption rates, interaction volume, and session length. The numbers are up and to the right. Everyone nods.</p>
<p>Then I ask: &ldquo;How many of those interactions ended with the user getting the right answer?&rdquo; Silence.</p>
<p>This is the metrics gap that keeps burning teams. Usage tells you people showed up. It tells you nothing about whether they left with what they needed. An AI feature can be heavily used and actively harmful at the same time. Users try it, get a wrong answer, correct it manually, and keep coming back because they&rsquo;re optimistic. Your dashboard shows engagement. Your product is eroding trust.</p>
<h2 id="what-to-actually-measure">What to Actually Measure</h2>
<p>Three things. That&rsquo;s it.</p>
<p><strong>Did the output help?</strong> Not &ldquo;was it generated.&rdquo; Did it contribute to the user completing their task? Define what successful completion looks like for your specific workflow, then measure whether AI-assisted completions happen more often, faster, or with fewer errors than the baseline. If you can&rsquo;t tie AI output to a task outcome, you&rsquo;re measuring wind.</p>
<p><strong>Was it correct?</strong> Combine  <a href="/blog/2024-02-19-evaluating-llm-applications/"
   
   >automated checks</a>
 with periodic human review. Automated checks catch format violations, hallucinated entities, and  <a href="/blog/2024-11-11-ai-safety-production/"
   
   >safety issues</a>
. Human review catches the subtle stuff: answers that are technically correct but misleading, or correct for the wrong version. Sample 5% of outputs weekly. That&rsquo;s enough to spot trends before they become incidents.</p>
<p><strong>Do users trust it?</strong> Trust is the leading indicator everyone ignores. Track it through implicit signals: how often users edit AI output before accepting it, how often they abandon a flow after seeing the AI response, and how often they re-prompt with the same question phrased differently. Rising edit rates or re-prompt rates mean trust is declining. By the time CSAT surveys catch this, you&rsquo;ve already lost months.</p>
<h2 id="the-dashboard-that-fits-on-one-screen">The Dashboard That Fits on One Screen</h2>
<p>Your AI scorecard should answer four questions at a glance:</p>
<ol>
<li>Are people using it? (adoption, retention &ndash; the basics)</li>
<li>Is the output good? (correctness rate, safety rate from automated + human review)</li>
<li>Is it helping? (task completion rate, time to completion vs. baseline)</li>
<li>Do they trust it? (edit rate, re-prompt rate, abandonment rate)</li>
</ol>
<p>Review weekly. Tie every metric to a decision. If a number moves and nobody changes anything, delete the number. Dashboards without decisions are theater.</p>
<p>When a metric dips, you should be able to trace it back to a model update, a retrieval change, or a product shift within the same week. If you can&rsquo;t, your  <a href="/blog/2025-03-31-ai-observability-deep/"
   
   >instrumentation</a>
 is too coarse.</p>
<h2 id="the-uncomfortable-truth">The Uncomfortable Truth</h2>
<p>Most teams avoid quality metrics because they&rsquo;re harder to collect and the numbers are less flattering than engagement counts. That&rsquo;s exactly why they matter. The teams that measure task success and trust alongside usage are the ones whose AI features survive past the demo phase.</p>
<p>Measure what the user felt. Everything else is vanity.</p>
]]></content:encoded></item><item><title>AI in 2025: The Year Discipline Wins</title><link>https://lawzava.com/blog/2025-01-06-ai-trends-2025/</link><pubDate>Mon, 06 Jan 2025 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2025-01-06-ai-trends-2025/</guid><description>The AI hype cycle is over. 2025 is about the teams who can make this stuff actually work in production &amp;amp;ndash; repeatably, measurably, and without burning money.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Stop chasing model announcements. The teams that win in 2025 are the ones building evals, monitoring quality, and treating AI like infrastructure instead of magic. Discipline over heroics.</p>
<hr>
<p>Every January, someone publishes a breathless AI predictions post. &ldquo;This will be the year of AGI.&rdquo; &ldquo;Agents will replace developers.&rdquo; &ldquo;Multimodal everything.&rdquo;</p>
<p>I&rsquo;m not going to do that.</p>
<p>What I can tell you is what I see working with teams that are actually shipping AI to production. The pattern is clear: 2024 was the year everyone built demos. 2025 is the year those demos have to work.</p>
<h2 id="the-demo-hangover">The demo hangover</h2>
<p>Here&rsquo;s what happened to most AI projects last year. Someone built a prototype in a weekend. It was impressive. Leadership got excited. Budget appeared. Then the prototype hit real users, real data, and real edge cases, and everything got complicated.</p>
<p>I watched this play out at three different companies. Same story every time. The model was fine. The engineering around the model wasn&rsquo;t.</p>
<p>Missing evaluation suites. No fallback paths. Prompts that drifted every time someone tweaked them. Cost tracking that amounted to &ldquo;we&rsquo;ll figure it out later.&rdquo; The model was the easy part. Operating discipline was the hard part.</p>
<p>That&rsquo;s the real trend for 2025. Not a new model. A new level of engineering rigor around models.</p>
<h2 id="reasoning-gets-interesting">Reasoning gets interesting</h2>
<p>Models that think before they answer are genuinely useful for a specific class of problems. Multi-step analysis. Code review. Debugging. Anything where you would rather wait 30 seconds for a correct answer than get a fast wrong one.</p>
<p>The trap is treating reasoning models as the default. They&rsquo;re slower, more expensive, and overkill for 80% of requests. The smart move is routing: fast model for simple tasks, reasoning model for complex ones. I&rsquo;ll write more about this in a couple of weeks.</p>
<h2 id="multimodal-is-real-but-boring">Multimodal is real but boring</h2>
<p>Image, audio, and text working together is no longer a research demo. It&rsquo;s a feature. Internal tools are the clearest win &ndash; think document-processing pipelines that can read scanned forms, or support systems that understand screenshots.</p>
<p>The value isn&rsquo;t in any single modality being amazing. It&rsquo;s in combining them so the system has richer context. Boring. Useful. Exactly the kind of thing that makes money.</p>
<h2 id="evaluation-first-development">Evaluation-first development</h2>
<p>The single biggest shift I keep pushing is simple:  <a href="/blog/2024-02-19-evaluating-llm-applications/"
   
   >define success before you write the first prompt</a>
.</p>
<p>This sounds obvious. Almost nobody does it. Teams will spend weeks tuning prompts and then measure success by vibes. &ldquo;It feels better.&rdquo; &ldquo;The CEO liked the demo.&rdquo; That isn&rsquo;t engineering. That&rsquo;s hope.</p>
<p>What works: a fixed eval set, tested on every change, with clear pass/fail criteria. Treat prompts like code. Version them. Review them. Test them. I won&rsquo;t ship a prompt change without running it against the eval suite. Period.</p>
<h2 id="governance-stops-being-optional">Governance stops being optional</h2>
<p>Regulation is firming up. The EU AI Act is real. Enterprise clients are asking for audit trails, documentation, and risk tiers before they&rsquo;ll sign contracts. If your AI system can&rsquo;t explain what it does, what data it touches, and who&rsquo;s responsible when it goes wrong, you&rsquo;re in for a bad year.</p>
<p>This isn&rsquo;t bureaucracy for its own sake. Good governance actually accelerates adoption because it turns &ldquo;can we use AI for this?&rdquo; from a six-week debate into a checklist. Risk tier low? Ship it. Risk tier high? Here&rsquo;s exactly what you need before you ship.</p>
<p>Governance that blocks delivery is broken governance. Governance that makes yes safe and fast is a competitive advantage.</p>
<h2 id="agents-promising-overhyped">Agents: promising, overhyped</h2>
<p>Agents that can execute multi-step tasks are improving fast. They&rsquo;re also still brittle. Context changes break them. Domain boundaries confuse them. The failure modes are subtle and hard to detect.</p>
<p>The near-term play is constrained agents with explicit checkpoints. Not open-ended autonomy. Not &ldquo;let the agent figure it out.&rdquo; Clear scope, clear permissions, clear rollback. We learned this lesson with microservices a decade ago: autonomy without contracts is chaos.</p>
<h2 id="what-im-ignoring">What I&rsquo;m ignoring</h2>
<ul>
<li>Any roadmap built on vendor keynote slides instead of product outcomes.</li>
<li>Prompt engineering tricks that can&rsquo;t be tested, versioned, or reproduced.</li>
<li>&ldquo;Autonomous&rdquo; systems with no permission model, no audit trail, and no kill switch.</li>
<li>Anyone who says &ldquo;just add AI&rdquo; without specifying what success looks like.</li>
</ul>
<h2 id="what-matters">What matters</h2>
<p>The capabilities are real. The models will keep getting better. But the gap between &ldquo;this works in a demo&rdquo; and &ldquo;this works in production at 3am on a Saturday&rdquo; is where careers and companies are made.</p>
<p>Ruthless focus on the boring stuff. Evals. Monitoring. Cost tracking. Fallback paths. Governance. That&rsquo;s the 2025 playbook.</p>
<p>The teams that  <a href="/blog/2024-12-09-ai-infrastructure-scale/"
   
   >treat AI like infrastructure</a>
 &ndash; with the same rigor they bring to databases and deployment pipelines &ndash; will win. Everyone else will keep rebuilding demos.</p>
]]></content:encoded></item><item><title>2025 Will Reward the Boring Teams</title><link>https://lawzava.com/blog/2024-12-23-preparing-for-2025/</link><pubDate>Mon, 23 Dec 2024 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2024-12-23-preparing-for-2025/</guid><description>The AI advantage in 2025 goes to teams that ship measurable workflows, not teams that chase capabilities. The gap is discipline, not technology.</description><content:encoded><![CDATA[<p>The prediction game is easy. Models get better. Context windows get longer. Multimodal improves. Agents get more capable. Legal and compliance teams get more involved. None of this is surprising.</p>
<p>The harder question: what should you actually do differently?</p>
<p>Here’s my short answer, based on a year of working on AI across multiple organizations and watching the gap between teams that shipped and teams that stalled.</p>
<h2 id="stop-experimenting-start-measuring">Stop Experimenting. Start Measuring.</h2>
<p>If you&rsquo;ve been running AI &ldquo;experiments&rdquo; for more than a quarter without a clear evaluation framework, you aren&rsquo;t experimenting. You&rsquo;re procrastinating. Experiments have hypotheses, metrics, and endpoints. Pilots have owners, success criteria, and deadlines.</p>
<p>Pick two or three use cases closest to production. Define success in numbers, not narratives. Build an evaluation set. Ship to real users with monitoring. Learn from data, not opinions.</p>
<p>This isn&rsquo;t glamorous. It&rsquo;s effective.</p>
<h2 id="build-the-operational-foundation">Build the Operational Foundation</h2>
<p>The teams that will move fastest in 2025 are the ones building the plumbing now. Not new models. Not new frameworks. Plumbing.</p>
<ul>
<li>An evaluation loop that runs regularly, not when someone remembers</li>
<li>Cost tracking with per-feature attribution so you know where money goes</li>
<li>Security controls for model access and data handling that satisfy your legal team</li>
<li>Model-agnostic interfaces so you can swap providers without rewriting your stack</li>
</ul>
<p>Every one of these is boring. Every one of these is a prerequisite for scaling anything in 2025. Through Q4, I&rsquo;ve been helping teams set up exactly this kind of infrastructure, and the teams that have it in place are already iterating faster than teams that built flashy demos without it.</p>
<h2 id="governance-isnt-the-enemy">Governance Isn&rsquo;t the Enemy</h2>
<p>AI governance has a reputation problem. Engineers hear &ldquo;governance&rdquo; and think &ldquo;bureaucracy that slows us down.&rdquo; That framing is wrong.</p>
<p>Lightweight governance &ndash; clear ownership for use case intake, a simple review path for legal and security risks, a cadence for measuring value and retiring weak experiments &ndash; actually accelerates shipping. It removes the ambiguity that causes teams to stall waiting for implicit approval.</p>
<p>The companies that move fastest all have some version of this. Not a committee. Not a 50-page policy document. A clear owner, a simple process, and a regular review. That&rsquo;s it.</p>
<h2 id="what-im-betting-on">What I&rsquo;m Betting On</h2>
<p>Personally, I&rsquo;m betting that 2025 is the year AI stops being a separate initiative and becomes part of how software gets built. Not a team. Not a project. A capability that lives inside existing workflows, owned by existing teams, measured by existing standards.</p>
<p>The companies that treat AI as special will keep producing expensive demos. The companies that treat it as normal &ndash; same code review, same evaluation, same cost accountability, same ownership &ndash; will ship things that last.</p>
<p>Discipline over heroics. Same as always.</p>
]]></content:encoded></item><item><title>Why Your Enterprise AI Pilot Is Stuck</title><link>https://lawzava.com/blog/2024-06-03-enterprise-ai-adoption/</link><pubDate>Mon, 03 Jun 2024 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2024-06-03-enterprise-ai-adoption/</guid><description>Most enterprise AI projects die between the demo and production. The blockers aren&amp;amp;rsquo;t technical &amp;amp;ndash; they&amp;amp;rsquo;re organizational. Here&amp;amp;rsquo;s what I keep seeing.</description><content:encoded><![CDATA[<p>Every enterprise AI conversation I&rsquo;ve had this year follows the same arc. Someone builds a proof of concept. The demo goes well. Leadership gets excited. Then, three months later, the project is stuck in limbo: security reviews, data access requests, and nobody quite sure who actually owns it.</p>
<p>I see this pattern across telecom and fintech organizations. The demo-to-production gap isn&rsquo;t a technology problem. It&rsquo;s an organizational one.</p>
<h2 id="the-demo-was-the-easy-part">The demo was the easy part</h2>
<p>A POC can skip everything that makes enterprise software hard. It runs on a developer&rsquo;s laptop with test data. It doesn&rsquo;t need to handle real user volumes. During a demo, nobody asks about audit trails or data retention policies.</p>
<p>Then the project moves toward production and reality hits. Security wants a threat model. Legal wants to know where the data goes. The platform team wants to know who pays for compute. The data science team discovers the training data is messier than expected. None of this is surprising. These are the same problems every enterprise system faces, plus a few new AI-specific ones: model drift, prompt management, and probabilistic outputs.</p>
<p>The teams that get stuck are the ones that treated the POC as the starting line instead of a feasibility check.</p>
<h2 id="start-boring-stay-boring">Start boring, stay boring</h2>
<p>The single best predictor of success I&rsquo;ve seen is picking a first use case that&rsquo;s low-risk and internal. Something where a human reviews the output before anything happens. Document summarization for internal teams. Draft generation for support responses that get edited before sending. Classification of inbound requests to route them to the right queue.</p>
<p>These aren&rsquo;t exciting. That&rsquo;s the point. You want a use case where a bad output is an inconvenience, not a liability. One where you can iterate on prompts and evaluate quality without a customer ever seeing an unpolished result.</p>
<p>I keep telling teams the same thing: your first AI feature should be invisible to customers. Ship it internally, prove it works, build the muscle memory for operating AI in production, then expand.</p>
<h2 id="build-the-platform-before-the-pilots-multiply">Build the platform before the pilots multiply</h2>
<p>Here&rsquo;s what happens when you don&rsquo;t have a shared platform: every team builds its own integration. They pick different models, prompt patterns, and logging approaches. Six months later, you have eight AI features and no way to compare quality, manage costs, or enforce policies across them.</p>
<p>The fix is unglamorous. Build a thin shared layer early. It needs three things:</p>
<ol>
<li><strong>Centralized model access</strong> with authentication, rate limiting, and cost tracking.</li>
<li><strong>A prompt registry</strong> so prompts are versioned, reviewable, and not buried in application code.</li>
<li><strong>Evaluation tooling</strong> that every team can use to measure output quality against a golden set.</li>
</ol>
<p>This doesn&rsquo;t need to be perfect or fully featured. It needs to exist before the third team starts building their own AI integration. I&rsquo;ve watched organizations try to consolidate after the fact. It&rsquo;s painful and expensive.</p>
<h2 id="governance-that-enables-instead-of-blocks">Governance that enables instead of blocks</h2>
<p>The worst governance models I see are designed by committee without input from the engineering teams that have to live with them. They produce a 40-page policy document, a six-week review cycle, and a strong incentive for teams to quietly build things without telling anyone.</p>
<p>Good governance is lightweight and fast. A one-page use case template. A clear risk-tier system: low risk gets self-service approval, high risk gets review. A standing meeting where legal, security, and engineering are in the same room instead of a months-long email chain.</p>
<p>One organization I worked with reduced its AI approval cycle from eight weeks to five days by switching from a document-based review to a 30-minute live walkthrough with all stakeholders. Same rigor. Fraction of the time.</p>
<h2 id="the-uncomfortable-truth">The uncomfortable truth</h2>
<p>Most enterprise AI projects don&rsquo;t fail because the technology isn&rsquo;t ready. They fail because the organization isn&rsquo;t ready. The AI works fine in the demo. The procurement process takes four months. The data team can&rsquo;t provide clean training data. The legal review has no precedent to follow, so it defaults to &ldquo;no&rdquo; until someone escalates.</p>
<p>If you want to ship AI in an enterprise, spend less time evaluating models and more time clearing organizational roadblocks. Get a budget owner. Get a security sponsor. Get data access sorted before you write the first prompt.</p>
<p>Process beats talent. Every time.</p>
]]></content:encoded></item><item><title>Stop Starting With the Model: AI Product Strategy That Works</title><link>https://lawzava.com/blog/2023-09-04-ai-product-strategy/</link><pubDate>Mon, 04 Sep 2023 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2023-09-04-ai-product-strategy/</guid><description>Every roadmap I&amp;amp;rsquo;ve seen this quarter has an AI feature. Most of them start with the wrong question. Start with the user problem, not the model.</description><content:encoded><![CDATA[<p>Every product meeting I&rsquo;ve been in for the past three months starts the same way: &ldquo;How can we use AI for this?&rdquo; Wrong question. The right question is: &ldquo;What problem are we solving, and is AI the best tool for it?&rdquo;</p>
<p>I know, I know. That sounds like something a consultant says to justify a strategy offsite. But having built AI features at a fintech company and watched a dozen teams try to do the same, the pattern is clear: teams that start with the model build demos. Teams that start with the problem build products.</p>
<h2 id="the-boring-problem-statement-trick">The boring problem statement trick</h2>
<p>Here&rsquo;s a filter I&rsquo;ve started using. Take your AI feature idea and describe it without mentioning AI, ML, LLMs, or GPT. If what&rsquo;s left sounds like a useful product improvement, you probably have something. If what&rsquo;s left sounds empty, you&rsquo;re building a solution looking for a problem.</p>
<p>&ldquo;We use GPT-4 to generate personalized financial insights&rdquo; becomes &ldquo;We show users relevant patterns in their spending.&rdquo; The second version is a product. The first version is a technology demo. The second version also tells you what to measure: are the patterns actually relevant? Do users act on them?</p>
<p>At a fintech company, we almost fell into this trap. The initial pitch was &ldquo;AI-powered transaction intelligence.&rdquo; That means nothing. We reframed it as &ldquo;reduce the time finance teams spend manually categorizing transactions from 4 hours to 20 minutes.&rdquo; Now we had a measurable goal, a clear user, and a way to know if it worked.</p>
<h2 id="where-ai-actually-helps">Where AI actually helps</h2>
<p>AI is good at tedious, pattern-heavy work where approximate answers are acceptable. Classification, summarization, drafting, extraction. It saves humans from the work they hate doing and are bad at doing consistently.</p>
<p>AI is bad at precision, accountability, and reasoning about edge cases. It&rsquo;s bad at knowing what it doesn&rsquo;t know. It&rsquo;s bad at anything where &ldquo;usually right&rdquo; is worse than &ldquo;always following a rule.&rdquo;</p>
<p>The sweet spot for AI features in most products right now: draft-and-review workflows. The AI produces a first pass. The human reviews, corrects if needed, and approves. Both sides do what they&rsquo;re good at.</p>
<p>The failure pattern: giving the model autonomy over decisions that matter. Auto-categorizing low-value transactions? Fine. Auto-approving expense reports over $10K? No. The line is about consequences. If the model gets it wrong, how bad is it? If the answer is &ldquo;not great,&rdquo; add a human checkpoint.</p>
<h2 id="the-five-questions-that-kill-bad-ideas-early">The five questions that kill bad ideas early</h2>
<p>Before a team starts building an AI feature, I make them answer these. In writing. No hand-waving.</p>
<ol>
<li><strong>What user problem is this solving?</strong> Not &ldquo;what could we do with AI&rdquo; but &ldquo;what is painful today, and who feels that pain?&rdquo;</li>
<li><strong>What does a wrong answer cost?</strong> If incorrect output has serious consequences, you need guardrails and fallbacks before you need features.</li>
<li><strong>How do we measure success?</strong> Not usage. Not engagement. What concrete outcome improves? Time saved, errors reduced, revenue impacted.</li>
<li><strong>What happens when the model is unavailable or uncertain?</strong> Your feature needs to work without AI. If it can&rsquo;t, it&rsquo;s too tightly coupled to a service you don&rsquo;t control.</li>
<li><strong>What&rsquo;s our edge over someone else calling the same API?</strong> If the answer is &ldquo;our prompts,&rdquo; you don&rsquo;t have an edge. Proprietary data, workflow integration, and feedback loops create defensibility. Prompt engineering doesn&rsquo;t.</li>
</ol>
<p>If any answer is weak, run a time-boxed experiment instead of a full build. Two weeks, clear success criteria, then decide.</p>
<h2 id="the-moat-isnt-the-model">The moat isn&rsquo;t the model</h2>
<p>I wrote about this in July, but it bears repeating because I keep seeing it: access to an API isn&rsquo;t a competitive advantage. OpenAI&rsquo;s API is available to everyone. So is Anthropic&rsquo;s. So are the open-source models.</p>
<p>Your moat is the system around the model. The data you collect from user interactions. The workflow integration that makes ripping you out painful. The quality metrics that let you improve faster than competitors. The fallback system that means your product works even when the model doesn&rsquo;t.</p>
<p>At a startup accelerator, I watched a batch of companies try to build moats around technology that was commoditizing in real time. The ones that survived had distribution or data advantages. The ones that died had clever technology and nothing else. Same pattern applies here.</p>
<h2 id="ship-the-ugly-version">Ship the ugly version</h2>
<p>The final piece of advice is the least popular: ship something simple and ugly before you ship something impressive and polished. A feature that works reliably for 80% of cases with a clear fallback for the other 20% beats a feature that works brilliantly for 95% of cases and fails catastrophically for the other 5%.</p>
<p>Users forgive limitations when the product is honest about them. They don&rsquo;t forgive confidence that turns out to be wrong.</p>
<p>Build the product. Not the demo.</p>
]]></content:encoded></item><item><title>What I Learned Building AI Features Into a Fintech Product</title><link>https://lawzava.com/blog/2023-08-07-building-ai-features/</link><pubDate>Mon, 07 Aug 2023 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2023-08-07-building-ai-features/</guid><description>Building AI features at a fintech taught me the hard part isn&amp;amp;rsquo;t the model: it&amp;amp;rsquo;s defining quality, handling failures, and not shipping a demo as a product.</description><content:encoded><![CDATA[<p>Three weeks ago, we shipped an AI-powered transaction categorization feature for a fintech infrastructure company. The demo took two days to build. Getting it production-ready took six weeks. That ratio tells you everything about AI feature development.</p>
<p>The demo was impressive. You paste in a batch of transactions, the model categorizes them, and the output looks clean. The CEO loved it. The PM loved it. I loved it too, right up until we started testing edge cases.</p>
<p>A wire transfer labeled &ldquo;SEPA CT REF-8847291 ACME GMBH&rdquo; landed in &ldquo;Entertainment.&rdquo; A recurring subscription payment got labeled differently every time we ran it. And the model confidently categorized a clearly fraudulent transaction as &ldquo;Regular business expense&rdquo; without any hesitation.</p>
<p>This is the gap between demo and product, and it&rsquo;s where most AI features die.</p>
<h2 id="the-thing-nobody-talks-about-defining-good">The thing nobody talks about: defining &ldquo;good&rdquo;</h2>
<p>Before writing any production code, I made the team answer one question: what does a correct categorization look like, and what happens when it&rsquo;s wrong?</p>
<p>For traditional features, this is obvious. A button either works or it doesn&rsquo;t. A calculation is either right or wrong. For AI features, &ldquo;right&rdquo; is fuzzy. Is 85% accuracy good enough? Depends. If you&rsquo;re categorizing expense reports for internal review, maybe. If you&rsquo;re categorizing transactions for regulatory reporting, absolutely not.</p>
<p>We built an eval set of 200 transactions with hand-labeled categories. Not exciting work. Took the team three days. But that eval set became the foundation for every decision that followed: prompt changes, model selection, fallback logic, launch criteria.</p>
<p>The rule I enforce now: if you can&rsquo;t write down what &ldquo;good&rdquo; looks like in concrete examples before you start building, you&rsquo;re not ready to build.</p>
<h2 id="architecture-for-uncertainty">Architecture for uncertainty</h2>
<p>AI features sit inside the same product architecture as everything else. But they need extra layers to manage the fundamental uncertainty of probabilistic output. Here&rsquo;s what the production feature stack actually looked like:</p>
<p><strong>Input validation.</strong> Transactions go through a normalizer before the model sees them. Strip reference codes, standardize currency formats, expand abbreviations. The cleaner the input, the more consistent the output. This sounds boring. It improved accuracy by 8 points.</p>
<p><strong>The model call.</strong> GPT-3.5-turbo with a tight prompt, structured JSON output, and a confidence score. We tried GPT-4 initially: better accuracy, but 10x the cost and 3x the latency. For this use case, 3.5-turbo plus good input normalization was the better trade-off.</p>
<p><strong>Output validation.</strong> Every categorization gets checked against the valid category list. If the model returns a category that doesn&rsquo;t exist (and it does, about 2% of the time), we fall back. If the confidence score is below our threshold, we fall back.</p>
<p><strong>The fallback.</strong> This is the part most teams skip. Our fallback is a rules-based categorizer that handles ~40% of transactions using keyword matching and counterparty lookup. It&rsquo;s not as good as the model, but it&rsquo;s deterministic and always available. When the model is uncertain, the user gets the rules-based result with a flag saying &ldquo;review suggested.&rdquo;</p>
<p><strong>Feedback loop.</strong> Users can correct categorizations. Those corrections feed back into our eval set and, eventually, into prompt improvements. This is the part that compounds over time.</p>
<h2 id="testing-probabilistic-systems">Testing probabilistic systems</h2>
<p>Unit tests cover the deterministic parts: input normalization, output validation, the rules-based fallback. Those work exactly like traditional testing.</p>
<p>Model behavior gets tested differently. We run the full eval set (200 transactions) on every prompt change and weekly against the live model. The metrics we track:</p>
<ul>
<li>Accuracy against our labeled set (target: &gt;88%)</li>
<li>Category distribution (catches when the model starts favoring certain categories)</li>
<li>Confidence score distribution (catches when the model becomes less certain overall)</li>
<li>Fallback rate (how often we&rsquo;re bypassing the model)</li>
</ul>
<p>We don&rsquo;t assert on individual outputs. That&rsquo;s a trap. The same input might get slightly different wording each time, and that&rsquo;s fine as long as the category is right. We assert on aggregate metrics across the eval set.</p>
<p>One thing that bit us: OpenAI changed something in the model (not a version bump, just inference-time behavior) and our accuracy dropped 3 points overnight. We caught it because we run the eval set daily. Teams that don&rsquo;t do this are flying blind.</p>
<h2 id="launching-without-embarrassment">Launching without embarrassment</h2>
<p>We launched to 5% of users first. Internal users, specifically the finance team. They&rsquo;re the harshest critics and the most forgiving audience: they understand the constraints and give precise feedback.</p>
<p>Two things came out of the soft launch that we hadn&rsquo;t anticipated:</p>
<ol>
<li>
<p>Multi-currency transactions confused the model more than we expected. A USD payment from a EUR account got categorized based on the currency, not the merchant. We added currency normalization to the input pipeline.</p>
</li>
<li>
<p>Users didn&rsquo;t trust the output even when it was correct. Adding the confidence score as a visual indicator (&ldquo;high confidence&rdquo; / &ldquo;review suggested&rdquo;) dramatically improved trust. People are fine with AI output when they know the system is honest about its uncertainty.</p>
</li>
</ol>
<p>After two weeks with the finance team, we expanded to 50%, then 100%. The feedback loop was running, the eval metrics were stable, and the fallback was handling edge cases gracefully.</p>
<h2 id="what-id-tell-another-team">What I&rsquo;d tell another team</h2>
<p><strong>Define quality before you write code.</strong> Build the eval set first. It&rsquo;s tedious, and it&rsquo;s the most important thing you&rsquo;ll do.</p>
<p><strong>Design the fallback before the happy path.</strong> What happens when the model is wrong or unavailable? If that experience is terrible, your feature is fragile. If it&rsquo;s graceful, you can ship confidently.</p>
<p><strong>Instrument everything.</strong> Not just errors and latency. Quality metrics, confidence distributions, fallback rates, user corrections. You need to see the model&rsquo;s behavior as a continuous signal, not a binary pass/fail.</p>
<p><strong>Resist the demo.</strong> The demo is a trap. It shows the best case. Production shows the average case and every edge case you didn&rsquo;t think of. Ship the demo only after you&rsquo;ve built the scaffolding that makes the average case acceptable and the worst case survivable.</p>
<p>AI features are software. Treat them that way, with tests, rollout controls, monitoring, and healthy skepticism. The model is the least interesting part of the system. The interesting part is everything around it.</p>
]]></content:encoded></item><item><title>Most AI Startups Are Wrappers. That's the Problem.</title><link>https://lawzava.com/blog/2023-07-03-ai-startup-landscape/</link><pubDate>Mon, 03 Jul 2023 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2023-07-03-ai-startup-landscape/</guid><description>Everyone has an AI startup now. Having been through two accelerators and founded two companies, I can tell you: most of these will not survive the year.</description><content:encoded><![CDATA[<p>Every pitch deck I&rsquo;ve seen in the last three months has an AI slide. Every single one. I sat through a batch of startup pitches last month &ndash; founder-program alumni network, so the bar is usually decent &ndash; and eight out of ten were some variation of &ldquo;we put GPT-4 in front of [industry] data.&rdquo;</p>
<p>I&rsquo;ve founded two startups. I&rsquo;ve been through a deep-tech founder program and a startup accelerator. I know what a defensible business looks like and what a science fair project looks like. Most of what&rsquo;s being built right now is the latter.</p>
<h2 id="the-wrapper-problem">The wrapper problem</h2>
<p>Here&rsquo;s the test I apply: if OpenAI adds your feature to ChatGPT next Tuesday, do you still have a company? If the answer requires more than three seconds of thought, you&rsquo;re a wrapper.</p>
<p>Wrappers are easy to build, impressive in demos, and worthless when the platform moves. I saw this exact pattern in mobile (2012-2013), chatbots (2016), and crypto (2021). The playbook is always the same: new technology enables easy prototypes, hundreds of startups launch, the platform or an incumbent absorbs the feature, and the startups die.</p>
<p>The AI version is faster because the prototyping is faster. You can build a convincing demo in a weekend. That&rsquo;s the trap.</p>
<h2 id="where-the-actual-value-is">Where the actual value is</h2>
<p>The startups that will survive have at least one of these:</p>
<p><strong>Proprietary data loops.</strong> If every user interaction makes the product better in a way competitors can&rsquo;t replicate, you have something. This is boring and slow to build. Good.</p>
<p><strong>Deep workflow integration.</strong> If ripping you out requires a migration project, switching costs protect you. This means being embedded in existing processes, not sitting as a standalone tool people can ignore.</p>
<p><strong>Domain expertise that reduces error rates.</strong> In regulated industries &ndash; finance, healthcare, legal &ndash; the model being 90% right is a liability, not a feature. The value is in the last 10%, and that requires domain knowledge the model doesn&rsquo;t have.</p>
<p><strong>Distribution.</strong> If you already have the customers and you&rsquo;re adding AI to an existing product, you win against the startup that&rsquo;s trying to acquire customers and build AI simultaneously.</p>
<p>Notice what&rsquo;s not on the list: &ldquo;better prompts.&rdquo; Prompt engineering isn&rsquo;t a moat. It&rsquo;s barely a speed bump.</p>
<h2 id="what-id-actually-build">What I&rsquo;d actually build</h2>
<p>If I were starting a company today &ndash; and I think about this more than I should &ndash; I&rsquo;d focus on the infrastructure layer. The tooling for evaluation, observability, and cost management. The picks-and-shovels play during a gold rush is a cliche for a reason: it actually works.</p>
<p>Or I&rsquo;d go deep vertical in a domain I know. Fintech, specifically. (I&rsquo;m biased &ndash; I&rsquo;ve spent years in the space.) The financial services industry has massive data complexity, real compliance requirements, and budgets. An AI product that can navigate GAAP, regulatory reporting, and multi-currency reconciliation isn&rsquo;t getting replaced by ChatGPT anytime soon.</p>
<h2 id="the-honest-take">The honest take</h2>
<p>Most AI startups being funded right now won&rsquo;t exist in 18 months. Not because AI isn&rsquo;t real &ndash; it absolutely is. But because having access to an API isn&rsquo;t a business. The companies that survive will be the ones that did something hard with the technology, not something easy.</p>
<p>If your pitch starts with &ldquo;we use GPT-4 to&hellip;&rdquo; you&rsquo;ve already lost. Start with the problem. The model is a detail.</p>
]]></content:encoded></item><item><title>How We Track and Prioritize Tech Debt at a Fintech Startup</title><link>https://lawzava.com/blog/2018-12-10-tech-debt-tracking-and-prioritization/</link><pubDate>Mon, 10 Dec 2018 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2018-12-10-tech-debt-tracking-and-prioritization/</guid><description>A framework for cataloging technical debt, scoring it by impact and risk, and scheduling paydown without stalling feature delivery.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Stop pretending tech debt will fix itself. Put it in a registry, score it, and schedule it like real work. We did this at the fintech startup and it turned a chaotic backlog of &ldquo;we should really fix that&rdquo; into something we actually ship against every sprint.</p>
<hr>
<p>Every engineering team I&rsquo;ve worked with has some version of the same conversation. Someone mentions a gnarly part of the codebase. Everyone nods. Someone says &ldquo;we should really clean that up.&rdquo; Nobody does.</p>
<p>At the fintech startup, we hit a point where this was genuinely hurting us. Our fintech data pipeline had accumulated enough shortcuts and half-finished migrations that feature work was getting slower every quarter. Not dramatically. Just enough friction that estimates kept creeping up and nobody could point to exactly why.</p>
<p>So I built a system for it. Nothing fancy. But it changed how we think about debt.</p>
<h2 id="what-tech-debt-actually-is">What tech debt actually is</h2>
<p>Tech debt is a future cost you created with a present decision. Sometimes that decision was smart &ndash; you shipped faster and the tradeoff was worth it. Sometimes it was accidental &ndash; the design was fine until requirements shifted. Either way, the cost is real.</p>
<p>The forms we see most often:</p>
<ul>
<li><strong>Deliberate shortcuts.</strong> You knew it was a hack. You shipped anyway. Fair enough.</li>
<li><strong>Accidental debt.</strong> Looked fine at the time. Scale or new requirements proved otherwise.</li>
<li><strong>Environmental shifts.</strong> A dependency gets deprecated. A compliance rule changes. Not your fault, still your problem.</li>
<li><strong>Operational gaps.</strong> Missing monitoring, thin tests, no runbook. The kind of thing that bites you at 2am.</li>
</ul>
<h2 id="our-debt-registry">Our debt registry</h2>
<p>Here is what actually worked for us. We created a debt registry &ndash; basically a shared list of known liabilities that lives right in our issue tracker. Not a separate doc. Not a wiki page nobody reads. Same board, same sprint planning, same visibility as feature work.</p>
<p>Each entry looks roughly like this:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-yaml" data-lang="yaml"><span style="display:flex;"><span><span style="color:#f92672">id</span>: <span style="color:#ae81ff">TD-021</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">title</span>: <span style="color:#ae81ff">Legacy auth flow lacks rate limiting</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">area</span>: <span style="color:#ae81ff">auth-service</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">impact</span>: <span style="color:#ae81ff">Security, reliability</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">risk</span>: <span style="color:#ae81ff">High</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">effort</span>: <span style="color:#ae81ff">3-4</span> <span style="color:#ae81ff">weeks</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">owner</span>: <span style="color:#ae81ff">Platform</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">status</span>: <span style="color:#ae81ff">Proposed</span>
</span></span></code></pre></div><p>The key insight was treating debt entries as first-class work items. When debt lives in a separate spreadsheet, it gets ignored. When it sits next to feature tickets in the same planning session, it gets discussed.</p>
<p>We review the registry every two weeks during sprint planning. Takes ten minutes. That alone changed how we worked.</p>
<h2 id="scoring-keep-it-dead-simple">Scoring: keep it dead simple</h2>
<p>We tried complex scoring matrices. They didn&rsquo;t survive contact with reality. What stuck was three dimensions and a formula you can do in your head:</p>
<ul>
<li><strong>Impact.</strong> How much does this slow us down or degrade the product?</li>
<li><strong>Risk.</strong> What happens if it gets worse or fails outright?</li>
<li><strong>Effort.</strong> How much work to fix it?</li>
</ul>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-text" data-lang="text"><span style="display:flex;"><span>priority = (impact * 2) + risk - effort
</span></span></code></pre></div><p>Each on a 1-5 scale. Is it perfect? No. But it forces you to compare things that otherwise feel impossible to rank. &ldquo;The auth system is scary&rdquo; versus &ldquo;the build is slow&rdquo; suddenly becomes a conversation with numbers instead of gut feelings.</p>
<h2 id="when-to-act-now-versus-later">When to act now versus later</h2>
<p>Some things can&rsquo;t wait:</p>
<ul>
<li>Security exposure. Full stop.</li>
<li>The same root cause behind repeated incidents.</li>
<li>A dependency about to hit end-of-life.</li>
<li>Workarounds that tax every single release.</li>
</ul>
<p>Everything else goes into the prioritized backlog. High impact, low effort items get picked up opportunistically. Large refactors get scheduled as dedicated work with clear milestones.</p>
<h2 id="how-we-schedule-paydown">How we schedule paydown</h2>
<p>We tried three approaches before settling on a hybrid:</p>
<p><strong>Capacity allocation.</strong> We reserve roughly 20% of each sprint for maintenance and debt work. Non-negotiable. Product knows about it. This is the baseline.</p>
<p><strong>Debt sprints.</strong> Every six weeks, we run a focused sprint on the highest-priority debt items. Engineers pick from the top of the registry. These sprints have been some of the most satisfying work the team does.</p>
<p><strong>Opportunistic paydown.</strong> If you&rsquo;re already in the code, improve it. Boy scout rule. We just ask people to tag the cleanup commits so we can track the effort.</p>
<h2 id="incremental-migration-over-big-rewrites">Incremental migration over big rewrites</h2>
<p>We learned this the hard way. Big-bang rewrites fail. They just do. Every large debt item at the fintech startup now follows the same pattern:</p>
<ul>
<li>Run old and new paths in parallel.</li>
<li>Migrate the highest-traffic paths first.</li>
<li>Set a cutover date. Remove old code on that date. No exceptions.</li>
</ul>
<p>Feature flags make this safe:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#66d9ef">if</span> flags<span style="color:#f92672">.</span>new_auth_flow:
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> new_auth<span style="color:#f92672">.</span>authenticate(user)
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">return</span> legacy_auth<span style="color:#f92672">.</span>authenticate(user)
</span></span></code></pre></div><p>The flag gives you a kill switch. Sleep better at night.</p>
<h2 id="selling-debt-work-to-non-engineers">Selling debt work to non-engineers</h2>
<p>This is where most teams fail. You can&rsquo;t walk into a planning meeting and say &ldquo;the auth code is messy.&rdquo; Nobody cares.</p>
<p>What works: translate debt into business language. &ldquo;Feature X takes twice as long to ship because every change requires manual regression testing in the auth module.&rdquo; Now it&rsquo;s a delivery speed conversation. Product managers understand delivery speed.</p>
<p>At the fintech startup, we started including debt impact in our sprint velocity reports. When the team could show that velocity dropped 15% quarter-over-quarter and tie it to specific debt items, getting time allocated stopped being a fight.</p>
<h2 id="prevention">Prevention</h2>
<p>A few guardrails go further than any cleanup sprint:</p>
<ul>
<li>Lightweight design reviews for anything touching critical paths.</li>
<li>Code review checklists that ask about maintainability, not just correctness.</li>
<li>Quarterly dependency upgrades so you never face a multi-year jump.</li>
<li>Documentation expectations for anything another team will touch.</li>
</ul>
<h2 id="closing">Closing</h2>
<p>Tech debt isn&rsquo;t a failure. It&rsquo;s an accounting problem. You took a loan, now you need a repayment plan.</p>
<p>The registry changed everything for us. Not because it was sophisticated &ndash; it was a handful of fields in Jira. But because it made invisible costs visible. And once everyone can see the cost, the conversation about paying it down gets a lot easier.</p>
]]></content:encoded></item><item><title>Stop Trying to Fix All Your Tech Debt</title><link>https://lawzava.com/blog/2017-12-18-technical-debt-triage-framework-for-prioritization/</link><pubDate>Mon, 18 Dec 2017 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2017-12-18-technical-debt-triage-framework-for-prioritization/</guid><description>A two-number scoring system for tech debt that tells you what to fix now, what to schedule, and what to quietly accept.</description><content:encoded><![CDATA[<h2 id="quick-take">Quick take</h2>
<p>Multiply pain by how often you touch it. Fix the top. Schedule the middle. Ignore the rest guilt-free.</p>
<h3 id="the-problem-with-we-should-really-fix-that">The Problem With &ldquo;We Should Really Fix That&rdquo;</h3>
<p>At the fintech startup we had a spreadsheet. Forty-seven items on it. Stuff ranging from &ldquo;our deploy script is held together with duct tape&rdquo; to &ldquo;that one table column named <code>data2</code>.&rdquo; Every retro, someone added more. Nothing came off.</p>
<p>The list was useless because everything on it felt important to whoever wrote it down. No ranking. No shared language for severity. Just a growing monument to good intentions.</p>
<p>So I built a framework. Dead simple. Two numbers.</p>
<h3 id="two-numbers-one-multiplication">Two Numbers, One Multiplication</h3>
<p>For every piece of tech debt, score two things on a 1-to-5 scale:</p>
<p><strong>Impact</strong> &ndash; how much does this actually hurt? A 5 means outages, data bugs, or features you literally can&rsquo;t ship. A 1 means it annoys you when you see it but has zero customer-facing consequence.</p>
<p><strong>Change frequency</strong> &ndash; how often does someone touch this code? A 5 means daily. A 1 means you forgot the directory existed.</p>
<p>Multiply them. That&rsquo;s your priority score.</p>
<p><code>priority = impact x change frequency</code></p>
<ul>
<li><strong>15-25</strong>: Fix this soon. Like, put it in the next sprint soon.</li>
<li><strong>8-14</strong>: Schedule it. Quarter planning, debt sprints, whatever your cadence is.</li>
<li><strong>1-7</strong>: Accept it. Seriously. Move on.</li>
</ul>
<h3 id="running-the-numbers-on-real-stuff">Running the Numbers on Real Stuff</h3>
<p>I&rsquo;ll use examples close to what we actually triaged at the fintech startup.</p>
<p>Our user service had no tests and we were shipping changes to it three times a week. Impact 4, frequency 5. Score: <strong>20</strong>. Obviously top of the list.</p>
<p>The deploy process was painful and slow. Impact 4, frequency 4. Score: <strong>16</strong>. Right behind it.</p>
<p>We had a gnarly payment integration that scared everyone, but we only touched it once a quarter. Impact 5, frequency 2. Score: <strong>10</strong>. Scary, but not urgent.</p>
<p>An admin dashboard that loaded slowly? Impact 2, frequency 2. Score: <strong>4</strong>. Nobody cared enough. And that was the correct call.</p>
<p>The math removes the emotion. That&rsquo;s the point.</p>
<h3 id="not-all-debt-ages-the-same">Not All Debt Ages the Same</h3>
<p>This matters and most people miss it. Some debt compounds. Every time you work around the problem, the workaround becomes the new baseline. The next person works around the workaround. You know exactly what I&rsquo;m talking about.</p>
<p>Other debt just sits there. Ugly, stable, inert. The weird naming convention in a module nobody touches? It&rsquo;ll be weird next year too, but it won&rsquo;t be worse.</p>
<p>And some debt evaporates. We had a messy integration with a third-party API we were planning to drop. Spending time cleaning it up would have been pure waste.</p>
<p>When you&rsquo;re prioritizing, ask: is this getting worse? Compounding debt should jump the queue even if the current score is middling.</p>
<h3 id="four-ways-to-actually-pay-it-down">Four Ways to Actually Pay It Down</h3>
<p><strong>Boy Scout rule.</strong> You&rsquo;re already in the file. Leave it slightly better. Rename that variable. Extract that function. This handles the small stuff and keeps entropy from winning.</p>
<p><strong>Protected time.</strong> Block real capacity for debt work. We did one day a week. Some teams do a cooldown sprint after each release. Doesn&rsquo;t matter what rhythm you pick, what matters is that the time is sacred and not the first thing cut when a deadline looms.</p>
<p><strong>Bundle it with features.</strong> If a feature touches a debt-heavy area, pad the estimate to include cleanup. Product managers accept this more easily than standalone debt tickets because the work is tied to something they already want.</p>
<p><strong>Bite the bullet.</strong> Some debt needs a dedicated rewrite. A subsystem replacement. A migration. These need a business case, an owner, and a timeline. Not a Jira ticket that says &ldquo;refactor payments&rdquo; with no assignee sitting in the backlog for eight months.</p>
<h3 id="selling-debt-work-upward">Selling Debt Work Upward</h3>
<p>Engineers talk about debt in terms of code quality. Leadership doesn&rsquo;t care about code quality. They care about shipping speed, incident frequency, and hiring retention.</p>
<p>So translate. &ldquo;This area has no tests and we ship to it constantly, which is why we&rsquo;ve had three production incidents this quarter&rdquo; lands differently than &ldquo;we need to improve test coverage.&rdquo; Same problem. Different framing.</p>
<p>At the fintech startup I started tying debt items to incident reports. When leadership could see a direct line between a piece of debt and a customer-facing problem, the prioritization conversation got a lot shorter.</p>
<h3 id="accepting-debt-is-a-decision-not-a-failure">Accepting Debt Is a Decision, Not a Failure</h3>
<p>Here&rsquo;s what changed once we had the scoring system: we stopped feeling guilty about the bottom of the list. A score of 4 means you&rsquo;ve looked at it, evaluated it, and decided your time is better spent elsewhere. That&rsquo;s not neglect. That&rsquo;s judgment.</p>
<p>Low-score debt in a system you&rsquo;re replacing? Ignore it. Cosmetic issues in stable code? Ignore them. Theoretical problems with no evidence of real pain? Keep an eye on them, but don&rsquo;t burn a sprint.</p>
<h3 id="the-takeaway">The Takeaway</h3>
<p>Forty-seven items became six that mattered. The rest we either accepted or scheduled for later with clear triggers for when to revisit. The team stopped arguing about what to fix because the math settled it.</p>
<p>Two numbers. One multiplication. That&rsquo;s the whole framework.</p>
]]></content:encoded></item><item><title>Pitching Infrastructure to People Who Don't Care About Infrastructure</title><link>https://lawzava.com/blog/2017-09-04-the-business-case-for-infrastructure-investment/</link><pubDate>Mon, 04 Sep 2017 00:00:00 +0000</pubDate><guid>https://lawzava.com/blog/2017-09-04-the-business-case-for-infrastructure-investment/</guid><description>Your board doesn&amp;amp;rsquo;t care about Kubernetes. They care about money, risk, and speed. Here&amp;amp;rsquo;s how I learned to pitch infra investment at the fintech startup.</description><content:encoded><![CDATA[<p>Nobody on your board woke up thinking about database migrations. They woke up thinking about revenue, burn rate, and whether the next funding round will close. If you walk into that room talking about latency percentiles, you&rsquo;ve already lost.</p>
<p>I learned this the hard way at the fintech startup. We needed a serious infrastructure overhaul. I knew it. My team knew it. But our investors didn&rsquo;t speak in uptime and throughput. They spoke in money. So I had to learn their language.</p>
<h3 id="talk-money-or-dont-talk">Talk money or don&rsquo;t talk</h3>
<p>Every infra pitch boils down to one of four things: it makes money, saves money, reduces risk, or lets us ship faster. That&rsquo;s it. Pick the frame that fits and lead with it. Not with the technology. Not with the architecture diagram. With the business outcome.</p>
<p>At the fintech startup, we were heading into a period where marketing had big traffic projections. Our system couldn&rsquo;t handle it. I didn&rsquo;t pitch &ldquo;we need to re-architect our data pipeline.&rdquo; I pitched &ldquo;we&rsquo;ll lose paying users in Q3 if we don&rsquo;t act now.&rdquo; Different sentence. Same project. Completely different reaction in the room.</p>
<h3 id="risk-is-the-easiest-sell">Risk is the easiest sell</h3>
<p>When downside is obvious, risk framing writes itself. Quantify the probability of an outage. Estimate the cost. Multiply. That number gets attention fast because nobody wants to be the one who ignored a preventable disaster.</p>
<p>Growth constraints work too, but you need credible projections. Vague &ldquo;we might need to scale&rdquo; doesn&rsquo;t cut it. Show the wall you&rsquo;re about to hit and the revenue sitting on the other side.</p>
<p>Velocity framing is underrated. If your deploys take two hours and engineers sit idle waiting, that&rsquo;s burned salary. Translate it. &ldquo;We&rsquo;re losing 40 engineer-hours a week to deployment bottlenecks&rdquo; hits harder than &ldquo;our CI/CD pipeline needs work.&rdquo;</p>
<h3 id="do-the-math-even-if-its-rough">Do the math, even if it&rsquo;s rough</h3>
<p>A proposal without numbers is just an opinion. I&rsquo;ve seen smart engineers lose budget requests because they couldn&rsquo;t answer &ldquo;what does this cost us today?&rdquo; Document the current state in dollars and hours. Then describe what changes, in the same units.</p>
<p>Keep the ROI dead simple. Investment. Expected annual benefit. Payback period. If your assumptions are honest and explicit, rough math beats no math every single time.</p>
<h3 id="structure-for-skimmers">Structure for skimmers</h3>
<p>Executives don&rsquo;t read proposals. They skim them. One paragraph up top: what you want, why, and what it&rsquo;s worth. Then a short problem statement in business terms. Proposed solution with minimal technical detail. Quantified impact. Cost. Timeline. Risks and mitigations. Alternatives you considered and rejected.</p>
<p>That order matters. Bury the ask on page three and it never gets read.</p>
<h3 id="where-ive-seen-pitches-die">Where I&rsquo;ve seen pitches die</h3>
<p>Being too technical. Your CFO doesn&rsquo;t need to understand container orchestration. They need to understand what it changes for the business.</p>
<p>Skipping quantification. &ldquo;Better reliability&rdquo; means nothing. &ldquo;$200K in avoided incident costs per year&rdquo; means something.</p>
<p>Ignoring tradeoffs. If this delays a feature, say so. Explain why the trade is still worth it. Pretending there&rsquo;s no cost makes you look naive.</p>
<p>Asking for everything at once. Phase it. Get a win. Use that win to fund the next phase.</p>
<p>And the killer: not following up. If you don&rsquo;t report results after getting funded, you&rsquo;ve just burned your credibility for the next ask.</p>
<h3 id="keep-the-conversation-going">Keep the conversation going</h3>
<p>Don&rsquo;t wait for budget season. Share metrics regularly. When an incident happens, quantify the damage. When delivery speeds up after an investment, show the numbers. Make infrastructure value visible all the time, so the next request feels like an obvious yes instead of a surprise ask.</p>
<p>Infrastructure is a business investment. Treat it like one and it gets funded. Treat it like a technical hobby and it doesn&rsquo;t.</p>
]]></content:encoded></item></channel></rss>