Writing / 2019
Designing for Failure: Timeouts, Circuit Breakers, Failover
Timeouts, blast-radius isolation, fallbacks, circuit breakers, and tested failover: the rules I follow so one slow dependency can't take down everything.
I’m halfway through my founder-program cohort, building my infrastructure startup, and I keep having the same conversation with other founders here: “We’ll handle reliability later.” Later. The word that has personally cost me more sleep than any production bug.
At the fintech startup, I watched a single slow Elasticsearch query cascade through our entire API layer. One degraded dependency. Total platform outage. The fix took ten minutes. The recovery took four hours. All because nothing in the request path had a timeout.
At a mobility startup, a forgotten WAL retention setting filled a disk at 3 AM and I discovered our “tested” failover was eleven hours behind. Forty minutes of locked vehicles across the city. The monitoring said everything was fine. The monitoring was wrong.
These weren’t exotic failures. They were boring, preventable ones. The kind that happen when you assume dependencies work and never verify what happens when they don’t.
Three rules I follow
Set a deadline on everything. Every outbound call gets a timeout. Every request gets a budget. If a dependency can’t answer in time, you move on without it. Slow failure is worse than fast failure because it holds resources hostage while it dies.
Isolate the blast radius. A slow search index should never starve your payment flow. Separate connection pools. Separate queues. The goal is simple: one problem stays one problem.
Know your fallback before you need it. A stale cache hit is better than a 500. A default list of popular items is better than a blank page. But the fallback has to be intentional. Accidental fallbacks are just bugs you haven’t noticed yet.
The pattern that keeps saving me
Circuit breakers. Dead simple concept, and one of the lessons about failure I keep relearning. If a dependency is failing, stop calling it. Serve the fallback. Check back later. It turns a cascading outage into a graceful degradation that most users never notice.
An open breaker is the system working as designed. It chose fast, predictable behavior over slow, unpredictable death.
What I got wrong early on
I used to think resilience meant more redundancy. Add a replica. Add a region. Add a retry. But redundancy without testing is just a more expensive single point of failure. That mobility startup’s replica was a perfect example. It existed. It was running. It was useless.
Now I test the recovery path, not just the happy path. If you haven’t promoted your replica under realistic conditions in the last quarter, you don’t have a failover. You have a hope. Breaking things on purpose is how you find out which one you have.
Reliability is a prioritization problem
Designing for failure is mostly a prioritization problem. Every founder and every CTO knows they should do it. Most don’t because the next feature feels more urgent. It always feels more urgent.
Until 3 AM on a Thursday, when it doesn’t.