Writing / 2019

Zero-Downtime Deployments: Migrations, Probes, and Habits

Zero-downtime deploys depend less on tooling than on expand-and-contract migrations, backward-compatible code, readiness probes, and graceful shutdown.

Two production outages I caused still stick with me. One was at the fintech startup, when a database migration locked a table for eleven minutes during market open. The other was during my founder-program cohort, when I pushed a config change that broke backward compatibility with in-flight requests. Both times, the deploy tooling worked perfectly. The problem was me.

No deploy tool gives you zero downtime on its own. It’s a discipline across code, database migrations, and infrastructure, and the hardest part is the schema changes, not the YAML.

What Zero Downtime Means

Two things. That’s it.

  1. Capacity never drops below what traffic needs.
  2. Running code stays compatible with in-flight requests and existing data.

If both hold, your deploy is invisible to users. If either breaks, you have an outage regardless of how fancy your canary setup is.

Database Migrations Are the Hard Part

Rolling updates, blue-green, canary: pick your favorite. They all handle the application layer fine. The hard part is always the database.

A schema change that locks a table or drops a column still in use will take you down no matter what deployment pattern you’re running. I learned this the painful way with Postgres at the fintech startup, where our users table had enough rows that a naive ALTER TABLE would hold a lock for minutes.

The fix is boring, and I covered it in more depth in database migrations without downtime . Expand, migrate, contract. In that order, across separate deploys.

-- Deploy 1: Expand (add the new column)
ALTER TABLE users ADD COLUMN phone VARCHAR(20);

-- Deploy 2: Migrate (backfill in batches by id range, not one giant UPDATE)
UPDATE users SET phone = legacy_phone
WHERE id > 0 AND id <= 1000 AND phone IS NULL;
-- ...repeat for the next id range until done

-- Deploy 3: Contract (drop old column after all code uses the new one)
ALTER TABLE users DROP COLUMN legacy_phone;

Three deploys for one column rename. Annoying? Yes. But nobody notices. That’s the point.

On MySQL, tools like gh-ost or pt-online-schema-change do background table copies to avoid long locks. On Postgres, the work is mostly careful DDL: nullable columns without defaults, CREATE INDEX CONCURRENTLY, a short lock_timeout, and backfills in batches (Postgres UPDATE has no LIMIT, so batch by primary key range). Worth learning if you’re running anything with real traffic.

Pick a Rollout Pattern and Keep It Simple

Rolling Updates

The default and usually the right choice. Set maxUnavailable: 0 so you never lose capacity during the rollout.

spec:
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0
      maxSurge: 25%

This works when old and new versions can run side by side. Which means your APIs need to be backward compatible. If that sounds like extra work, it is. But it’s the kind of extra work that prevents 3 a.m. pages.

Canary

Send a small slice of traffic to the new version. Watch your error rate. If it’s clean, ramp up. If it spikes, roll back automatically.

The key word there is “automatically.” A canary release without automated rollback is just a smaller blast radius with manual intervention. You might as well flip a coin.

Blue-Green

Run two environments, switch traffic when the new one is verified. Fast rollback, but you’re paying for double infrastructure. Fine for critical services. Overkill for most things.

The Stuff Between the Deploys

The patterns above are table stakes. What actually makes deploys boring (in a good way) is the plumbing around them.

Readiness vs. liveness probes. These aren’t the same thing. Liveness says “is this process alive.” Readiness says “can this process handle traffic right now.” If your readiness check passes before caches are warm and dependencies are verified, you’re sending users to a half-alive instance.

Graceful shutdown. When Kubernetes sends SIGTERM (see my Kubernetes production checklist ), your app needs to stop accepting new connections, finish in-flight requests, then exit. Sounds obvious. Roughly half the Go services I’ve reviewed get this wrong, usually by not waiting long enough for the load balancer to drain.

Connection draining. Set a realistic termination grace period. If your longest request takes 30 seconds, a 5-second grace period is going to drop connections.

DNS TTLs. Lower them before a cutover, raise them after. I’ve seen teams spend hours debugging a “failed deploy” that was actually stale DNS.

My Pre-Deploy Checklist

The questions I ask before every deploy:

  • Can the new code read data written by the old code? Can the old code read data written by the new code?
  • Does the database migration need a lock? How long?
  • Do the readiness checks actually verify dependencies, or do they just return 200?
  • If this deploy goes wrong, can I roll back in under a minute?

If any answer is “I’m not sure,” that’s where the work is. Not in the deploy tool.

Why this stays hard

Zero downtime deploys aren’t hard technically. They’re hard culturally. They require everyone on the team to think about backward compatibility, schema evolution, and graceful degradation before they write the code. Not after.

The teams I’ve seen do this well, at the fintech startup, at startups in my founder-program cohort, and on other teams I’ve worked with, all share one trait. Deploys are boring. Nobody watches them. Nobody holds their breath. The monitoring catches problems and the rollback is automatic.

References