// Topics / Reliability
Reliability
40 entries tagged “Reliability”
- The Review Queue Is Your Real Agent Limit
· 4 min
Plan agent rollouts like capacity: risk-weighted review demand against effective reviewer-hours. Past the constraint, seats buy inventory, not throughput.
ai
operations
teams
A Backup Model You Never Exercise Is Not a Backup
· 4 min
Vendor count is inventory; measured failover time is resilience. Shadow replay, a bounded canary, and a recovery clock with defined start and stop.
ai
reliability
strategy
Autonomy Without a Demotion Path Is Permission Creep
· 4 min
Teams can describe how agents earn autonomy. Few can take it back. Write the transition contract: trigger, authority, mechanics, re-entry.
ai
governance
reliability
The Eval Inversion
· 5 min
Two frontier labs lost containment this summer, in eval environments. The more adversarial the workload, the weaker the walls around it. Invert that.
ai
security
reliability
Garbage Context, Confident Answer
· 4 min
Most AI failures are context failures wearing a model's face. Score retrieval with groundedness and margin, check freshness against live state.
ai
reliability
architecture
AI Insurance Will Ask for Evidence, Not Intent
· 4 min
As insurers exclude AI, protection tracks evidence, not intent. Your operating cadence is your audit trail.
governance
ai
executive
Reliability Is the Autonomy Ceiling
· 4 min
How much autonomy to give an AI agent: bound its failure rate with the rule of three, shadow mode, and fault injection, then price it by failure cost.
ai
reliability
operations
Agent Identity Is the New Control Plane
· 4 min
AI agents need workload identity, not shared API keys: SPIFFE SVIDs, RFC 8693 token exchange, and Vault leases make access scoped, attributable, revocable.
ai
reliability
security
The Benchmark You Didn't Build
· 4 min
Use public LLM benchmarks as a shortlist filter, then decide on an owned eval: programmatic assertions, a versioned LLM judge, and paired per-case diffs.
ai
reliability
metrics
Agentic Systems at Scale: The New Reliability Contract
· 4 min
An AI agent reliability contract is real only where the control plane enforces it: scoped credentials, a deny-by-default tool gateway, a sandbox.
ai
reliability
operations
The Anti-Fragile AI Organization
· 4 min
A warm LLM fallback is a standing cost that needs an owner. Price portability per feature, and turn each vendor shock into a capability you keep.
teams
ai
reliability
Designing the AI Leadership Bench: Roles, Interfaces, and Failure Boundaries
Canon post —
· 2 min
Scaling AI needs a leadership bench: named owners for product, platform, applied AI, and governance, with failure handoffs rehearsed before incidents.
leadership
teams
ai
How to Run an AI Incident Review That Changes Architecture, Not Slides
· 2 min
An AI incident review is done when it changes architecture, evals, alerting, or ownership. An eight-part template that ends in fixes, owners, and dates.
reliability
ai
governance
AI Evaluation and Production Governance: A Maturity Model
· 4 min
A five-level maturity model for AI evaluation and production governance, from vibes-based deploys to CI eval gates, production sampling, and rollback.
governance
ai
reliability
Why Most Enterprise AI Architecture Fails in Year One
· 3 min
Enterprise AI architecture fails in year one when teams expect deterministic behavior from a statistical engine. Build failure boundaries and telemetry.
architecture
ai
reliability
Red-Teaming Distributed Databases Before the Black Swan
· 7 min
Most catastrophic distributed database incidents are compound failures nobody practiced. How to red-team partitions, clock skew, and operator error.
distributed-systems
databases
reliability
Building Reliable AI Agents in Go
· 5 min
How I build reliable AI agents in Go: bounded tools, schema validation at the boundary, idempotent state, and a supervisor loop with hard limits.
agents
reliability
ai
AI Incident Response: Failures That Don't Look Like Outages
· 4 min
AI systems can return 200 OK while confidently wrong. How to detect, contain, and learn from AI incidents using proven incident response principles.
incident-management
ai
reliability
Agentic Workflows in Production: Constrain the Blast Radius
· 5 min
AI agents that take actions carry real blast radius. Policy allowlists, structured workflows, idempotent steps, tracing, and a shadow-mode rollout.
agents
ai
production
Multi-Model LLM Routing and Fallbacks in Production
· 4 min
Betting on one LLM provider is a database with no failover. How I route between models, chain fallbacks, and log which model served each request.
ai
architecture
llm
Resilient Engineering Teams Are Boring Teams
· 4 min
The engineering teams that got through 2022 best had the least drama: no heroics, no single points of failure, sane on-call, and bad news raised early.
teams
leadership
reliability
Tech Layoffs 2022: What I Saw From the Inside
· 5 min
What I saw during the 2022 tech layoff wave, and what helps engineering teams survive contraction without burning out.
leadership
teams
hiring
The December 2021 AWS us-east-1 Outage Was Predictable
· 5 min
The December 7 us-east-1 outage hit AWS control planes, not running workloads. Why it keeps catching teams out, and the multi-region basics that help.
cloud
incident-management
reliability
What a 3 AM Outage Taught Me About Incident Management
· 6 min
A 47-minute outage fixed by a 90-second rollback. How I run incidents now: symptom alerts, clear roles, mitigate first, short runbooks, real postmortems.
incident-management
reliability
SRE Team Structures: Stop Renaming Your Ops Team
· 5 min
Centralized, embedded, or platform SRE: how each model fails, the engagement tiers and entry criteria I use, and why renaming ops to SRE changes nothing.
reliability
teams
leadership
Postgres Reliability: Restores, Failover, Safe Migrations
· 8 min
Practical database reliability from running Postgres in production: configs, safe migration patterns, and the operational habits that prevent outages.
databases
reliability
infrastructure
Observability-Driven Development: Instrument Before You Ship
· 4 min
Observability-driven development without the jargon: structured logs, RED metrics, traces, and SLO alerts that ship with each feature.
observability
development
reliability
Chaos Engineering Without Hypotheses Is Theater
· 3 min
Teams love saying they do chaos engineering. Few actually have hypotheses. Even fewer fix what they find.
reliability
opinion
Business Continuity for Engineering Teams: Beyond the Binder
· 5 min
Most BCP documents are shelf-ware. What keeps engineering teams running in a crisis: no human single points of failure, real runbooks, and drills.
reliability
engineering
remote-work
Zero-Downtime Deployments: Migrations, Probes, and Habits
· 5 min
Zero-downtime deploys depend less on tooling than on expand-and-contract migrations, backward-compatible code, readiness probes, and graceful shutdown.
ci-cd
devops
kubernetes
Load Testing Strategies That Find Real Breaking Points
· 3 min
Most load tests produce comforting numbers instead of answers. Soak, spike, and baseline tests with production-shaped data, think time, and percentiles.
testing
performance
reliability
SLOs and Error Budgets That Change How Teams Ship
· 5 min
Most SLOs are dashboards nobody acts on. Pick indicators that reflect real users, set targets from data, and make error budgets change how your team ships.
reliability
observability
engineering
Designing for Failure: Timeouts, Circuit Breakers, Failover
· 3 min
Timeouts, blast-radius isolation, fallbacks, circuit breakers, and tested failover: the rules I follow so one slow dependency can't take down everything.
reliability
architecture
distributed-systems
Async Job Processing Patterns: Queues, Retries, Idempotency
· 6 min
Background job patterns from a fintech data pipeline: priority queues, idempotent workers, backoff with jitter, dead letter queues, and the outbox.
backend
architecture
reliability
Partial Failure in Distributed Systems: Lessons From Fintech
· 6 min
Designing distributed systems for partial failure at a fintech startup: timeouts, retries with jitter, circuit breakers, bulkheads, and idempotency.
distributed-systems
reliability
architecture
SRE Principles for Small Teams: Skip the Cargo Cult
· 4 min
Teams copy Google's SRE playbook without asking if it fits. What matters for small teams: one SLO, error budgets, toil, alerts, and postmortems.
reliability
devops
operations
Incident Management for Growing Teams: What to Change
· 5 min
How incident response changed as our fintech startup outgrew five people: incident roles, severity levels, sane on-call, mitigation first, real follow-up.
incident-management
devops
reliability
Chaos Engineering for Small Teams: No Netflix Required
· 4 min
Chaos engineering for a small team: start with whiteboard game days, then staging experiments with kill and tc. Always a hypothesis, always a stop button.
reliability
testing
devops
Data Pipelines That Survive Production: How I Build Them
· 6 min
Every data pipeline I built at a fintech startup broke. Raw storage, idempotent writes, boundary validation, and output-shape alerts made them recoverable.
data
reliability
architecture
Building Resilient Systems: Lessons from Production Failures
· 7 min
Production incidents show where architecture bends and breaks. Lessons on designing for failure, limiting blast radius, and making recovery routine.
reliability
architecture
engineering