// Topics / Production
Production
27 entries tagged “Production”
- AI Evaluation and Production Governance: A Maturity Model
· 4 min
A five-level maturity model for AI evaluation and production governance, from vibes-based deploys to CI eval gates, production sampling, and rollback.
governance
ai
reliability
AI Security in 2026: Prompt Injection, Agents, and Defenses
· 6 min
Multi-step prompt injection and tool-based exfiltration target AI agents in 2026. The layered defenses, monitoring, and security review I use.
security
ai
production
AI-Native Architecture Patterns 2026: Production Guide
· 6 min
AI-native architecture patterns for production: an AI gateway, a retrieval layer, an evaluation pipeline, fallbacks, cost control, and clear ownership.
architecture
ai
engineering
Video AI in Practice: Pipelines, Evaluation, Human Review
· 4 min
Video AI is practical for scoped workflows. A reference pipeline, the evaluation metrics that matter for video, and where human review still belongs.
llm
ai
production
AI Incident Response: Failures That Don't Look Like Outages
· 4 min
AI systems can return 200 OK while confidently wrong. How to detect, contain, and learn from AI incidents using proven incident response principles.
incident-management
ai
reliability
AI Workflow Automation: Let the Model Decide, Let Code Act
· 4 min
The trick to AI workflow automation is simple: let the model decide, let deterministic code act, and never confuse the two.
devops
ai
agents
AI Customer Support: Design for Handoff, Not Deflection
· 4 min
Most AI support systems are built to deflect tickets. The ones that work are built around escalation, grounding, and the idea that customers aren't idiots.
ai
production
AI Security: Prompt Injection and the New Attack Surface
· 5 min
AI systems are exposed APIs with real blast radius. The threats are injection, leakage, and tool misuse. The defenses are the ones we've always needed.
security
ai
production
Testing AI in Production: Shadow Mode, Canaries, Holdouts
· 6 min
Offline evals aren't enough. How I test AI features in production with shadow mode, canaries, holdouts, and automatic fallback, with Go code.
testing
ai
production
AI Observability: Green Dashboards, Wrong Answers
· 4 min
Uptime monitoring won't show an AI service returning confident garbage. The five signals I track instead: traces, quality, cost per outcome, safety, ops.
observability
ai
production
Reasoning Models in Production: A Practical Guide
· 7 min
Reasoning models are slow and expensive. How I run them in Go services: complexity routing, async jobs, per-request budgets, and result caching.
llm
production
ai
AI Infrastructure at Scale Is Just Infrastructure
· 4 min
AI infrastructure at scale is just infrastructure. Gateways, caching, workload separation, budgets, and circuit breakers solve the same old problems.
ai
infrastructure
architecture
AI Safety in Production Is Defense in Depth
· 5 min
Production AI safety works like cyber defense: assume breach, layer input, output, and system controls, and watch every boundary with real monitoring.
ai
governance
production
LLM Function Calling Patterns That Survive Production
· 7 min
Function calling is how LLMs touch real systems. Treat tools like APIs, arguments like untrusted input, and the model like an intern with root access.
llm
ai
go
Agentic Workflows in Production: Constrain the Blast Radius
· 5 min
AI agents that take actions carry real blast radius. Policy allowlists, structured workflows, idempotent steps, tracing, and a shadow-mode rollout.
agents
ai
production
Multi-Model LLM Routing and Fallbacks in Production
· 4 min
Betting on one LLM provider is a database with no failover. How I route between models, chain fallbacks, and log which model served each request.
ai
architecture
llm
AI Engineering Is Its Own Discipline Now
· 4 min
AI engineering is not ML research with a product hat. It is the work of making models behave in production, and it needs its own hiring profile.
ai
hiring
llm
LLM Observability: Monitoring Quality, Not Just Uptime
· 5 min
Traditional monitoring says the service is up. It won't tell you the model started returning garbage last Tuesday. How to actually observe LLM systems.
observability
llm
ai
AI in Production Is an Engineering Problem
· 3 min
ChatGPT reset expectations overnight. Shipping LLM features that hold up takes timeouts, validation, caching, cost tracking, and fallbacks.
ai
production
engineering
Testing in Production: Why Staging Misses the Real Bugs
· 5 min
Staging misses the bugs that matter. How I test in production safely: feature flags, 1% rollouts, canaries, shadow traffic, and synthetic checks.
testing
production
ci-cd
Kubernetes Production Checklist: The Boring Basics
· 5 min
Most Kubernetes outages come from skipped basics: limits, probes, network policies, RBAC, upgrades, etcd restores. The checklist I run on every cluster.
kubernetes
devops
infrastructure
GraphQL in Production: N+1 Queries, Caching, and Tracing
· 4 min
GraphQL in production at a fintech startup: N+1 queries, query cost limits, caching, field-level tracing, and what the conference talks leave out.
api
backend
production
Two Years of Kubernetes in Production: The Boring Parts
· 6 min
Year two of running Kubernetes in production: network policies, DNS, resource requests, PodDisruptionBudgets, upgrades, and RBAC.
kubernetes
containers
devops
Kubernetes in Production: What Paid Off and What Bit Us
· 6 min
Running Kubernetes in production: what paid off, what bit us (networking, secrets, YAML sprawl), and who should adopt it.
kubernetes
containers
devops
Production Monitoring: Why We Deleted 42 Grafana Panels
· 3 min
We cut 47 Grafana panels to five metrics and three paging alerts. The production metrics that matter for a startup backend, and how to prune the rest.
observability
devops
production
Building Resilient Systems: Lessons from Production Failures
· 7 min
Production incidents show where architecture bends and breaks. Lessons on designing for failure, limiting blast radius, and making recovery routine.
reliability
architecture
engineering
Docker in Production: Lessons from Running Containers
· 8 min
What running Docker in production at a mobility startup taught us about image builds, tagging, networking, logs, resource limits, and non-root security.
containers
devops
infrastructure