// Topics / Testing

Testing

    Testing AI in Production: Shadow Mode, Canaries, Holdouts Offline evals aren't enough. How I test AI features in production with shadow mode, canaries, holdouts, and automatic fallback, with Go code. testing ai production AI Code Review Is Mostly Noise Months of AI code review on real PRs: about 22% of comments get accepted. How I scope prompts, track hit rate, and keep it out of merge gates. engineering ai development How I Test LLM Features: Three Layers, One Cadence LLM outputs are non-deterministic. That doesn't mean you can't test them rigorously. Here's the layered testing approach I use in production. llm testing ai LLM Evaluation: Stop Shipping on Vibes LLM evaluation that catches regressions before users do: a realistic test set, layered checks, offline and online evals, and deploy gates. ai llm testing AI Code Review: What It Catches and What It Misses Three months of AI-assisted code review on Go services: it catches unchecked errors and leaks, misses context, and only helps once you filter the noise. ai engineering developer-experience Testing Microservices: Contract Tests Over End-to-End Suites Microservices fail at the seams. A layered test strategy that keeps feedback fast and catches integration issues before production. testing microservices go Terraform Testing: Static Checks, Policy, and Integration I tested Terraform modules with static checks, policy engines, and integration runs side by side: what each layer catches, what it misses, when to run it. infrastructure testing infrastructure-as-code Load Testing Strategies That Find Real Breaking Points Most load tests produce comforting numbers instead of answers. Soak, spike, and baseline tests with production-shaped data, think time, and percentiles. testing performance reliability Testing in Production: Why Staging Misses the Real Bugs Staging misses the bugs that matter. How I test in production safely: feature flags, 1% rollouts, canaries, shadow traffic, and synthetic checks. testing production ci-cd Code Review Quality: Stop Counting Reviews, Start Reading Most code reviews are rubber stamps. What made ours useful at a fintech startup: read the change, skip style nits, label comments, small PRs, automation. engineering testing teams Chaos Engineering for Small Teams: No Netflix Required Chaos engineering for a small team: start with whiteboard game days, then staging experiments with kill and tc. Always a hypothesis, always a stop button. reliability testing devops