// Topics / DevOps

DevOps

    AI Workflow Automation: Let the Model Decide, Let Code Act The trick to AI workflow automation is simple: let the model decide, let deterministic code act, and never confuse the two. devops ai agents Agentic Workflows in Production: Constrain the Blast Radius AI agents that take actions carry real blast radius. Policy allowlists, structured workflows, idempotent steps, tracing, and a shadow-mode rollout. agents ai production Terraform Patterns That Scale: Modules, State, Promotion Practical Terraform patterns for teams past the tutorial stage: module design, state management, environment promotion, and policy enforcement. infrastructure-as-code devops cloud Platform Engineering: DevOps Grew Up Platform engineering is what happens when you realize 'you build it, you run it' does not scale past a handful of teams. platform-engineering devops developer-experience Monorepo vs. Polyrepo: A Practical Decision Guide Monorepo or polyrepo depends on coupling, team shape, and your appetite for build tooling. Here is how to decide without getting religious about it. architecture development developer-experience Kubernetes Requests and Limits: Lessons From an Outage CPU is compressible and memory is not. How to set Kubernetes requests and limits from real usage, when to skip CPU limits, and which guardrails to add. kubernetes infrastructure devops Kubernetes Security Hardening: Pods, RBAC, Network, Secrets Kubernetes defaults favor convenience over security. A layered hardening guide covering pods, RBAC, network policies, secrets, and the control plane. kubernetes security devops DORA Metrics: Keep Them Off Performance Reviews DORA metrics work until someone puts them on a performance review. How to define, collect, and use them at team level without gaming. metrics devops productivity Terraform at Scale: What Changed Since 2019 Two years after my first Terraform-at-scale post: what held up, what broke, and what I do differently now with modules, workspaces, CI, and policy. infrastructure-as-code devops cloud Platform Engineering Maturity: Why Teams Stall at Level 2 Most platform teams build tools nobody asked for while developers wait in ticket queues. Lessons from platform teams at large companies. platform-engineering devops developer-experience Feature Flags at Scale: Ownership, Expiry, and Cleanup 847 feature flags, about 200 with owners. Lessons from production codebases on flag types, ownership rules, fail modes, and cleanup. ci-cd devops go Observability-Driven Development: Instrument Before You Ship Observability-driven development without the jargon: structured logs, RED metrics, traces, and SLO alerts that ship with each feature. observability development reliability DevSecOps in Practice: The Pipeline Controls I Implement Pre-commit hooks, CI security stages, OPA Gatekeeper policies, and a triage system: the DevSecOps controls I set up so developers don't route around them. security devops ci-cd MLOps Fundamentals: Fix the Basics Before Buying Tools Most teams buying MLOps tools can't version their training data. Start with dataset versioning, a release checklist, and one monitoring signal. ai devops data Platform Engineering vs DevOps: Mostly a Rebrand The industry loves renaming things. Platform teams and internal developer platforms are DevOps done properly, and most companies still won't do it right. platform-engineering devops infrastructure Observability for Small Distributed Teams: A Minimal Stack Most observability advice targets 500-engineer orgs. For a small distributed team: structured logs, a request_id, one dashboard per service, few alerts. observability distributed-systems devops GitHub Actions in Production: Matrix, Caching, Secrets Matrix builds, dependency caching, gated deploys, and the security gotchas I hit building a startup's CI/CD pipeline on GitHub Actions. ci-cd devops Kubernetes Requests and Limits: How I Size Resources Most K8s clusters I audit are either wildly overprovisioned or one bad deploy away from eviction storms. Here's how I set requests, limits, and guardrails. kubernetes devops infrastructure Terraform Testing: Static Checks, Policy, and Integration I tested Terraform modules with static checks, policy engines, and integration runs side by side: what each layer catches, what it misses, when to run it. infrastructure testing infrastructure-as-code Kubernetes Predictions for 2020: Day-Two Operations Win The adoption debate is over. 2020 is about operating Kubernetes well: managed control planes, GitOps by default, and policy enforcement. kubernetes trends cloud Zero-Downtime Deployments: Migrations, Probes, and Habits Zero-downtime deploys depend less on tooling than on expand-and-contract migrations, backward-compatible code, readiness probes, and graceful shutdown. ci-cd devops kubernetes Terraform at Scale: Splitting the State Monolith How to split a monolithic Terraform state into something teams can work with: state layout by domain, small pinned modules, plan review, and drift checks. infrastructure-as-code infrastructure devops Developer Experience: Internal Platforms vs Ad-Hoc Tooling Purpose-built internal platforms versus the scripts and Makefiles teams grow themselves, and the team size and pain at which each one wins. developer-experience platform-engineering devops Security Incident Response: Drills Beat Written Plans Most incident response plans are shelf-ware. What matters in a breach: simple severity levels, the first 30 minutes, four roles, evidence, and drills. security incident-management devops Internal Developer Platforms: Treat Them as Products Most internal developer platforms fail because nobody treated them as a product. Lessons from building, and scrapping, platform tooling at three startups. platform-engineering devops developer-experience GitOps with Flux and Argo CD: Stop Deploying From a Laptop How to move a team off ad-hoc kubectl deploys to Git-driven Kubernetes with Flux and Argo CD: repo layout, secrets, rollbacks, and my mistakes. ci-cd devops kubernetes Kubernetes Production Checklist: The Boring Basics Most Kubernetes outages come from skipped basics: limits, probes, network policies, RBAC, upgrades, etcd restores. The checklist I run on every cluster. kubernetes devops infrastructure Terraform in Production: Repo Layout, Modules, and State Opinionated Terraform patterns from a fintech startup: repo layout, small modules, per-environment state, CI plans, drift detection, and secrets. infrastructure infrastructure-as-code devops Container Security in 2018: PSP, Distroless, Image Signing Eight months after my first container security post: PodSecurityPolicy, distroless images, image signing, and Vault at a fintech startup. security containers kubernetes Monitoring vs Observability: Lessons From a Silent Outage After an outage our dashboards couldn't explain, I rebuilt a fintech startup's telemetry around metrics, logs, and traces tied together by request IDs. observability devops distributed-systems SRE Principles for Small Teams: Skip the Cargo Cult Teams copy Google's SRE playbook without asking if it fits. What matters for small teams: one SLO, error budgets, toil, alerts, and postmortems. reliability devops operations Kubernetes Operators: Powerful, but Overhyped Kubernetes operators are useful for Day 2 operations, but writing a good one is hard. When to adopt an existing operator, when to build your own. kubernetes devops infrastructure Two Years of Kubernetes in Production: The Boring Parts Year two of running Kubernetes in production: network policies, DNS, resource requests, PodDisruptionBudgets, upgrades, and RBAC. kubernetes containers devops Building a Platform Team: Lessons from Our First Year Standing up a small platform team at a fintech startup: tight scope, infrastructure run as a product, paved roads over mandates, and what I'd change. platform-engineering teams engineering Container Security Beyond the Basics: What We Hardened Containers share the host kernel, so they aren't a security boundary. How we hardened images, runtimes, network policies, and RBAC at a fintech startup. containers kubernetes security Incident Management for Growing Teams: What to Change How incident response changed as our fintech startup outgrew five people: incident roles, severity levels, sane on-call, mitigation first, real follow-up. incident-management devops reliability Chaos Engineering for Small Teams: No Netflix Required Chaos engineering for a small team: start with whiteboard game days, then staging experiments with kill and tc. Always a hypothesis, always a stop button. reliability testing devops Security Automation in CI: Stop Reviewing by Hand Manual security review can't keep up with continuous delivery. How I moved secret scanning, dependency audits, SAST, and DAST into a fintech CI pipeline. security devops ci-cd Observability vs Monitoring: Why Dashboards Aren't Enough Monitoring answers questions you predicted. After splitting into microservices, we needed structured logs, trace IDs, and tracing to debug what broke. observability devops distributed-systems Kubernetes in Production: What Paid Off and What Bit Us Running Kubernetes in production: what paid off, what bit us (networking, secrets, YAML sprawl), and who should adopt it. kubernetes containers devops Production Monitoring: Why We Deleted 42 Grafana Panels We cut 47 Grafana panels to five metrics and three paging alerts. The production metrics that matter for a startup backend, and how to prune the rest. observability devops production Container Orchestration: Docker Swarm vs Kubernetes vs Mesos Docker Swarm, Kubernetes, and Mesos compared side by side at a mobility startup in late 2016. Kubernetes will win, but its operational tax is real. containers kubernetes agents Building a Security-First Engineering Culture Security culture is habits leadership enforces: no secrets in code, a security question on every PR, least privilege, patch deadlines, and champions. security engineering teams Log Aggregation at Scale: ELK vs Alternatives ELK is powerful and a second full-time job. What running it at a mobility startup taught me, and the hosted or simpler options I'd pick instead. observability databases devops Zero-Downtime PostgreSQL Migrations: Expand and Contract How to change PostgreSQL schemas without maintenance windows: expand and contract, batched backfills, NOT VALID constraints, and concurrent indexes. databases devops engineering Why I Moved Our AWS Infrastructure to Terraform We moved from console-driven, script-heavy infrastructure to Terraform so changes are reviewed, reproducible, and recoverable from code. infrastructure-as-code devops cloud Continuous Deployment Without the Chaos Continuous deployment is a discipline problem. How we ship a mobility startup's backend many times a day: trusted tests, small changes, fast rollback. ci-cd devops engineering Security Incident Response Playbook for Startups An incident response playbook for small teams, learned the hard way: define incidents, name owners, contain carefully, and communicate clearly. security incident-management startups Ansible vs Puppet vs Chef: Why Ansible Won for Us Puppet, Chef, and Ansible compared from production use. Ansible wins on least ceremony; Puppet and Chef still fit large or Ruby-fluent infra teams. infrastructure-as-code devops infrastructure Building a DevOps Culture from Scratch Hiring a DevOps engineer won't end the dev vs ops fight. Put developers on the pager, pilot with one team, and fix incentives before tools. devops teams engineering Docker in Production: Lessons from Running Containers What running Docker in production at a mobility startup taught us about image builds, tagging, networking, logs, resource limits, and non-root security. containers devops infrastructure