// Topics / Performance

Performance

    AI Inference Cost Trends 2026: LLM Prices, September Update LLM API prices now run from $0.10 to $10 per million input tokens. A sourced September 2026 price table, the trend since GPT-4, and where savings come from. cost ai performance LLM Cost Benchmarking: Measure Cost per Completed Task Price-per-token is the least useful number on your AI bill. Real cost benchmarking starts with your workload, not a provider's pricing page. ai cost llm Small LLMs in Production: Use the Smallest Model That Works Most production LLM tasks are classification and extraction that don't need GPT-4. Where small models win, where they fail, and how to route. llm ai performance LLM Response Caching in Go: Cut Costs Without Breaking Things LLM response caching in Go: versioned cache keys, TTLs by data freshness, event-driven invalidation, and what never to cache. llm performance go Cloud Cost Optimization Starts With Visibility Most cloud cost problems are visibility problems. Fix tagging, kill idle resources, right-size the rest, and make cost a regular engineering conversation. cost cloud infrastructure Caching Strategies: Adding It Is Easy, Invalidation Is Hard Cache-aside, write-through, invalidation strategies, and the failure modes that will wake you up at night. With Go examples. performance databases go PostgreSQL Performance: Measure First, Tune Second Most Postgres performance problems are indexing problems. The rest are vacuum problems. Here's how to find and fix both. databases performance backend Rust vs Go for Cloud Services: A Go Developer's View A Go developer on Rust in early 2021: where it wins (tail latency, memory, compile-time safety), where Go still wins, and when a rewrite pays off. engineering go cloud eBPF for Observability: Interesting, but I'm Not Sold Yet eBPF gives kernel-level observability without kernel modules. Kernel version gaps, patchy BTF support, and scarce expertise make it hard to run today. observability engineering performance Load Testing Strategies That Find Real Breaking Points Most load tests produce comforting numbers instead of answers. Soak, spike, and baseline tests with production-shaped data, think time, and percentiles. testing performance reliability PostgreSQL Performance Tuning: The Playbook I Use PostgreSQL tuning in the order I do it: find slow queries with pg_stat_statements, pool connections, size memory, fix indexes, tune autovacuum and WAL. databases performance backend Making Go Services Fast: Profiling, Allocations, Timeouts Go performance tuning from production: pprof profiling, cutting allocations, bounded concurrency, HTTP timeouts, and database pool settings. go performance backend Rust vs Go for Backend Services: A Go Developer's First Look A Go developer tries Rust for backend work in early 2018: what impressed me, where it still hurts, and the one service where it might fit. engineering go backend PostgreSQL Performance Tuning: Stop Guessing, Measure First How I diagnose slow PostgreSQL queries: slow query log, EXPLAIN ANALYZE, stale statistics, targeted indexes, connection pooling, one change at a time. databases performance