// Topics / LLMs

LLMs

    Video AI in Practice: Pipelines, Evaluation, Human Review Video AI is practical for scoped workflows. A reference pipeline, the evaluation metrics that matter for video, and where human review still belongs. llm ai production Running LLMs Locally: A Team Guide With Ollama and Go Local AI is no longer a hobby project. How to set it up properly: provider abstraction, versioned models, eval harnesses, and a cloud fallback. llm development privacy AI Documentation Search: Versioned Retrieval and Citations AI docs systems often retrieve the wrong version and never admit uncertainty. Build one with hybrid search, version filters, citations, and freshness jobs. engineering ai llm When to Fine-Tune an LLM: Prompt and Retrieve First Fine-tuning is the go-to move for teams who skipped the basics. Most of the time, better prompts and proper retrieval solve the actual problem. llm ai opinion AI Data Pipelines for RAG: ETL With a Retrieval Layer AI data pipelines are ETL with a retrieval layer bolted on. The discipline is the same as always: detect change, chunk intelligently, keep indexes fresh. data ai infrastructure Model Context Protocol in Go: Building an MCP Tool Server I built a Model Context Protocol server in Go with mcp-go. The protocol layer is clean. Auth, permissions, and write safety are still on you. agents ai go Video Understanding AI: Transcript First, Frames Second I built a video understanding pipeline for support call review. What worked: transcript first, scene-based frame sampling, timestamps on every claim. llm ai data Reasoning Models in Production: A Practical Guide Reasoning models are slow and expensive. How I run them in Go services: complexity routing, async jobs, per-request budgets, and result caching. llm production ai Picking an AI Model for Production (Late 2024) No LLM is best for everything. Compare task fit, tail latency, cost per success, and schema compliance in a bake-off on your workload, then route. ai llm reflection LLM Cost Benchmarking: Measure Cost per Completed Task Price-per-token is the least useful number on your AI bill. Real cost benchmarking starts with your workload, not a provider's pricing page. ai cost llm RAG Retrieval in Go: Hybrid Search, Chunking, Reranking Most RAG failures are retrieval failures. Hybrid search, structural chunking, query expansion, and reranking, each measured separately from generation. llm go How I Test LLM Features: Three Layers, One Cadence LLM outputs are non-deterministic. That doesn't mean you can't test them rigorously. Here's the layered testing approach I use in production. llm testing ai Small LLMs in Production: Use the Smallest Model That Works Most production LLM tasks are classification and extraction that don't need GPT-4. Where small models win, where they fail, and how to route. llm ai performance Context Window Budgets: Stop Stuffing Your LLM Prompts Bigger context windows don't excuse sloppy prompts. Budget tokens per section, pin anchors, retrieve less, and measure quality against context size. llm ai architecture LLM Function Calling Patterns That Survive Production Function calling is how LLMs touch real systems. Treat tools like APIs, arguments like untrusted input, and the model like an intern with root access. llm ai go Claude 3.5 Sonnet Analysis: Cost, Coding, and Model Routing Four days with Claude 3.5 Sonnet: coding quality, where Opus still wins, what Artifacts adds, and why it changes model routing for production. llm ai reflection Building Voice AI: Latency, Interruptions, and Narrow Scope Voice AI is ready to ship. The hard parts are latency, interruptions, and knowing when voice is the wrong interface. Here's how I approach it. llm ai go GPT-4o Changed the Interface, Not the Hard Part GPT-4o puts text, vision, and audio in one model. That removes pipeline glue, but transport, devices, and consent stay hard. How I'd evaluate it. llm architecture ai LLM Structured Output in Go: JSON Schema, Validation, Retries How to get reliable JSON from LLMs in Go with schemas, validation, repair loops, and typed contracts. llm api go LLM Response Caching in Go: Cut Costs Without Breaking Things LLM response caching in Go: versioned cache keys, TTLs by data freshness, event-driven invalidation, and what never to cache. llm performance go Multi-Model LLM Routing and Fallbacks in Production Betting on one LLM provider is a database with no failover. How I route between models, chain fallbacks, and log which model served each request. ai architecture llm Claude 3 First Impressions: Picking Opus, Sonnet, or Haiku Claude 3 shipped as three models. First notes on Opus, Sonnet, and Haiku quality, vision, and context, and how I route work between the tiers. llm ai LLM Evaluation: Stop Shipping on Vibes LLM evaluation that catches regressions before users do: a realistic test set, layered checks, offline and online evals, and deploy gates. ai llm testing Architecting AI-Native Applications (Without the Delusion) AI-native apps differ from a model bolted onto a CRUD app. The layers, confidence routing, fallbacks, and feedback loops I use, with Go code. architecture ai engineering Local LLMs for Development: Stop Paying to Test Prompts Local LLMs are good enough for development now. Iterate on prompts against Ollama, and save the API bill for evals and production traffic. llm engineering go AI Engineering Is Its Own Discipline Now AI engineering is not ML research with a product hat. It is the work of making models behave in production, and it needs its own hiring profile. ai hiring llm GPT-4V Use Cases: Five That Work, Three That Don't Weeks of testing GPT-4V on real workflows: receipts, UI review, alt text, and diagrams work; counting, measurements, and tiny text don't. Plus cost tips. ai llm business OpenAI Assistants API: Two Weeks of Real Use Two weeks building internal tools on OpenAI's Assistants API: quick wins with retrieval and code interpreter, opaque internals, and runs that hang. llm ai go OpenAI DevDay 2023: What It Means for Build vs Buy GPT-4 Turbo, the Assistants API, and GPTs moved OpenAI further up the stack. What that means for teams deciding which AI plumbing to build or buy. llm ai architecture LLM Security: A Field Guide for People Who Ship Things LLMs bring security failure modes most teams aren't defending against. Prompt injection, data leakage, tool abuse, and cost attacks are exploitable today. security llm ai AI Technical Debt: Prompts, Models, and Evals Nobody Tracks AI features pile up debt in unversioned prompts, unpinned models, unvalidated data, and missing evals. Warning signs and how to pay it down. ai technical-debt engineering AI Agent Architecture Patterns for Production Agent demos impress. Production agents mostly don't. Planning, memory, least-privilege tool access, and evals: the systems design that decides what ships. ai agents llm LLM Observability: Monitoring Quality, Not Just Uptime Traditional monitoring says the service is up. It won't tell you the model started returning garbage last Tuesday. How to actually observe LLM systems. observability llm ai What I Learned Building AI Features Into a Fintech Product Shipping AI transaction categorization in a fintech product: two-day demo, six weeks to production. Eval sets, input normalization, fallbacks, rollout. ai strategy business LLM Cost Optimization: The Basics Most Teams Skip A 12-person startup got a $14,000 OpenAI bill. The fix wasn't a FinOps tool: routing, caching, max_tokens, prompt trimming, per-feature cost tracking. ai cost llm Embedding Models Compared: Quality, Cost, and Latency ada-002, instructor-large, and all-MiniLM-L6-v2 on a 150-query retrieval eval: precision, MRR, latency, index size, cost, and when to self-host. llm ai go Building Semantic Search in Go with OpenAI and pgvector Building semantic search with Go, OpenAI embeddings, and pgvector: structure-aware chunking, hybrid retrieval, an eval set, and the mistakes I made. ai llm go Fine-Tuning vs. Prompting: A Decision Framework Most teams should exhaust prompting before they even think about fine-tuning. Here's how to decide which lever to pull. ai llm LangChain Is the New ORM: Convenient Until It Is Not LangChain promises to simplify LLM development. Instead it adds abstraction layers you will fight against the moment your use case gets real. llm ai opinion RAG in Production: Patterns That Survive Real Traffic RAG quality is retrieval quality. Chunking, hybrid search, query shaping, reranking, and evals for grounding LLMs in private data, with Go examples. llm ai go Vector Databases Explained: What They Are, When You Need One What vector databases store, how similarity search and ANN indexes work, and when pgvector is enough versus a dedicated vector database. llm ai go Claude vs GPT-4: Two Weeks Using Both Claude and GPT-4 both went public on March 14. How they compare on instructions, refusals, and reasoning, and why teams shouldn't marry one model. ai llm My First Week Building with GPT-4 My first week with GPT-4 after its March 14 launch: where it beats GPT-3.5, what it costs, and why I route requests instead of switching everything over. ai llm architecture Prompt Engineering Is Not Engineering The term 'prompt engineering' oversells what is mostly clear writing. What works in a prompt, where teams waste time, and where the hard work is. ai llm opinion LLM Integration Patterns That Survive Production LLM calls are slow, costly, and non-deterministic. Patterns for prompt versioning, structured output validation, RAG, tool guardrails, and fallbacks in Go. ai llm go AI in Production Is an Engineering Problem ChatGPT reset expectations overnight. Shipping LLM features that hold up takes timeouts, validation, caching, cost tracking, and fallbacks. ai production engineering Five Days With ChatGPT: First Impressions From an Engineer Five days with ChatGPT as a working engineer: fast first drafts, confidently wrong answers, and why verifying the output now matters more than typing it. ai llm developer-experience