// Topics / LLMs
LLMs
47 entries tagged “LLMs”
- Video AI in Practice: Pipelines, Evaluation, Human Review
· 4 min
Video AI is practical for scoped workflows. A reference pipeline, the evaluation metrics that matter for video, and where human review still belongs.
llm
ai
production
Running LLMs Locally: A Team Guide With Ollama and Go
· 6 min
Local AI is no longer a hobby project. How to set it up properly: provider abstraction, versioned models, eval harnesses, and a cloud fallback.
llm
development
privacy
AI Documentation Search: Versioned Retrieval and Citations
· 4 min
AI docs systems often retrieve the wrong version and never admit uncertainty. Build one with hybrid search, version filters, citations, and freshness jobs.
engineering
ai
llm
When to Fine-Tune an LLM: Prompt and Retrieve First
· 4 min
Fine-tuning is the go-to move for teams who skipped the basics. Most of the time, better prompts and proper retrieval solve the actual problem.
llm
ai
opinion
AI Data Pipelines for RAG: ETL With a Retrieval Layer
· 5 min
AI data pipelines are ETL with a retrieval layer bolted on. The discipline is the same as always: detect change, chunk intelligently, keep indexes fresh.
data
ai
infrastructure
Model Context Protocol in Go: Building an MCP Tool Server
· 7 min
I built a Model Context Protocol server in Go with mcp-go. The protocol layer is clean. Auth, permissions, and write safety are still on you.
agents
ai
go
Video Understanding AI: Transcript First, Frames Second
· 4 min
I built a video understanding pipeline for support call review. What worked: transcript first, scene-based frame sampling, timestamps on every claim.
llm
ai
data
Reasoning Models in Production: A Practical Guide
· 7 min
Reasoning models are slow and expensive. How I run them in Go services: complexity routing, async jobs, per-request budgets, and result caching.
llm
production
ai
Picking an AI Model for Production (Late 2024)
· 5 min
No LLM is best for everything. Compare task fit, tail latency, cost per success, and schema compliance in a bake-off on your workload, then route.
ai
llm
reflection
LLM Cost Benchmarking: Measure Cost per Completed Task
· 4 min
Price-per-token is the least useful number on your AI bill. Real cost benchmarking starts with your workload, not a provider's pricing page.
ai
cost
llm
RAG Retrieval in Go: Hybrid Search, Chunking, Reranking
· 7 min
Most RAG failures are retrieval failures. Hybrid search, structural chunking, query expansion, and reranking, each measured separately from generation.
llm
go
How I Test LLM Features: Three Layers, One Cadence
· 5 min
LLM outputs are non-deterministic. That doesn't mean you can't test them rigorously. Here's the layered testing approach I use in production.
llm
testing
ai
Small LLMs in Production: Use the Smallest Model That Works
· 3 min
Most production LLM tasks are classification and extraction that don't need GPT-4. Where small models win, where they fail, and how to route.
llm
ai
performance
Context Window Budgets: Stop Stuffing Your LLM Prompts
· 4 min
Bigger context windows don't excuse sloppy prompts. Budget tokens per section, pin anchors, retrieve less, and measure quality against context size.
llm
ai
architecture
LLM Function Calling Patterns That Survive Production
· 7 min
Function calling is how LLMs touch real systems. Treat tools like APIs, arguments like untrusted input, and the model like an intern with root access.
llm
ai
go
Claude 3.5 Sonnet Analysis: Cost, Coding, and Model Routing
· 5 min
Four days with Claude 3.5 Sonnet: coding quality, where Opus still wins, what Artifacts adds, and why it changes model routing for production.
llm
ai
reflection
Building Voice AI: Latency, Interruptions, and Narrow Scope
· 5 min
Voice AI is ready to ship. The hard parts are latency, interruptions, and knowing when voice is the wrong interface. Here's how I approach it.
llm
ai
go
GPT-4o Changed the Interface, Not the Hard Part
· 4 min
GPT-4o puts text, vision, and audio in one model. That removes pipeline glue, but transport, devices, and consent stay hard. How I'd evaluate it.
llm
architecture
ai
LLM Structured Output in Go: JSON Schema, Validation, Retries
· 7 min
How to get reliable JSON from LLMs in Go with schemas, validation, repair loops, and typed contracts.
llm
api
go
LLM Response Caching in Go: Cut Costs Without Breaking Things
· 6 min
LLM response caching in Go: versioned cache keys, TTLs by data freshness, event-driven invalidation, and what never to cache.
llm
performance
go
Multi-Model LLM Routing and Fallbacks in Production
· 4 min
Betting on one LLM provider is a database with no failover. How I route between models, chain fallbacks, and log which model served each request.
ai
architecture
llm
Claude 3 First Impressions: Picking Opus, Sonnet, or Haiku
· 4 min
Claude 3 shipped as three models. First notes on Opus, Sonnet, and Haiku quality, vision, and context, and how I route work between the tiers.
llm
ai
LLM Evaluation: Stop Shipping on Vibes
· 5 min
LLM evaluation that catches regressions before users do: a realistic test set, layered checks, offline and online evals, and deploy gates.
ai
llm
testing
Architecting AI-Native Applications (Without the Delusion)
· 7 min
AI-native apps differ from a model bolted onto a CRUD app. The layers, confidence routing, fallbacks, and feedback loops I use, with Go code.
architecture
ai
engineering
Local LLMs for Development: Stop Paying to Test Prompts
· 3 min
Local LLMs are good enough for development now. Iterate on prompts against Ollama, and save the API bill for evals and production traffic.
llm
engineering
go
AI Engineering Is Its Own Discipline Now
· 4 min
AI engineering is not ML research with a product hat. It is the work of making models behave in production, and it needs its own hiring profile.
ai
hiring
llm
GPT-4V Use Cases: Five That Work, Three That Don't
· 5 min
Weeks of testing GPT-4V on real workflows: receipts, UI review, alt text, and diagrams work; counting, measurements, and tiny text don't. Plus cost tips.
ai
llm
business
OpenAI Assistants API: Two Weeks of Real Use
· 4 min
Two weeks building internal tools on OpenAI's Assistants API: quick wins with retrieval and code interpreter, opaque internals, and runs that hang.
llm
ai
go
OpenAI DevDay 2023: What It Means for Build vs Buy
· 4 min
GPT-4 Turbo, the Assistants API, and GPTs moved OpenAI further up the stack. What that means for teams deciding which AI plumbing to build or buy.
llm
ai
architecture
LLM Security: A Field Guide for People Who Ship Things
· 6 min
LLMs bring security failure modes most teams aren't defending against. Prompt injection, data leakage, tool abuse, and cost attacks are exploitable today.
security
llm
ai
AI Technical Debt: Prompts, Models, and Evals Nobody Tracks
· 4 min
AI features pile up debt in unversioned prompts, unpinned models, unvalidated data, and missing evals. Warning signs and how to pay it down.
ai
technical-debt
engineering
AI Agent Architecture Patterns for Production
· 6 min
Agent demos impress. Production agents mostly don't. Planning, memory, least-privilege tool access, and evals: the systems design that decides what ships.
ai
agents
llm
LLM Observability: Monitoring Quality, Not Just Uptime
· 5 min
Traditional monitoring says the service is up. It won't tell you the model started returning garbage last Tuesday. How to actually observe LLM systems.
observability
llm
ai
What I Learned Building AI Features Into a Fintech Product
· 5 min
Shipping AI transaction categorization in a fintech product: two-day demo, six weeks to production. Eval sets, input normalization, fallbacks, rollout.
ai
strategy
business
LLM Cost Optimization: The Basics Most Teams Skip
· 4 min
A 12-person startup got a $14,000 OpenAI bill. The fix wasn't a FinOps tool: routing, caching, max_tokens, prompt trimming, per-feature cost tracking.
ai
cost
llm
Embedding Models Compared: Quality, Cost, and Latency
· 6 min
ada-002, instructor-large, and all-MiniLM-L6-v2 on a 150-query retrieval eval: precision, MRR, latency, index size, cost, and when to self-host.
llm
ai
go
Building Semantic Search in Go with OpenAI and pgvector
· 7 min
Building semantic search with Go, OpenAI embeddings, and pgvector: structure-aware chunking, hybrid retrieval, an eval set, and the mistakes I made.
ai
llm
go
Fine-Tuning vs. Prompting: A Decision Framework
· 4 min
Most teams should exhaust prompting before they even think about fine-tuning. Here's how to decide which lever to pull.
ai
llm
LangChain Is the New ORM: Convenient Until It Is Not
· 4 min
LangChain promises to simplify LLM development. Instead it adds abstraction layers you will fight against the moment your use case gets real.
llm
ai
opinion
RAG in Production: Patterns That Survive Real Traffic
· 8 min
RAG quality is retrieval quality. Chunking, hybrid search, query shaping, reranking, and evals for grounding LLMs in private data, with Go examples.
llm
ai
go
Vector Databases Explained: What They Are, When You Need One
· 6 min
What vector databases store, how similarity search and ANN indexes work, and when pgvector is enough versus a dedicated vector database.
llm
ai
go
Claude vs GPT-4: Two Weeks Using Both
· 3 min
Claude and GPT-4 both went public on March 14. How they compare on instructions, refusals, and reasoning, and why teams shouldn't marry one model.
ai
llm
My First Week Building with GPT-4
· 4 min
My first week with GPT-4 after its March 14 launch: where it beats GPT-3.5, what it costs, and why I route requests instead of switching everything over.
ai
llm
architecture
Prompt Engineering Is Not Engineering
· 3 min
The term 'prompt engineering' oversells what is mostly clear writing. What works in a prompt, where teams waste time, and where the hard work is.
ai
llm
opinion
LLM Integration Patterns That Survive Production
· 6 min
LLM calls are slow, costly, and non-deterministic. Patterns for prompt versioning, structured output validation, RAG, tool guardrails, and fallbacks in Go.
ai
llm
go
AI in Production Is an Engineering Problem
· 3 min
ChatGPT reset expectations overnight. Shipping LLM features that hold up takes timeouts, validation, caching, cost tracking, and fallbacks.
ai
production
engineering
Five Days With ChatGPT: First Impressions From an Engineer
· 4 min
Five days with ChatGPT as a working engineer: fast first drafts, confidently wrong answers, and why verifying the output now matters more than typing it.
ai
llm
developer-experience