Everyone demos an agent in a notebook. Few teams run one safely in production beside paying users. The gap is engineering: boundaries, evaluation, cost control, and failure modes that do not wake you up at 3 a.m.
Orchestration, Not Magic
Production agents are graphs — retrieve, plan, act, verify. Each step has timeouts, retry policies, and structured outputs validated before the next hop. Treat LLM calls like external APIs with SLAs and circuit breakers.
RAG Done Properly
Chunking strategy, metadata filters, hybrid search (vector + keyword), and citation requirements reduce hallucination. Re-embed when source documents change; version your index like you version schema.
Human-in-the-Loop
High-impact actions — refunds, deletes, outbound emails — need approval queues. Log prompts, retrieved context, model version, and token usage for audit and debugging.
Evaluation Before Scale
Golden datasets, regression evals on prompt changes, and red-team tests for injection attacks. If you cannot measure quality, you cannot improve it — or trust it in prod.