Agent Production Engineering
How does an agent run reliably in a real environment?
Add evals, observability, budgets, human-in-the-loop approval, and recovery.
📌 Learning goals
- Tell apart where the helper works, how it repeats, and the full route with branches.
- Turn real failures into rerunnable tests instead of trusting one beautiful answer.
- Find every step, error, timing, and cost inside one task run.
- Make high-risk actions pause for a person, and resume from the right position.
- Use the same evidence to judge whether the system can be handed to someone else.
Entry conditions
Finish Stages 3–4; you need Python and Git (plus Docker for deployment exercises). The shipping order: define success → keep a work record → ask a person before risky actions → prove it can continue after a fall → only then hand it to other people.
🧭 Lessons on this site
Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.- 01LLM Evaluation and Selection Guide: How to Scientifically Choose Models in 2026
A systematic LLM selection framework: Benchmark interpretation, real-world scenario testing methods, cost/performance matrix, DeepSeek V4 vs Qwen vs Llama comparison.
4 min - 02From Fabrication to Verification: The Trust Architecture of Agent Verification
When an LLM Agent deteriorated from 'skipping verification' to 'fabricating verification records' across 6 runs, we learned a fundamental lesson: plain-text rules cannot constrain an LLM. An External Supervisor is the only reliable solution.
6 min - 03Comparison of AI Agent Verification Architectures: From Prompt Self-Awareness to Architectural Enforcement, a Source-Level Analysis of Seven Approaches
Who verifies that every step of an AI Agent is correct? This compares the execution-verification separation mechanisms of seven approaches: Claude Code, MetaGPT, OpenHands, VIGIL, PEV, ReVeal, and odot. It finds that all reliable approaches follow the same iron rule: the executor cannot also be the judge.
27 min - 04A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
In September 2026, System One decision models formed a new category within two weeks: the closed-source JEV API, then LAYA, KEV, and CLM-8B open-sourced one after another. This article does more than summarize the differences among the four; it uses six everyday scenarios to explain how they are actually used, and places the official marketing side by side with our same-question measurements: on the same set of questions, full-coverage accuracy was only 54%, but with confidence gating it reached 91.7%.
20 min - 05An Empirical Analysis of LLM Agents Autonomously Bypassing Process Constraints: The Deterioration Path from 'Skip' to 'Fabricate'
In the production environment, we observed an LLM agent skipping a mandatory verification step 6 times in a row, escalating to fabricating verification records on the 6th run. This article provides the complete experimental data, a five-layer root cause analysis, cross-model predictive analysis, and an architecture-level solution.
17 min - 06AI Security and Red Teaming: Prompt Injection, Jailbreak Defense
A full landscape of security threats to AI systems: Prompt Injection, Jailbreak, Data Poisoning, and practical defense strategies and red team testing methods.
7 min - 07Loop Engineering: A Deep Retrospective on a Field Failure, When the Design Documents Cannot Be Executed
A complete failure log from building a Loop Engineering system from scratch: we designed a perfect architecture, 36 files, and 10 cron jobs, but not a single line of code ever actually ran the inner loop.
4 min - 08Ten-Day Pitfall Log: 16 Fatal Lessons in Building an AI Assistant System
Over the past 10 days, our AI assistant evolved from a conversational bot into a system-level AI assistant. This is the complete pitfall record: 16 pitfalls, 5 root cause patterns, 12 mandatory rules, and the meta-principles for system building distilled from them.
20 min
📚 Required reading
- 1.Anthropic — Demystifying evals for AI agents⭐⭐⭐⭐⭐Separate outcome from full trajectory; an agent saying "done" does not mean the external result is done.
- 2.OpenAI Agents SDK — Tracing⭐⭐⭐⭐⭐See how trace, span, tool, handoff, and guardrail events chain into one run.
- 3.OpenAI Agents SDK — Human-in-the-loop⭐⭐⭐⭐⭐Pause sensitive tools, save the run state, approve or reject, then resume.
- 4.LangGraph — Persistence⭐⭐⭐⭐Tell checkpoints apart from cross-thread stores; interrupt, recovery, and long-term memory are different things.
- 5.LangGraph — Interrupts⭐⭐⭐⭐See how human approval pauses and resumes; side effects before an interrupt must be idempotent.
- 6.Anthropic — Building Effective Agents⭐⭐⭐⭐⭐Start with simple compositions; add autonomy or multi-agent only when division of labor is truly needed.
🎯 Curated resources
| Resource | Who it's for | Priority | Why |
|---|---|---|---|
Evals promptfoo | Putting evals into CI | ⭐⭐⭐⭐⭐ | Build success criteria and graders; a config file is no substitute for a good rubric. |
Evals Anthropic — Develop tests and evaluations | Writing measurable success criteria first | ⭐⭐⭐⭐⭐ | Cases still come from your own real work and failures. |
Observability langfuse/langfuse | Need traces, evals, and prompt management | ⭐⭐⭐⭐⭐ | Self-hosting still needs operations and data governance. |
Observability Arize-ai/phoenix | OpenTelemetry with local analysis | ⭐⭐⭐⭐ | Design sensitive-data masking first. |
Observability OpenTelemetry GenAI conventions | Learning portable trace fields | ⭐⭐⭐⭐ | The spec is still evolving; platform support varies. |
HITL / recovery LangGraph — Interrupts | Approval, checkpoints, resume | ⭐⭐⭐⭐⭐ | Production needs a durable checkpointer, not memory alone; side effects must be idempotent. |
Harness / sandbox anthropics/claude-agent-sdk-python | Reading the tool loop, permissions, and subagent implementation | ⭐⭐⭐⭐⭐ | Centred on the Claude runtime. |
Harness / sandbox OpenAI — Harness engineering | Environment, feedback loops, and machine rules | ⭐⭐⭐⭐ | See how the environment, feedback loop, and machine rules keep an agent stable. |
Deploy bentoml/BentoML | Packaging apps as services and containers | ⭐⭐⭐⭐ | A deployment framework does not add evals and guardrails for you. |
Multi-agent cases crewAIInc/crewAI | Understanding role-based division of labor | ⭐⭐⭐⭐ | More roles do not guarantee better answers. |
Multi-agent cases stablyai/orca | Running coding agents in parallel worktrees | ⭐⭐⭐⭐ | Parallel results still need human review and selection. |
Benchmarks SWE-bench | Learning what benchmark sets look like | ⭐⭐⭐⭐ | Read alongside tau2-bench, GAIA, OSWorld, and terminal-bench; scores never replace your own cases. |
🛠 Hands-on practice (upstream)
Full exercises & starter codeSummaries from the upstream curriculum; full code, cost, and latency estimates live upstream.
- One story for the whole chapter: an AI helper checks three sources, writes a summary, and asks a person before sending — you gradually add its workspace, repeat rhythm, branching routes, quality checks, and a safety brake.
- Exercise spine: fixed cases and success rules → step-level traces and cost records → pause-and-ask for risky actions → checkpoint resume with idempotent side effects.
- Every exercise starts with a test that needs no API key; set a small budget before calling paid models.
- Work records may contain prompts and model replies: never send passwords, personal data, or customer data straight to a tracing platform.
- Each extra agent adds model calls, latency, and debugging — stabilize one helper first.
✅ Self-check
- I can tell outcome from trajectory and check the same case with both.
- I have fixed eval cases built from real failures, not one beautiful output.
- I can find the trace, errors, latency, and tokens of one run.
- High-risk tools have least privilege and human approval; without approval they fail closed.
- I can resume from a checkpoint and prove the same idempotency key causes no duplicate side effects.
- I can show an execution receipt and explain when to stop, recover, or roll back.
Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.