Agentic Research
Agent Learning Roadmap · Stage 7

Agent Production Engineering

How does an agent run reliably in a real environment?

Add evals, observability, budgets, human-in-the-loop approval, and recovery.

12–20 hours8 mapped lessonsUpstream edition

📌 Learning goals

  • Tell apart where the helper works, how it repeats, and the full route with branches.
  • Turn real failures into rerunnable tests instead of trusting one beautiful answer.
  • Find every step, error, timing, and cost inside one task run.
  • Make high-risk actions pause for a person, and resume from the right position.
  • Use the same evidence to judge whether the system can be handed to someone else.

Entry conditions

Finish Stages 3–4; you need Python and Git (plus Docker for deployment exercises). The shipping order: define success → keep a work record → ask a person before risky actions → prove it can continue after a fall → only then hand it to other people.

🧭 Lessons on this site

Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.
Progress here
0/8
Saved in your browser only
  1. 01
    LLM Evaluation and Selection Guide: How to Scientifically Choose Models in 2026

    A systematic LLM selection framework: Benchmark interpretation, real-world scenario testing methods, cost/performance matrix, DeepSeek V4 vs Qwen vs Llama comparison.

    4 min
  2. 02
    From Fabrication to Verification: The Trust Architecture of Agent Verification

    When an LLM Agent deteriorated from 'skipping verification' to 'fabricating verification records' across 6 runs, we learned a fundamental lesson: plain-text rules cannot constrain an LLM. An External Supervisor is the only reliable solution.

    6 min
  3. 03
    Comparison of AI Agent Verification Architectures: From Prompt Self-Awareness to Architectural Enforcement, a Source-Level Analysis of Seven Approaches

    Who verifies that every step of an AI Agent is correct? This compares the execution-verification separation mechanisms of seven approaches: Claude Code, MetaGPT, OpenHands, VIGIL, PEV, ReVeal, and odot. It finds that all reliable approaches follow the same iron rule: the executor cannot also be the judge.

    27 min
  4. 04
    A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula

    In September 2026, System One decision models formed a new category within two weeks: the closed-source JEV API, then LAYA, KEV, and CLM-8B open-sourced one after another. This article does more than summarize the differences among the four; it uses six everyday scenarios to explain how they are actually used, and places the official marketing side by side with our same-question measurements: on the same set of questions, full-coverage accuracy was only 54%, but with confidence gating it reached 91.7%.

    20 min
  5. 05
    An Empirical Analysis of LLM Agents Autonomously Bypassing Process Constraints: The Deterioration Path from 'Skip' to 'Fabricate'

    In the production environment, we observed an LLM agent skipping a mandatory verification step 6 times in a row, escalating to fabricating verification records on the 6th run. This article provides the complete experimental data, a five-layer root cause analysis, cross-model predictive analysis, and an architecture-level solution.

    17 min
  6. 06
    AI Security and Red Teaming: Prompt Injection, Jailbreak Defense

    A full landscape of security threats to AI systems: Prompt Injection, Jailbreak, Data Poisoning, and practical defense strategies and red team testing methods.

    7 min
  7. 07
    Loop Engineering: A Deep Retrospective on a Field Failure, When the Design Documents Cannot Be Executed

    A complete failure log from building a Loop Engineering system from scratch: we designed a perfect architecture, 36 files, and 10 cron jobs, but not a single line of code ever actually ran the inner loop.

    4 min
  8. 08
    Ten-Day Pitfall Log: 16 Fatal Lessons in Building an AI Assistant System

    Over the past 10 days, our AI assistant evolved from a conversational bot into a system-level AI assistant. This is the complete pitfall record: 16 pitfalls, 5 root cause patterns, 12 mandatory rules, and the meta-principles for system building distilled from them.

    20 min

📚 Required reading

  1. 1.Anthropic — Demystifying evals for AI agents⭐⭐⭐⭐⭐Separate outcome from full trajectory; an agent saying "done" does not mean the external result is done.
  2. 2.OpenAI Agents SDK — Tracing⭐⭐⭐⭐⭐See how trace, span, tool, handoff, and guardrail events chain into one run.
  3. 3.OpenAI Agents SDK — Human-in-the-loop⭐⭐⭐⭐⭐Pause sensitive tools, save the run state, approve or reject, then resume.
  4. 4.LangGraph — Persistence⭐⭐⭐⭐Tell checkpoints apart from cross-thread stores; interrupt, recovery, and long-term memory are different things.
  5. 5.LangGraph — Interrupts⭐⭐⭐⭐See how human approval pauses and resumes; side effects before an interrupt must be idempotent.
  6. 6.Anthropic — Building Effective Agents⭐⭐⭐⭐⭐Start with simple compositions; add autonomy or multi-agent only when division of labor is truly needed.

🎯 Curated resources

ResourceWho it's forPriorityWhy
Evals
promptfoo
Putting evals into CI⭐⭐⭐⭐⭐Build success criteria and graders; a config file is no substitute for a good rubric.
Evals
Anthropic — Develop tests and evaluations
Writing measurable success criteria first⭐⭐⭐⭐⭐Cases still come from your own real work and failures.
Observability
langfuse/langfuse
Need traces, evals, and prompt management⭐⭐⭐⭐⭐Self-hosting still needs operations and data governance.
Observability
Arize-ai/phoenix
OpenTelemetry with local analysis⭐⭐⭐⭐Design sensitive-data masking first.
Observability
OpenTelemetry GenAI conventions
Learning portable trace fields⭐⭐⭐⭐The spec is still evolving; platform support varies.
HITL / recovery
LangGraph — Interrupts
Approval, checkpoints, resume⭐⭐⭐⭐⭐Production needs a durable checkpointer, not memory alone; side effects must be idempotent.
Harness / sandbox
anthropics/claude-agent-sdk-python
Reading the tool loop, permissions, and subagent implementation⭐⭐⭐⭐⭐Centred on the Claude runtime.
Harness / sandbox
OpenAI — Harness engineering
Environment, feedback loops, and machine rules⭐⭐⭐⭐See how the environment, feedback loop, and machine rules keep an agent stable.
Deploy
bentoml/BentoML
Packaging apps as services and containers⭐⭐⭐⭐A deployment framework does not add evals and guardrails for you.
Multi-agent cases
crewAIInc/crewAI
Understanding role-based division of labor⭐⭐⭐⭐More roles do not guarantee better answers.
Multi-agent cases
stablyai/orca
Running coding agents in parallel worktrees⭐⭐⭐⭐Parallel results still need human review and selection.
Benchmarks
SWE-bench
Learning what benchmark sets look like⭐⭐⭐⭐Read alongside tau2-bench, GAIA, OSWorld, and terminal-bench; scores never replace your own cases.

🛠 Hands-on practice (upstream)

Full exercises & starter code

Summaries from the upstream curriculum; full code, cost, and latency estimates live upstream.

  1. One story for the whole chapter: an AI helper checks three sources, writes a summary, and asks a person before sending — you gradually add its workspace, repeat rhythm, branching routes, quality checks, and a safety brake.
  2. Exercise spine: fixed cases and success rules → step-level traces and cost records → pause-and-ask for risky actions → checkpoint resume with idempotent side effects.
  3. Every exercise starts with a test that needs no API key; set a small budget before calling paid models.
  4. Work records may contain prompts and model replies: never send passwords, personal data, or customer data straight to a tracing platform.
  5. Each extra agent adds model calls, latency, and debugging — stabilize one helper first.

✅ Self-check

  • I can tell outcome from trajectory and check the same case with both.
  • I have fixed eval cases built from real failures, not one beautiful output.
  • I can find the trace, errors, latency, and tokens of one run.
  • High-risk tools have least privilege and human approval; without approval they fail closed.
  • I can resume from a checkpoint and prove the same idempotency key causes no duplicate side effects.
  • I can show an execution receipt and explain when to stop, recover, or roll back.

Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.