Agentic Research
Agent Learning Roadmap · Stage 7.5

Advanced Agentic Choices

Which advanced patterns are worth knowing?

Pick from 12 concepts — PAR loop, agent-as-judge, and more — as needed.

6–10 hours (selective)7 mapped lessonsUpstream edition

📌 Learning goals

  • Explain four advanced concepts in plain words and the failure each one handles.
  • Pick one candidate technique from a short table instead of installing every pattern at once.
  • Decide whether to keep the complexity using a baseline, the same eval set, cost, and safety results.
  • Make Keep/Simplify/Remove calls on harness components while keeping a way back.

Entry conditions

Complete Stage 7 with a rerunnable eval set, stop conditions, and a single-agent baseline. No coding or API purchase required. The default answer is "don't add yet": measure the baseline first, and keep a new technique only if it brings reproducible improvement on the same evals.

🧭 Lessons on this site

Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.
Progress here
0/7
Saved in your browser only
  1. 01
    Pre-mortem: Why AI Agents Need to Imagine Their Own Failure Before They Start

    The application of the Pre-mortem methodology in AI Agent systems. Explore how Agent Previsor uses multi-scenario divergent path forecasting to move the regret of "I should not have done it that way" from after execution to before execution, based on chess move calculation, military war-gaming, and a complete analysis across four forecasting dimensions.

    19 min
  2. 02
    Three Departments and Six Ministries vs Loop Engineering: A Technical Dissection of Institutional Process Enforcement

    A comparison of how the two GitHub 'Three Departments and Six Ministries' systems achieve process enforcement using a State Machine, Permission Matrix, Review Gate, and 4-layer Gateway, plus the implications for refactoring our Loop Engineering system.

    19 min
  3. 03
    Subagent Isolation Architecture: Reliability Lessons for AI Financial Applications, From the our HK research pipeline Data Contamination Incident to the Clean Context Design Pattern, A Complete Journey

    While analyzing Stock A and Stock B back to back, Stock A's property-fund holding (carrying value HK$368 million, impairment HK$154 million — both teaching assumptions) leaked into the Stock B report. Stock B's business is consumer-goods trading and money lending; it holds no such fund at all. This is not a hallucination — it is a systemic context-contamination problem.

    24 min
  4. 04
    Infrastructure-izing Context Compression: How Headroom Turns Token Cost from a Tactical Problem into System Architecture

    Headroom 25.8K⭐ · A new species of Agent infrastructure, not a cost-saving tool but an infrastructure layer that makes long-running AI Agents economically viable. A 6-layer compression pipeline + the reversible CCR design + a 16x academic breakthrough validating that the route is correct.

    5 min
  5. 05
    The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice

    TypeSafe AI emerges from stealth with a $40M seed and releases Jev, a System One model that generates no text and only answers typed questions (choice / score / noul). It answers multiple questions in parallel in a single forward pass, returns calibrated probabilities, with ~100ms latency, priced at $0.042/MTok.

    28 min
  6. 06
    200+ Skills, One Router: Engineering Practice of Agent Skill Routing

    After installing 200+ skills, does the AI Agent actually get dumber? How category-by-stage matrix routing raises skill discovery from 35% to 90%, reduces incorrect tool usage by 80%, and lets any task automatically match the right skill combination.

    20 min
  7. 07
    Skills Triggering Deep Dive: Why Your AI Agent Has 200 Skills but Can't Use Even One

    A comprehensive audit of 242 skills reveals: 95% of open source skill descriptions are English only, and non-English trigger success rate is just 20%. How a three-layer keyword strategy raised skill discovery rate from 35% to 90% by changing just one line.

    23 min

📚 Required reading

  1. 1.Anthropic — Building Effective Agents⭐⭐⭐⭐⭐Start with the simplest workflow; add agents only when division of labor proves valuable.
  2. 2.Anthropic — Demystifying Evals for AI Agents⭐⭐⭐⭐⭐Turn failures into rerunnable cases, then compare outcome, trajectory, cost, and stability.
  3. 3.OpenAI — Harness Engineering⭐⭐⭐⭐See how a real codebase uses clear boundaries, docs, and mechanical gates to keep agents stable; a case study, not the only architecture.

🎯 Curated resources

ResourceWho it's forPriorityWhy
Papers
Reflexion(arXiv 2303.11366)
Want the source of self-feedback⭐⭐⭐⭐The original paper on writing failures, feedback, and next strategies into reusable records.
Papers
Constitutional AI(arXiv 2212.08073)
Studying principle-based critique⭐⭐⭐⭐Training method for principle-based critique and revision; not the same thing as an eval.
Papers
More Agents Is All You Need(arXiv 2303.17760)
Testing "more agents is always better"⭐⭐⭐A specific benchmark experiment; do not extrapolate exact magnitudes.
Engineering posts
Anthropic — Effective context engineering
Systems whose context keeps overflowing⭐⭐⭐⭐⭐Anthropic's systematic take on context engineering.
Engineering posts
Anthropic — Multi-agent research system
Building parallel research systems⭐⭐⭐⭐Multi-agent suits breadth-first, parallelizable research; that system used roughly 15× the tokens of an average chat — not a universal rule.
Engineering posts
The Bitter Lesson(Rich Sutton)
The long-run tension between search and scale⭐⭐⭐⭐The historical lesson on why hand-built complexity often loses to simple methods plus scale.
Benchmarks
SWE-bench/SWE-bench
Starting from a small reproducible failure set⭐⭐⭐⭐Teaches observable actions without demanding private CoT; a research setup is not every production loop.
Benchmarks
sierra-research/tau2-bench
Evaluating tool use and policy compliance⭐⭐⭐⭐A dual-track benchmark for tool calls plus user interaction.
Chinese / hands-on
datawhalechina/hello-agents
Want a complete Chinese agent textbook⭐⭐⭐⭐⭐A complete Chinese agent textbook; long — pick chapters by the concepts here.
Chinese / hands-on
李宏毅生成式 AI 課程
Chinese course with research background⭐⭐⭐⭐⭐Pick topics by year; check official docs for product surfaces.
Chinese / hands-on
microsoft/ai-agents-for-beginners
An 18-lesson intro in many languages⭐⭐⭐⭐More vendor examples; learn the concepts before choosing an SDK.

🛠 Hands-on practice (upstream)

Full exercises & starter code

Summaries from the upstream curriculum; full code, cost, and latency estimates live upstream.

  1. No new code required: the main exercise is a decision drill — baseline → pick one candidate → compare on the same evals → Keep/Simplify/Remove.
  2. Four advanced concepts: evaluator-optimizer (agent-as-judge), failure injection (chaos evals), autonomy gradients (trust layers), and model-harness fit.
  3. Judges make mistakes too and cannot decide truth alone; inject failures only in isolated environments; high-risk actions never escalate their own permissions.
  4. Reading path: the three required pieces first, then papers and cases picked by the failure type you actually hit.

✅ Self-check

  • I can explain the four advanced concepts in plain words and name the boundaries of three community terms.
  • I add one pattern from the table only after having a single-agent baseline and rerunnable failures.
  • I compare quality, safety, cost, and latency on the same eval set, not one beautiful answer.
  • I can make a Keep/Simplify/Remove call on a harness component and point to the evidence and the way back.
  • I know judges need checking too, failure injection stays in isolated environments, and high-risk actions never escalate their own permissions.

Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.