Advanced Agentic Choices
Which advanced patterns are worth knowing?
Pick from 12 concepts — PAR loop, agent-as-judge, and more — as needed.
📌 Learning goals
- Explain four advanced concepts in plain words and the failure each one handles.
- Pick one candidate technique from a short table instead of installing every pattern at once.
- Decide whether to keep the complexity using a baseline, the same eval set, cost, and safety results.
- Make Keep/Simplify/Remove calls on harness components while keeping a way back.
Entry conditions
Complete Stage 7 with a rerunnable eval set, stop conditions, and a single-agent baseline. No coding or API purchase required. The default answer is "don't add yet": measure the baseline first, and keep a new technique only if it brings reproducible improvement on the same evals.
🧭 Lessons on this site
Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.- 01Pre-mortem: Why AI Agents Need to Imagine Their Own Failure Before They Start
The application of the Pre-mortem methodology in AI Agent systems. Explore how Agent Previsor uses multi-scenario divergent path forecasting to move the regret of "I should not have done it that way" from after execution to before execution, based on chess move calculation, military war-gaming, and a complete analysis across four forecasting dimensions.
19 min - 02Three Departments and Six Ministries vs Loop Engineering: A Technical Dissection of Institutional Process Enforcement
A comparison of how the two GitHub 'Three Departments and Six Ministries' systems achieve process enforcement using a State Machine, Permission Matrix, Review Gate, and 4-layer Gateway, plus the implications for refactoring our Loop Engineering system.
19 min - 03Subagent Isolation Architecture: Reliability Lessons for AI Financial Applications, From the our HK research pipeline Data Contamination Incident to the Clean Context Design Pattern, A Complete Journey
While analyzing Stock A and Stock B back to back, Stock A's property-fund holding (carrying value HK$368 million, impairment HK$154 million — both teaching assumptions) leaked into the Stock B report. Stock B's business is consumer-goods trading and money lending; it holds no such fund at all. This is not a hallucination — it is a systemic context-contamination problem.
24 min - 04Infrastructure-izing Context Compression: How Headroom Turns Token Cost from a Tactical Problem into System Architecture
Headroom 25.8K⭐ · A new species of Agent infrastructure, not a cost-saving tool but an infrastructure layer that makes long-running AI Agents economically viable. A 6-layer compression pipeline + the reversible CCR design + a 16x academic breakthrough validating that the route is correct.
5 min - 05The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
TypeSafe AI emerges from stealth with a $40M seed and releases Jev, a System One model that generates no text and only answers typed questions (choice / score / noul). It answers multiple questions in parallel in a single forward pass, returns calibrated probabilities, with ~100ms latency, priced at $0.042/MTok.
28 min - 06200+ Skills, One Router: Engineering Practice of Agent Skill Routing
After installing 200+ skills, does the AI Agent actually get dumber? How category-by-stage matrix routing raises skill discovery from 35% to 90%, reduces incorrect tool usage by 80%, and lets any task automatically match the right skill combination.
20 min - 07Skills Triggering Deep Dive: Why Your AI Agent Has 200 Skills but Can't Use Even One
A comprehensive audit of 242 skills reveals: 95% of open source skill descriptions are English only, and non-English trigger success rate is just 20%. How a three-layer keyword strategy raised skill discovery rate from 35% to 90% by changing just one line.
23 min
📚 Required reading
- 1.Anthropic — Building Effective Agents⭐⭐⭐⭐⭐Start with the simplest workflow; add agents only when division of labor proves valuable.
- 2.Anthropic — Demystifying Evals for AI Agents⭐⭐⭐⭐⭐Turn failures into rerunnable cases, then compare outcome, trajectory, cost, and stability.
- 3.OpenAI — Harness Engineering⭐⭐⭐⭐See how a real codebase uses clear boundaries, docs, and mechanical gates to keep agents stable; a case study, not the only architecture.
🎯 Curated resources
| Resource | Who it's for | Priority | Why |
|---|---|---|---|
Papers Reflexion(arXiv 2303.11366) | Want the source of self-feedback | ⭐⭐⭐⭐ | The original paper on writing failures, feedback, and next strategies into reusable records. |
Papers Constitutional AI(arXiv 2212.08073) | Studying principle-based critique | ⭐⭐⭐⭐ | Training method for principle-based critique and revision; not the same thing as an eval. |
Papers More Agents Is All You Need(arXiv 2303.17760) | Testing "more agents is always better" | ⭐⭐⭐ | A specific benchmark experiment; do not extrapolate exact magnitudes. |
Engineering posts Anthropic — Effective context engineering | Systems whose context keeps overflowing | ⭐⭐⭐⭐⭐ | Anthropic's systematic take on context engineering. |
Engineering posts Anthropic — Multi-agent research system | Building parallel research systems | ⭐⭐⭐⭐ | Multi-agent suits breadth-first, parallelizable research; that system used roughly 15× the tokens of an average chat — not a universal rule. |
Engineering posts The Bitter Lesson(Rich Sutton) | The long-run tension between search and scale | ⭐⭐⭐⭐ | The historical lesson on why hand-built complexity often loses to simple methods plus scale. |
Benchmarks SWE-bench/SWE-bench | Starting from a small reproducible failure set | ⭐⭐⭐⭐ | Teaches observable actions without demanding private CoT; a research setup is not every production loop. |
Benchmarks sierra-research/tau2-bench | Evaluating tool use and policy compliance | ⭐⭐⭐⭐ | A dual-track benchmark for tool calls plus user interaction. |
Chinese / hands-on datawhalechina/hello-agents | Want a complete Chinese agent textbook | ⭐⭐⭐⭐⭐ | A complete Chinese agent textbook; long — pick chapters by the concepts here. |
Chinese / hands-on 李宏毅生成式 AI 課程 | Chinese course with research background | ⭐⭐⭐⭐⭐ | Pick topics by year; check official docs for product surfaces. |
Chinese / hands-on microsoft/ai-agents-for-beginners | An 18-lesson intro in many languages | ⭐⭐⭐⭐ | More vendor examples; learn the concepts before choosing an SDK. |
🛠 Hands-on practice (upstream)
Full exercises & starter codeSummaries from the upstream curriculum; full code, cost, and latency estimates live upstream.
- No new code required: the main exercise is a decision drill — baseline → pick one candidate → compare on the same evals → Keep/Simplify/Remove.
- Four advanced concepts: evaluator-optimizer (agent-as-judge), failure injection (chaos evals), autonomy gradients (trust layers), and model-harness fit.
- Judges make mistakes too and cannot decide truth alone; inject failures only in isolated environments; high-risk actions never escalate their own permissions.
- Reading path: the three required pieces first, then papers and cases picked by the failure type you actually hit.
✅ Self-check
- I can explain the four advanced concepts in plain words and name the boundaries of three community terms.
- I add one pattern from the table only after having a single-agent baseline and rerunnable failures.
- I compare quality, safety, cost, and latency on the same eval set, not one beautiful answer.
- I can make a Keep/Simplify/Remove call on a harness component and point to the evidence and the way back.
- I know judges need checking too, failure injection stays in isolated environments, and high-risk actions never escalate their own permissions.
Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.