Agentic Research
Agent Learning Roadmap · Stage ★

Capstone

Can you ship a small agent project safely, on your own?

Complete the assigned project and pass the scoring rubric.

3–20 hours (two variants)4 mapped lessonsUpstream edition

📌 Learning goals

  • Turn "I finished the roadmap" into "I have a demonstrable artifact plus a rubric score I gave myself."
  • Track A: assemble a CLI-agent workflow you will reuse, automating one thing you do by hand (3–8 hours).
  • Track B: design, build, and evaluate a small system — either multi-agent (≥2 cooperating agents) or a RAG pipeline (8–20 hours).
  • Self-grade honestly on the four-level rubric (below basic / basic / good / excellent); an honest score beats a high one.

Entry conditions

Track A: Stages 0–2 + A1 + A2 + the Stage 5 Track-A core 5.1–5.4 + A3, all past their self-checks. Track B: Stages 0–8, all past their self-checks. Pick a problem you actually have — at work, in research, or in life; the capstone's value comes from being real.

🧭 Lessons on this site

Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.
Progress here
0/4
Saved in your browser only
  1. 01
    Build Your First AI Agent in Seven Steps: from environment to a minimal production pass

    Adapted from the MIT-licensed awesome-agentic-ai-zh integrated tutorial: one 'paper assistant' project wires up environment setup, your first LLM call, the four-part prompt, tool use, the agent loop with reflection, RAG memory, evals with observability, and the smallest interface. Every step ships minimal runnable code and safety rails; the local Ollama path costs 0 in API fees.

    13 min
  2. 02
    Agent Framework Capability Matrix: Hermes → OpenClaw → Claude Code Full Capability Overview

    One matrix to understand the full capability distribution of the three-layer Agent framework: search API, LLM backend, MCP protocol, skill system, role division, communication channels.

    10 min
  3. 03
    Building Your Own Investment-Research System: The Complete Design from Skill Library to Verification Gates

    Gathers the methods of all the previous articles into one complete blueprint: the five-layer architecture of data, skills, orchestration, verification, and output; the input/output contract design of the skill library; five gates from format validation to low-confidence abstention; why coverage and accuracy must be counted separately; the seven-step method for building an evaluation set; post-launch operations mechanisms; and the four-stage evolution path from a solo tool to a team system.

    24 min
  4. 04
    Build a Working Prototype Website in a Weekend: From Idea to Shareable Link

    To validate a product idea you do not have to wait for engineering capacity. Using AI generation tools like v0, Lovable, and Bolt, this article walks you through converging the idea into a one-page spec, producing the first version, wiring a database and login, deploying to a shareable link, and putting it in front of real users — and spells out which requirements mean you should stop and find people instead.

    10 min

🎯 Curated resources

ResourceWho it's forPriorityWhy
Official definition
CAPSTONE.md(上游完整定義)
Everyone finishing a track⭐⭐⭐⭐⭐The original definitions of topics, requirements, and both rubrics; each stage's pass conditions remain in its own file.
Progress tracking
PROGRESS.md(上游進度表)
Want to track completion stop by stop⭐⭐⭐⭐The upstream progress sheet; this site's roadmap checkboxes also live in your own browser.

🛠 Hands-on practice (upstream)

Full exercises & starter code

Summaries from the upstream curriculum; full code, cost, and latency estimates live upstream.

  1. Non-negotiables (both tracks): a clear input → usable output; rerunnable by someone else (install, configure, run, expected output); at least one failure case handled.
  2. Track A adds: one CLI agent at the core and at least one MCP server or self-written skill/command. Deliverable: repo + README + evidence of one real run + a reflection under 150 words.
  3. Track B adds: tool use, an external interface (CLI/API/chat), a real evaluation (≥5 development + 3 frozen holdout cases with success criteria, grader, trial counts, and baseline), failure-mode analysis, and an architecture sketch.
  4. Swap the scenario by role: researchers build literature QA, developers a CI review agent, teachers quiz helpers, knowledge workers meeting-notes-to-actions, everyday users chore automation.
  5. Show it with the artifact plus your rubric self-grade; describe concrete facts (what you built, what you measured) and skip the "world's best" talk.

✅ Self-check

  • The problem is real, recurring, and I can state its input and expected output clearly.
  • Someone else can rerun it from my README and see the expected output.
  • At least one failure case is handled explicitly — the program does not fail silently.
  • Track B: the eval is versioned, the holdout was never tuned against, and the baseline comparison is written down.
  • The reflection is specific: it names the wrong architecture or component choice and what I would change.

Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.