Capstone
Can you ship a small agent project safely, on your own?
Complete the assigned project and pass the scoring rubric.
📌 Learning goals
- Turn "I finished the roadmap" into "I have a demonstrable artifact plus a rubric score I gave myself."
- Track A: assemble a CLI-agent workflow you will reuse, automating one thing you do by hand (3–8 hours).
- Track B: design, build, and evaluate a small system — either multi-agent (≥2 cooperating agents) or a RAG pipeline (8–20 hours).
- Self-grade honestly on the four-level rubric (below basic / basic / good / excellent); an honest score beats a high one.
Entry conditions
Track A: Stages 0–2 + A1 + A2 + the Stage 5 Track-A core 5.1–5.4 + A3, all past their self-checks. Track B: Stages 0–8, all past their self-checks. Pick a problem you actually have — at work, in research, or in life; the capstone's value comes from being real.
🧭 Lessons on this site
Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.- 01Build Your First AI Agent in Seven Steps: from environment to a minimal production pass
Adapted from the MIT-licensed awesome-agentic-ai-zh integrated tutorial: one 'paper assistant' project wires up environment setup, your first LLM call, the four-part prompt, tool use, the agent loop with reflection, RAG memory, evals with observability, and the smallest interface. Every step ships minimal runnable code and safety rails; the local Ollama path costs 0 in API fees.
13 min - 02Agent Framework Capability Matrix: Hermes → OpenClaw → Claude Code Full Capability Overview
One matrix to understand the full capability distribution of the three-layer Agent framework: search API, LLM backend, MCP protocol, skill system, role division, communication channels.
10 min - 03Building Your Own Investment-Research System: The Complete Design from Skill Library to Verification Gates
Gathers the methods of all the previous articles into one complete blueprint: the five-layer architecture of data, skills, orchestration, verification, and output; the input/output contract design of the skill library; five gates from format validation to low-confidence abstention; why coverage and accuracy must be counted separately; the seven-step method for building an evaluation set; post-launch operations mechanisms; and the four-stage evolution path from a solo tool to a team system.
24 min - 04Build a Working Prototype Website in a Weekend: From Idea to Shareable Link
To validate a product idea you do not have to wait for engineering capacity. Using AI generation tools like v0, Lovable, and Bolt, this article walks you through converging the idea into a one-page spec, producing the first version, wiring a database and login, deploying to a shareable link, and putting it in front of real users — and spells out which requirements mean you should stop and find people instead.
10 min
🎯 Curated resources
| Resource | Who it's for | Priority | Why |
|---|---|---|---|
Official definition CAPSTONE.md(上游完整定義) | Everyone finishing a track | ⭐⭐⭐⭐⭐ | The original definitions of topics, requirements, and both rubrics; each stage's pass conditions remain in its own file. |
Progress tracking PROGRESS.md(上游進度表) | Want to track completion stop by stop | ⭐⭐⭐⭐ | The upstream progress sheet; this site's roadmap checkboxes also live in your own browser. |
🛠 Hands-on practice (upstream)
Full exercises & starter codeSummaries from the upstream curriculum; full code, cost, and latency estimates live upstream.
- Non-negotiables (both tracks): a clear input → usable output; rerunnable by someone else (install, configure, run, expected output); at least one failure case handled.
- Track A adds: one CLI agent at the core and at least one MCP server or self-written skill/command. Deliverable: repo + README + evidence of one real run + a reflection under 150 words.
- Track B adds: tool use, an external interface (CLI/API/chat), a real evaluation (≥5 development + 3 frozen holdout cases with success criteria, grader, trial counts, and baseline), failure-mode analysis, and an architecture sketch.
- Swap the scenario by role: researchers build literature QA, developers a CI review agent, teachers quiz helpers, knowledge workers meeting-notes-to-actions, everyday users chore automation.
- Show it with the artifact plus your rubric self-grade; describe concrete facts (what you built, what you measured) and skip the "world's best" talk.
✅ Self-check
- The problem is real, recurring, and I can state its input and expected output clearly.
- Someone else can rerun it from my README and see the expected output.
- At least one failure case is handled explicitly — the program does not fail silently.
- Track B: the eval is versioned, the holdout was never tuned against, and the baseline comparison is written down.
- The reflection is specific: it names the wrong architecture or component choice and what I would change.
Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.