Agentic Research

A Complete Tutorial for Matt Pocock Skills: The 115K Star AI Coding Engineering Discipline System, a Full Walkthrough of 15 Skills from Installation to Practice

2026/06/0452 min readUltraClaw閱讀中文原文
TopicsAgent SkillsClaude CodeOpenClawAI Coding AgentTypeScript

Core proposition: While most AI coding skills teach you "how to write more code faster", Matt Pocock's skill system goes the other way, teaching you how to govern AI with engineering discipline and prevent the codebase from turning into mud. Data: 115,000+ Stars · 10,000+ Forks · 2,700,000+ total installs · 35 skills · 60,000+ newsletter subscribers · MIT License Methodology: installation record → point-by-point dissection of the 15 skills → demonstrations across 8 everyday scenarios → recommended complete workflows


Introduction: "Skills for Real Engineers, Not Vibe Coding"

Matt Pocock is a household name in the TypeScript ecosystem, the founder of Total TypeScript and a technical educator with 200K+ followers on X. His TypeScript courses are regarded by developers worldwide as scripture.

In February 2026, he did something that shook the AI coding world: he open-sourced the skills he uses every day in Claude Code directly to GitHub. Within just 4 months, the repo gained 115,000+ Stars, becoming the most popular skill collection in the skills.sh ecosystem.

His core idea is written directly in the README:

"Developing real applications is hard. Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process. But while doing so, they take away your control."

The design philosophy of these skills is: small, easy to modify, composable. Based on decades of software engineering experience, not on fleeting AI hype.


1. Installation Walkthrough

1.1 One-Command Install

npx skills@latest add mattpocock/skills

The installer guides you through three steps:

Step 1, choose skills: an interactive multi-select interface with 29 skills available. There are 14 core engineering skills (strongly recommended to select all):

SkillPurpose
cavemanCuts token usage by 75%, keeping only technical content
diagnoseA scientific six-phase debug loop
grill-meDeep requirements interrogation (non-code scenarios)
grill-with-docsRequirements interrogation + building a shared-language document
handoffGenerates an Agent task handoff document
improve-codebase-architectureDeep analysis of code architecture
prototypeA throwaway prototype (a terminal app or multiple UI variants)
setup-matt-pocock-skills⚙️ Initialization config (a prerequisite for the other skills)
tddRed-green-refactor test-driven development
to-issuesPRD → vertical-slice issues
to-prdConversation context → PRD → GitHub issue
triageIssue state-machine classification management
write-a-skillScaffolding for creating new skills
zoom-outA system-level bird's-eye view of the code

Step 2, choose the Agent target: supports 71 Agent platforms, including Claude Code, Codex, Cursor, OpenClaw, Amp, and more.

Step 3, install scope: Project (the current directory's .agents/skills/) or Global (~/.agents/skills/).

1.2 OpenClaw Compatibility Check

After installation, run openclaw skills check, and all 14 skills are successfully recognized and loaded:

✅ caveman
✅ diagnose
✅ grill-me
✅ grill-with-docs
✅ handoff
✅ improve-codebase-architecture
✅ prototype
✅ tdd
✅ to-issues
✅ to-prd
✅ triage
✅ write-a-skill
🔒 setup-matt-pocock-skills  (hidden command, available)
🔒 zoom-out                  (hidden command, available)

2. Design Philosophy: Solving the Four Major Failure Modes of AI Coding

Matt Pocock reduces the common failures of AI coding to four core problems, each mapped to a group of skills:

2.1 Failure Mode A: Misalignment

The Problem: You think the Agent understands your requirements, but what it builds is not at all what you wanted.

This is the most common failure mode. Matt Pocock quotes a famous line from "The Pragmatic Programmer": "No-one knows exactly what they want", and in the AI era, the communication gap only grows larger.

Corresponding skills:

  • /grill-with-docs: the most popular skill, deeply interrogating your plan
  • /grill-me: the equivalent for non-code scenarios

2.2 Failure Mode B: Verbosity

The Problem: The Agent uses 20 words to say what 1 word could, wasting tokens badly.

Corresponding skills:

  • The CONTEXT.md established in /grill-with-docs (a shared-language document)
  • /caveman: directly cutting 75% of token overhead

Example:

Before: "There's a problem when a lesson inside a section of a course is made 'real' (i.e. given a spot in the file system)"

After: "There's a problem with the materialization cascade"

Nine words replace thirty-three. Precision is not lost.

2.3 Failure Mode C: Broken Code

The Problem: You have agreement on the requirements, but the Agent still writes code that will not run, because there is no feedback loop.

Corresponding skills:

  • /tdd: the red-green-refactor loop
  • /diagnose: scientific debugging

2.4 Failure Mode D: Software Entropy

The Problem: AI accelerates coding, and it also accelerates software entropy. The codebase becomes complex and hard to modify at an unprecedented rate.

Quoting Kent Beck: "Invest in the design of the system every day."

Corresponding skills:

  • /improve-codebase-architecture: find deepening opportunities
  • /zoom-out: understand the code at a system level
  • The module design Q&A in /to-prd

3. Deep Dive into the Core Skills

3.1 /grill-with-docs: The Most Popular Skill

This is the crown jewel of the entire skill set. It is not just a requirements alignment tool; it is a process of establishing a shared language.

Workflow:

  1. Interrogate one by one: one question at a time, waiting for your answer before asking the next

  2. Term refinement: when you use vague wording, immediately propose a precise term

  3. Code cross-validation: when you describe system behavior, check whether the code is consistent

  4. Scenario stress testing: invent edge cases to challenge you

  5. Write to documents immediately:

    • CONTEXT.md: a domain glossary (containing no implementation details)
    • docs/adr/0001-xxx.md: an architecture decision record (only created when three conditions are met)

Field example:

👤 You: /grill-with-docs I want to add an automatic customer service expiry reminder feature to CSMS

🤖 AI (First question): You said "service": do you mean Audit, Company Secretary, or all?
👤 You: All five service categories are needed

🤖 AI (Second question): How many days before expiry should it be sent? You mentioned "30 days": is that working days or calendar days?
👤 You: Calendar days, but configurable

🤖 AI (Third question, after checking code): Your `Service` model already has an `endDate` field,
but no `reminderSent` flag. What happens if the same service triggers multiple reminders?
👤 You: Add a reminderSent flag, each service sends at most once

🤖 AI (Fourth question, writing to CONTEXT.md): I have defined "Service Reminder" as
"an automatic notification event N calendar days before service expiry" and added it to CONTEXT.md.
Now confirm: how to handle notification failure (invalid email)?

3.2 /tdd: The Red-Green-Refactor Loop

Core principle: Tests should verify behavior through public interfaces, not test implementation details.

Good tests vs bad tests:

Good Test ✅Bad Test ❌
"A user can check out with a valid cart""calculateTotal() returns the expected value"
Tested through API endpointsMocks internal collaborators
Does not break after refactoringBreaks as soon as an internal function is renamed

Horizontal slicing is forbidden:

❌ WRONG (horizontal):
  RED:   test1, test2, test3, test4, test5
  GREEN: impl1, impl2, impl3, impl4, impl5

✅ RIGHT (vertical):
  RED→GREEN: test1→impl1
  RED→GREEN: test2→impl2
  RED→GREEN: test3→impl3
  ...

Field example:

👤 You: /tdd add fuzzy matching to CSMS client search

🤖 AI (Plan): Endpoint: GET /api/clients?search=keyword → Client[]
Test scope: exact match, fuzzy match, no results, empty search, special characters
Confirm the above?

👤 You: OK

🤖 AI (RED): Write the first test: exact match "ABC Company" → returns that client
Test FAILED ❌ (fuzzy matching is not yet implemented in the search endpoint)

🤖 AI (GREEN): Add a contains condition to the Prisma query
Test PASSED ✅

🤖 AI (RED): Write the second test: fuzzy match "abc" → returns "ABC Company"
Test FAILED ❌ (case-sensitive)

🤖 AI (GREEN): Add mode: 'insensitive'
Test PASSED ✅

🤖 AI (REFACTOR): Extract searchQuery builder into a reusable module
All tests still PASSED ✅

3.3 /diagnose: A Scientific Six-Phase Debug

The core mantra: Phase 1 is 90% of the work. If you have a fast, deterministic, Agent-runnable pass/fail signal, the bug is already 90% solved.

The six-phase process:

PhaseContentKey Principle
1. Build Feedback LoopBuild a reproducible test signalSpend the most time here
2. ReproduceConfirm it reproduces reliablyIt must be the same bug the user described
3. HypothesisePropose 3-5 falsifiable hypothesesEach must state a prediction
4. InstrumentChange one variable at a timeEach change maps to one hypothesis
5. FixFix itVerify for regressions
6. Regression TestRegression testAdd it to the test suite

10 ways to build a feedback loop (ordered by priority):

  1. Failing test (unit/integration/E2E)
  2. Curl / HTTP script (against a running dev server)
  3. CLI invocation (fixed input vs known-correct output)
  4. Headless browser script (Playwright/Puppeteer)
  5. Replay a captured trace
  6. A throwaway test harness
  7. Property/fuzz testing loop (1000 random inputs)
  8. Bisection harness (git bisect run)
  9. Differential loop (compare output between the old and new versions)
  10. HITL bash script (a last resort, requiring human clicking)

Field example:

👤 You: /diagnose CSMS dashboard loading exceeds 10 seconds

🤖 AI (Phase 1): Establishing feedback loop...
curl -w "@curl-format.txt" http://localhost:3000/api/dashboard
→ time_total: 12.3s

🤖 AI: Loop established. Now make it faster:
Added --connect-timeout 5, ruling out DNS latency.
→ time_total: 11.8s. Not a DNS issue.

🤖 AI (Phase 2): 5 consecutive runs to confirm reproducibility.
→ 12.1s, 11.9s, 12.4s, 12.0s, 12.2s ✅ Stably reproducible

🤖 AI (Phase 3): Hypothesis ranking:
1. N+1 query: Prisma does not include relations in Dashboard query
2. Large data volume: clients table has over 100,000 rows, no pagination
3. middleware blocking: Supabase Auth validates token on every request
4. Frontend rendering: Chart.js renders large dataset on client
5. External API calls: Dashboard may depend on external services

🤖 AI (Phase 4): Test hypothesis #1 first...
Added eager loading → 12.1s, no improvement. Hypothesis #1 ruled out.
Test hypothesis #2... added LIMIT 20 → 0.8s! Bottleneck is unpaginated large dataset.

3.4 /improve-codebase-architecture: Architecture Deepening

Core concept: Deep Modules

Deep module = small interface, deep implementation Shallow module = interface nearly as complex as implementation

The Deletion Test:

Pretend to delete this module. If the complexity disappears, it is a pass-through (useless). If the complexity reappears across N callers, it deserves to exist.

Field example:

👤 You: /improve-codebase-architecture

🤖 AI: Exploring codebase... (Reading CONTEXT.md + ADRs)

Found 5 candidate refactoring opportunities:

1. [Strong] utils/formatters.ts
   Issue: 60 pure functions, but 40 are used only once
   Suggestion: merge into their respective consuming modules

2. [Worth exploring] middleware chain
   Issue: 7 middleware, but auth and logging are tightly coupled
   Suggestion: split into independent adapters

3. [Speculative] API route handlers
   Issue: each route repeats the same error handling pattern
   Suggestion: extract into a wrapper

📄 HTML report generated: /tmp/architecture-review-20260604.html

3.5 /caveman: The Token Efficiency Revolution

Rules: Cut articles, filler words, and pleasantries. Code blocks stay unchanged. Zero loss of technical accuracy.

Persistence: once triggered, it stays in effect until you say /stop caveman.

Field example:

👤 You: /caveman
👤 You: How to add soft delete to CSMS Prisma schema?

🤖 AI: Add deletedAt DateTime? to model.
Middleware filter where deletedAt is null.
See Prisma docs client extensions.

(Originally might require 200 tokens → now ~25 tokens, 87% reduction)

4. Recommended Complete Workflow

4.1 🚀 Standard Flow for New Feature Development

1️⃣  /grill-with-docs  →  Align requirements + establish shared language + ADR
2️⃣  /to-prd            →  Automatically synthesize PRD → GitHub Issue
3️⃣  /to-issues         →  Split into vertical slice Issues
4️⃣  /tdd               →  red-green-refactor, implement Issue by Issue
5️⃣  /improve-codebase-architecture → Run once every 2-3 days

4.2 🐛 Debug Flow

1️⃣  /diagnose          →  Scientific diagnosis (no explanations, no intuition)
2️⃣  /tdd               →  Use tests to lock in the fix

4.3 🔄 Taking Over Legacy Code Flow

1️⃣  /zoom-out          →  System-level bird's-eye view
2️⃣  /grill-with-docs   →  Establish domain language
3️⃣  /improve-codebase-architecture → Identify refactoring opportunities

4.4 📋 Daily Management Flow

Every morning:  /triage          →  Organize new Issues
Weekly:  /to-prd          →  Turn discussion content into formal specifications
Weekly:  /improve-codebase-architecture → Architecture health check
When necessary: /handoff         →  Hand off task to another Agent

5. Comparison with Other AI Coding Methodologies

DimensionMatt Pocock SkillsGSD / BMADSpec-KitVibe Coding
Control🟢 Full developer control🔴 Process controlled by the framework🟡 Shared control🔴 Agent-led
Composability🟢 Independent skills freely combined🔴 The overall process cannot be split🟡 Partially splittable🔴 No structure
Learning curve🟡 Requires an engineering foundation🔴 Requires learning a framework mindset🟡 Moderate🟢 Zero learning
Applicable scale🟢 Any scale🟡 Medium to large projects🟡 Medium to large projects🔴 Small prototypes
Long-term maintenance🟢 Built-in architecture governance🟡 Depends on execution🟡 Depends on execution🔴 Inevitable technical debt
Documentation output🟢 CONTEXT.md + ADR🟡 Framework docs🟢 Spec docs🔴 Usually none

6. Application Recommendations for Junze Zhiku

6.1 Immediately Usable Scenarios

  • CSMS development: /grill-with-docs + /tdd for new feature development
  • CS Management system: /improve-codebase-architecture to prevent rapid prototypes from turning to mud
  • MemoryHub: /diagnose for troubleshooting performance issues
  • Agentics website: /to-prd + /to-issues to manage content development

6.2 A CONTEXT.md Worth Establishing

Establish a domain glossary for every active project:

ProjectExamples of Core Terms
CSMSClient, Service, Case, Compliance, Renewal
MemoryHubSession, Observation, Consolidation, Dream, Capture
Sub2APIChannel, Pricing Rule, Model Pricing, Upstream

7. Summary

Matt Pocock Skills is the "The Pragmatic Programmer" of the AI coding era. It does not teach you to write code faster with AI, but to ensure code quality through engineering discipline.

Core value:

  1. Shared Language: replace filler with domain terms, so the Agent truly understands your project
  2. Feedback Loops: TDD + diagnosis, ensuring the code runs
  3. Architecture Governance: continuously deepen the code structure and prevent decay
  4. Process Standardization: PRD → Issues → TDD, a repeatable engineering process

In one sentence: these skills are not meant to replace software engineering, but to ensure that in an era where AI accelerates everything, we are still doing real engineering.


This article is based on mattpocock/skills v1.5.10. Skill repository: https://github.com/mattpocock/skills Install command: npx skills@latest add mattpocock/skills