A Complete Tutorial for Matt Pocock Skills: The 115K Star AI Coding Engineering Discipline System, a Full Walkthrough of 15 Skills from Installation to Practice
Core proposition: While most AI coding skills teach you "how to write more code faster", Matt Pocock's skill system goes the other way, teaching you how to govern AI with engineering discipline and prevent the codebase from turning into mud. Data: 115,000+ Stars · 10,000+ Forks · 2,700,000+ total installs · 35 skills · 60,000+ newsletter subscribers · MIT License Methodology: installation record → point-by-point dissection of the 15 skills → demonstrations across 8 everyday scenarios → recommended complete workflows
Introduction: "Skills for Real Engineers, Not Vibe Coding"
Matt Pocock is a household name in the TypeScript ecosystem, the founder of Total TypeScript and a technical educator with 200K+ followers on X. His TypeScript courses are regarded by developers worldwide as scripture.
In February 2026, he did something that shook the AI coding world: he open-sourced the skills he uses every day in Claude Code directly to GitHub. Within just 4 months, the repo gained 115,000+ Stars, becoming the most popular skill collection in the skills.sh ecosystem.
His core idea is written directly in the README:
"Developing real applications is hard. Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process. But while doing so, they take away your control."
The design philosophy of these skills is: small, easy to modify, composable. Based on decades of software engineering experience, not on fleeting AI hype.
1. Installation Walkthrough
1.1 One-Command Install
npx skills@latest add mattpocock/skills
The installer guides you through three steps:
Step 1, choose skills: an interactive multi-select interface with 29 skills available. There are 14 core engineering skills (strongly recommended to select all):
| Skill | Purpose |
|---|---|
caveman | Cuts token usage by 75%, keeping only technical content |
diagnose | A scientific six-phase debug loop |
grill-me | Deep requirements interrogation (non-code scenarios) |
grill-with-docs | Requirements interrogation + building a shared-language document |
handoff | Generates an Agent task handoff document |
improve-codebase-architecture | Deep analysis of code architecture |
prototype | A throwaway prototype (a terminal app or multiple UI variants) |
setup-matt-pocock-skills | ⚙️ Initialization config (a prerequisite for the other skills) |
tdd | Red-green-refactor test-driven development |
to-issues | PRD → vertical-slice issues |
to-prd | Conversation context → PRD → GitHub issue |
triage | Issue state-machine classification management |
write-a-skill | Scaffolding for creating new skills |
zoom-out | A system-level bird's-eye view of the code |
Step 2, choose the Agent target: supports 71 Agent platforms, including Claude Code, Codex, Cursor, OpenClaw, Amp, and more.
Step 3, install scope: Project (the current directory's .agents/skills/) or Global (~/.agents/skills/).
1.2 OpenClaw Compatibility Check
After installation, run openclaw skills check, and all 14 skills are successfully recognized and loaded:
✅ caveman
✅ diagnose
✅ grill-me
✅ grill-with-docs
✅ handoff
✅ improve-codebase-architecture
✅ prototype
✅ tdd
✅ to-issues
✅ to-prd
✅ triage
✅ write-a-skill
🔒 setup-matt-pocock-skills (hidden command, available)
🔒 zoom-out (hidden command, available)
2. Design Philosophy: Solving the Four Major Failure Modes of AI Coding
Matt Pocock reduces the common failures of AI coding to four core problems, each mapped to a group of skills:
2.1 Failure Mode A: Misalignment
The Problem: You think the Agent understands your requirements, but what it builds is not at all what you wanted.
This is the most common failure mode. Matt Pocock quotes a famous line from "The Pragmatic Programmer": "No-one knows exactly what they want", and in the AI era, the communication gap only grows larger.
Corresponding skills:
/grill-with-docs: the most popular skill, deeply interrogating your plan/grill-me: the equivalent for non-code scenarios
2.2 Failure Mode B: Verbosity
The Problem: The Agent uses 20 words to say what 1 word could, wasting tokens badly.
Corresponding skills:
- The
CONTEXT.mdestablished in/grill-with-docs(a shared-language document) /caveman: directly cutting 75% of token overhead
Example:
Before: "There's a problem when a lesson inside a section of a course is made 'real' (i.e. given a spot in the file system)"
After: "There's a problem with the materialization cascade"
Nine words replace thirty-three. Precision is not lost.
2.3 Failure Mode C: Broken Code
The Problem: You have agreement on the requirements, but the Agent still writes code that will not run, because there is no feedback loop.
Corresponding skills:
/tdd: the red-green-refactor loop/diagnose: scientific debugging
2.4 Failure Mode D: Software Entropy
The Problem: AI accelerates coding, and it also accelerates software entropy. The codebase becomes complex and hard to modify at an unprecedented rate.
Quoting Kent Beck: "Invest in the design of the system every day."
Corresponding skills:
/improve-codebase-architecture: find deepening opportunities/zoom-out: understand the code at a system level- The module design Q&A in
/to-prd
3. Deep Dive into the Core Skills
3.1 /grill-with-docs: The Most Popular Skill
This is the crown jewel of the entire skill set. It is not just a requirements alignment tool; it is a process of establishing a shared language.
Workflow:
-
Interrogate one by one: one question at a time, waiting for your answer before asking the next
-
Term refinement: when you use vague wording, immediately propose a precise term
-
Code cross-validation: when you describe system behavior, check whether the code is consistent
-
Scenario stress testing: invent edge cases to challenge you
-
Write to documents immediately:
CONTEXT.md: a domain glossary (containing no implementation details)docs/adr/0001-xxx.md: an architecture decision record (only created when three conditions are met)
Field example:
👤 You: /grill-with-docs I want to add an automatic customer service expiry reminder feature to CSMS
🤖 AI (First question): You said "service": do you mean Audit, Company Secretary, or all?
👤 You: All five service categories are needed
🤖 AI (Second question): How many days before expiry should it be sent? You mentioned "30 days": is that working days or calendar days?
👤 You: Calendar days, but configurable
🤖 AI (Third question, after checking code): Your `Service` model already has an `endDate` field,
but no `reminderSent` flag. What happens if the same service triggers multiple reminders?
👤 You: Add a reminderSent flag, each service sends at most once
🤖 AI (Fourth question, writing to CONTEXT.md): I have defined "Service Reminder" as
"an automatic notification event N calendar days before service expiry" and added it to CONTEXT.md.
Now confirm: how to handle notification failure (invalid email)?
3.2 /tdd: The Red-Green-Refactor Loop
Core principle: Tests should verify behavior through public interfaces, not test implementation details.
Good tests vs bad tests:
| Good Test ✅ | Bad Test ❌ |
|---|---|
| "A user can check out with a valid cart" | "calculateTotal() returns the expected value" |
| Tested through API endpoints | Mocks internal collaborators |
| Does not break after refactoring | Breaks as soon as an internal function is renamed |
Horizontal slicing is forbidden:
❌ WRONG (horizontal):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
✅ RIGHT (vertical):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
Field example:
👤 You: /tdd add fuzzy matching to CSMS client search
🤖 AI (Plan): Endpoint: GET /api/clients?search=keyword → Client[]
Test scope: exact match, fuzzy match, no results, empty search, special characters
Confirm the above?
👤 You: OK
🤖 AI (RED): Write the first test: exact match "ABC Company" → returns that client
Test FAILED ❌ (fuzzy matching is not yet implemented in the search endpoint)
🤖 AI (GREEN): Add a contains condition to the Prisma query
Test PASSED ✅
🤖 AI (RED): Write the second test: fuzzy match "abc" → returns "ABC Company"
Test FAILED ❌ (case-sensitive)
🤖 AI (GREEN): Add mode: 'insensitive'
Test PASSED ✅
🤖 AI (REFACTOR): Extract searchQuery builder into a reusable module
All tests still PASSED ✅
3.3 /diagnose: A Scientific Six-Phase Debug
The core mantra: Phase 1 is 90% of the work. If you have a fast, deterministic, Agent-runnable pass/fail signal, the bug is already 90% solved.
The six-phase process:
| Phase | Content | Key Principle |
|---|---|---|
| 1. Build Feedback Loop | Build a reproducible test signal | Spend the most time here |
| 2. Reproduce | Confirm it reproduces reliably | It must be the same bug the user described |
| 3. Hypothesise | Propose 3-5 falsifiable hypotheses | Each must state a prediction |
| 4. Instrument | Change one variable at a time | Each change maps to one hypothesis |
| 5. Fix | Fix it | Verify for regressions |
| 6. Regression Test | Regression test | Add it to the test suite |
10 ways to build a feedback loop (ordered by priority):
- Failing test (unit/integration/E2E)
- Curl / HTTP script (against a running dev server)
- CLI invocation (fixed input vs known-correct output)
- Headless browser script (Playwright/Puppeteer)
- Replay a captured trace
- A throwaway test harness
- Property/fuzz testing loop (1000 random inputs)
- Bisection harness (
git bisect run) - Differential loop (compare output between the old and new versions)
- HITL bash script (a last resort, requiring human clicking)
Field example:
👤 You: /diagnose CSMS dashboard loading exceeds 10 seconds
🤖 AI (Phase 1): Establishing feedback loop...
curl -w "@curl-format.txt" http://localhost:3000/api/dashboard
→ time_total: 12.3s
🤖 AI: Loop established. Now make it faster:
Added --connect-timeout 5, ruling out DNS latency.
→ time_total: 11.8s. Not a DNS issue.
🤖 AI (Phase 2): 5 consecutive runs to confirm reproducibility.
→ 12.1s, 11.9s, 12.4s, 12.0s, 12.2s ✅ Stably reproducible
🤖 AI (Phase 3): Hypothesis ranking:
1. N+1 query: Prisma does not include relations in Dashboard query
2. Large data volume: clients table has over 100,000 rows, no pagination
3. middleware blocking: Supabase Auth validates token on every request
4. Frontend rendering: Chart.js renders large dataset on client
5. External API calls: Dashboard may depend on external services
🤖 AI (Phase 4): Test hypothesis #1 first...
Added eager loading → 12.1s, no improvement. Hypothesis #1 ruled out.
Test hypothesis #2... added LIMIT 20 → 0.8s! Bottleneck is unpaginated large dataset.
3.4 /improve-codebase-architecture: Architecture Deepening
Core concept: Deep Modules
Deep module = small interface, deep implementation Shallow module = interface nearly as complex as implementation
The Deletion Test:
Pretend to delete this module. If the complexity disappears, it is a pass-through (useless). If the complexity reappears across N callers, it deserves to exist.
Field example:
👤 You: /improve-codebase-architecture
🤖 AI: Exploring codebase... (Reading CONTEXT.md + ADRs)
Found 5 candidate refactoring opportunities:
1. [Strong] utils/formatters.ts
Issue: 60 pure functions, but 40 are used only once
Suggestion: merge into their respective consuming modules
2. [Worth exploring] middleware chain
Issue: 7 middleware, but auth and logging are tightly coupled
Suggestion: split into independent adapters
3. [Speculative] API route handlers
Issue: each route repeats the same error handling pattern
Suggestion: extract into a wrapper
📄 HTML report generated: /tmp/architecture-review-20260604.html
3.5 /caveman: The Token Efficiency Revolution
Rules: Cut articles, filler words, and pleasantries. Code blocks stay unchanged. Zero loss of technical accuracy.
Persistence: once triggered, it stays in effect until you say /stop caveman.
Field example:
👤 You: /caveman
👤 You: How to add soft delete to CSMS Prisma schema?
🤖 AI: Add deletedAt DateTime? to model.
Middleware filter where deletedAt is null.
See Prisma docs client extensions.
(Originally might require 200 tokens → now ~25 tokens, 87% reduction)
4. Recommended Complete Workflow
4.1 🚀 Standard Flow for New Feature Development
1️⃣ /grill-with-docs → Align requirements + establish shared language + ADR
2️⃣ /to-prd → Automatically synthesize PRD → GitHub Issue
3️⃣ /to-issues → Split into vertical slice Issues
4️⃣ /tdd → red-green-refactor, implement Issue by Issue
5️⃣ /improve-codebase-architecture → Run once every 2-3 days
4.2 🐛 Debug Flow
1️⃣ /diagnose → Scientific diagnosis (no explanations, no intuition)
2️⃣ /tdd → Use tests to lock in the fix
4.3 🔄 Taking Over Legacy Code Flow
1️⃣ /zoom-out → System-level bird's-eye view
2️⃣ /grill-with-docs → Establish domain language
3️⃣ /improve-codebase-architecture → Identify refactoring opportunities
4.4 📋 Daily Management Flow
Every morning: /triage → Organize new Issues
Weekly: /to-prd → Turn discussion content into formal specifications
Weekly: /improve-codebase-architecture → Architecture health check
When necessary: /handoff → Hand off task to another Agent
5. Comparison with Other AI Coding Methodologies
| Dimension | Matt Pocock Skills | GSD / BMAD | Spec-Kit | Vibe Coding |
|---|---|---|---|---|
| Control | 🟢 Full developer control | 🔴 Process controlled by the framework | 🟡 Shared control | 🔴 Agent-led |
| Composability | 🟢 Independent skills freely combined | 🔴 The overall process cannot be split | 🟡 Partially splittable | 🔴 No structure |
| Learning curve | 🟡 Requires an engineering foundation | 🔴 Requires learning a framework mindset | 🟡 Moderate | 🟢 Zero learning |
| Applicable scale | 🟢 Any scale | 🟡 Medium to large projects | 🟡 Medium to large projects | 🔴 Small prototypes |
| Long-term maintenance | 🟢 Built-in architecture governance | 🟡 Depends on execution | 🟡 Depends on execution | 🔴 Inevitable technical debt |
| Documentation output | 🟢 CONTEXT.md + ADR | 🟡 Framework docs | 🟢 Spec docs | 🔴 Usually none |
6. Application Recommendations for Junze Zhiku
6.1 Immediately Usable Scenarios
- CSMS development:
/grill-with-docs+/tddfor new feature development - CS Management system:
/improve-codebase-architectureto prevent rapid prototypes from turning to mud - MemoryHub:
/diagnosefor troubleshooting performance issues - Agentics website:
/to-prd+/to-issuesto manage content development
6.2 A CONTEXT.md Worth Establishing
Establish a domain glossary for every active project:
| Project | Examples of Core Terms |
|---|---|
| CSMS | Client, Service, Case, Compliance, Renewal |
| MemoryHub | Session, Observation, Consolidation, Dream, Capture |
| Sub2API | Channel, Pricing Rule, Model Pricing, Upstream |
7. Summary
Matt Pocock Skills is the "The Pragmatic Programmer" of the AI coding era. It does not teach you to write code faster with AI, but to ensure code quality through engineering discipline.
Core value:
- Shared Language: replace filler with domain terms, so the Agent truly understands your project
- Feedback Loops: TDD + diagnosis, ensuring the code runs
- Architecture Governance: continuously deepen the code structure and prevent decay
- Process Standardization: PRD → Issues → TDD, a repeatable engineering process
In one sentence: these skills are not meant to replace software engineering, but to ensure that in an era where AI accelerates everything, we are still doing real engineering.
This article is based on mattpocock/skills v1.5.10.
Skill repository: https://github.com/mattpocock/skills
Install command: npx skills@latest add mattpocock/skills
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built