Agentic Research

Skill Curator: When Your Agent Has 125 Skills but Only 64 Survived

2026/06/1079 min readUltraClaw閱讀中文原文
TopicsOpenClawUltraClawAgentic Infrastructure

Premise: You installed 125 skills for your Agent. You think it is powerful. In reality, it can correctly invoke only 64 of them. Of the other 61, 6 can never be triggered, and 55 are basically ineffective for Chinese users.
Solution: Skill Curator, a six-stage automated curation engine. Scan → Diagnose → Adapt → Scenario → Report → Execute.
Result: 125 skills improved from 51% health to 99.2%. 6 critical → 1 (ghost skill), 55 warnings → 0, 64 healthy → 124.


🎭 Chapter 1: The Library Paradox

Imagine you walk into a library. The shelves are packed with books, 125 of them. But:

  • 60% of the books have no Chinese catalog, so you cannot find them when searching in Chinese.
  • 6 books do not even have covers, so you have no idea what is inside.
  • 15 books overlap completely in content, but you do not know which one to discard.
  • No one tells you how to use each book; you stand at the entrance, at a loss.

This is the reality every AI Agent faces.

We call this phenomenon the "Library Paradox": More skills ≠ greater capability. Without curation, the number of skills is inversely proportional to usability.

The Root of the Problem

The Agent ecosystem has three structural problems:

  1. Skills are downloaded, not adapted. In open-source skills, 95% of the description field is pure English. A Chinese user says, in Chinese, “help me build a website” → the description contains only "frontend design" and "UI engineering" → keyword matching fails → the skill is effectively useless.

  2. Formatting errors are inherent. Community skills are copied from one another, with misaligned frontmatter YAML, missing required fields, and encoding issues. A blank description can prevent a skill from ever being triggered.

  3. No one actively maintains skills. Users download skills once and never maintain them. You need an automated "skill curator."


🔬 Chapter 2: Audit, Blood Test Report for 125 Skills

In May 2026, we performed the first comprehensive health scan of UltraClaw's 125 skills. Here are the blood test results:

Health Distribution

6
🔴 Fatal

No description or corrupted format, permanently cannot be triggered. Equivalent to "this book has no cover".

55
🟡 Warning

Missing Chinese keywords, description is English only. Chinese users have a trigger success rate below 20%.

64
🟢 Healthy

Complete format + bilingual Chinese and English keywords. But there are still functional overlaps and untested risks.

Core Findings

MetricDataSeverity
Overall health rate51.2%Half of the skills are subhealthy or dead
Non-English trigger success rate~20%Only 1 in 5 Chinese commands can find the right skill
description language coverage1.2 languages/skill95% of skills have only English description
Functionally overlapping skill groups7 groups (15 skills)The same function has 2-3 overlapping skills that interfere with each other

Conclusion: You think your Agent is strong. In reality, one out of every two skills is slacking off.


🧬 Chapter 3: Six-Phase Lifecycle

The core design philosophy of Skill Curator is full lifecycle management. It is not a one-time fix, but ongoing health maintenance.

{["1️⃣ Scan", "2️⃣ Diagnose", "3️⃣ Adapt", "4️⃣ Scenarios", "5️⃣ Report", "6️⃣ Execute"].map((phase, i) => (
{phase}
))}

Phase 1: Scan, Full Health Inventory

The scanner traverses all SKILL.md files under the skills/ directory and extracts:

  • Frontmatter completeness: Whether the required fields name and description are present
  • Language coverage: How many languages are included in description
  • Keyword matrix: Which trigger words exist for each language
  • File metadata: Size, modification time, and license type
# Scan report example snippet
{
  "skill": "frontend-design",
  "health": "🟡 warning",
  "issues": [
    "description contains only English (1/6 language)",
    "Missing keywords: website, frontend, website building, webpage"
  ],
  "recommended_action": "inject_keywords"
}

Scan results automatically generate a JSON report for consumption by downstream phases.


Phase 2: Diagnose, Three-Tier Classification

This is the most valuable design in Skill Curator: not all problems are equally important.

🔴 Critical

The skill cannot be triggered at all. It requires immediate repair.

  • The description field is missing or empty.
  • Frontmatter YAML syntax errors cause parsing to fail.
  • The SKILL.md file is corrupted or empty.

The first scan found 6 critical skills. Of these, 4 were due to abnormal frontmatter formatting in community skills (incorrect indentation, use of non-standard fields causing the parser to skip required fields).

🟡 Warning

The skill exists but is extremely inefficient. It is nearly unusable for non-English users.

  • description contains only English (trigger words cover 1/6 of languages).
  • Missing keywords in Traditional Chinese, Simplified Chinese, Japanese, Korean, and Arabic.
  • Frontmatter is complete, but description is too short (< 50 characters).

The first scan found 55 warning skills. This is the largest problem set, accounting for 44% of the total. Each is "theoretically usable," but for a boss who uses Traditional Chinese, the trigger rate is below 20%.

🟢 Suggestion

The skill functions normally but has room for optimization. It requires user confirmation before execution.

  • Functional overlap: multiple skills implement the same or similar functionality (such as 3 different "frontend design" skills).
  • Never used: the skill has been installed for more than 30 days but has zero invocation records.
  • Token optimization: the skill's SKILL.md is too large (> 20KB), consuming a large amount of context on every load.

The first scan found 7 groups of functional overlap (involving 15 skills), as well as 12 "ghost skills" (installed but never invoked).


Phase 3: Adapt, Automated Repair Engine

This is the "surgical operation" phase of Skill Curator. It runs fully automatically, but always backs up first.

3.1 Six-Language Keyword Injection

Core algorithm: parse the semantic content of description → generate corresponding keywords for each missing language → inject them into the description field.

Before fix:
description: "Full-stack web application development with modern frameworks"

After fix:
description: "Full-stack development Website creation Frontend and backend Web application development Full-stack development"
تطوير الويب Full-stack development site web Full-stack web application development 
with modern frameworks"

Six-language coverage:

  • 🇹🇼 Traditional Chinese
  • 🇨🇳 Simplified Chinese
  • 🇯🇵 Japanese
  • 🇰🇷 한국어
  • 🇸🇦 العربية
  • 🇬🇧 English (original)

3.2 Frontmatter Repair

Automatically fix common formatting corruption:

  • YAML indentation error → Reformat
  • Missing description → Automatically extract and generate from the skill content
  • Non-standard fields → keep but move to comment area
  • Encoding issues → standardize to UTF-8

3.3 Safety Mechanism: Always Back Up First

# Run automatically before each fix
cp skills/frontend-design/SKILL.md skills/frontend-design/SKILL.md.bak.$(date +%s)

Result: In the first curation pass, 61 skills were automatically repaired. Zero errors. Because the original files were preserved before every modification.


Phase 4: Scenario Generation Scenarios, "This Is How You Should Use Me"

Fixing only makes skills discoverable. Scenario generation enables skills to be used correctly.

Automatically generate 3-5 trigger scenarios for each skill:

## 🧪 Trigger Scenarios

| What you want to do | What you should say | Trigger Skill |
|-----------|---------|---------|
| Build a company website | "Help me make an official website" | Frontend Design |
| Stock research | "Research Tencent" | AK-HK-Stock-DD |
| Generate an investment proposal | "Write an investment proposal" | AK-Investment-Proposal |
| Send daily report email | "Send today's report" | Email Report |
| Create a Feishu document | "Create meeting minutes" | Feishu Create Doc |

The Value of Scenario Generation:

  • Lower learning costs: Users do not need to know skill names, only "what I want to do"
  • Improve trigger accuracy: Trigger words in the scenario are synced to description, creating a positive feedback loop
  • Discover hidden features: Users may not know that the Agent has certain skills, and scenario generation surfaces them

Phase 5: Report, Transparency Is the Cornerstone of Trust

After curation is complete, a structured report is automatically generated and pushed to Feishu:

## 📊 Skill Curator Curation Report: 2026-06-10

### Health Trend
- Overall health rate: 51.2% → 99.2% (+48.0%)
- Fatal skills: 6 → 1 (-83.3%)
- Warning skills: 55 → 0 (-100%)
- Healthy skills: 64 → 124 (+93.8%)

### Actions Performed This Run
- 🔧 Auto repair: 61 skills
- 🌏 Keyword injection: 61 skills (average +15 keywords each)
- 🔨 Frontmatter repair: 4 skills
- 📋 Scenario generation: 125 skills (487 scenarios total)
- ⚠️ Recommended uninstall: 3 overlapping skills (pending approval)
- 👻 Ghost skills: 1 (cannot be repaired, marked)

### Language Coverage Improvement
- Before repair: average 1.2 languages/skill
- After repair: average 5.8 languages/skill

Report Design Principles:

  • Not a technical log for developers, but a health report for decision-makers
  • Every operation is traceable, including backup file paths, modified content, and timestamps
  • Advisory operations require explicit user approval, with no black-box decisions

Phase 6: Execute, Decision Boundaries for Human-AI Collaboration

This is Skill Curator's most critical design principle: the boundaries of automation.

Operation TypeExecution MethodReason
Keyword injection🤖 Fully automaticNo risk, reversible, backed up
Frontmatter repair🤖 Fully automaticFormat correction, no semantic changes
Scenario generation🤖 Fully automaticAdditional content, does not modify original files
Deduplication merge👤 Requires approvalInvolves functional changes
Skill uninstallation👤 Requires approvalIrreversible operation
Routing rule adjustment👤 Requires approvalAffects global behavior

Design philosophy: Automate what can be automated; clearly mark what cannot be automated and wait for the user to decide. An Agent should not make value judgments on behalf of humans.


⚔️ Chapter 4: Hands-On, the Complete Record of the First Curation

Timeline

14:02
Phase 1: Full Scan Started
Traverse 125 skills/ directories, extract frontmatter, and measure language coverage.
14:03
Phase 2: Diagnosis Complete
6 fatal, 55 warnings, 64 healthy. Generated a repair plan: 61 skills require keyword injection.
14:04
Phase 3: Automated Adaptation Started
Back up one by one → inject keywords → validate format. Processing about 2 skills per second.
14:35
Phase 3: Adaptation Complete
61 skills repaired. 0 errors. 4 fatal skills fixed due to frontmatter YAML misalignment.
14:36
Phase 4: Scenario Generation
Generated 487 trigger scenarios for all 125 skills.
14:38
Phase 5: Report Delivery
Pushed the full curation report to Feishu. Health score rose from 51.2% to 99.2%.

Total time: 36 minutes. How long would it take to do the same work manually?

  • Keyword injection for 61 skills: at least 5 minutes each = 305 minutes
  • Frontmatter troubleshooting and repair: about 30 minutes
  • Scenario generation: about 120 minutes
  • Total: about 7.5 hours → Skill Curator completed it in 36 minutes, 12.5 times faster.

Deep Dive into Repairs: Using "Frontend Design" as an Example

SKILL.md before repair:

---
name: frontend-design
description: "Production-quality frontend UI engineering with Next.js, Tailwind CSS, and modern component patterns"
---

Issue Diagnosis:

  • 🟡 Warning: description contains only English
  • Chinese users say "website," "frontend," "interface" → all fail to match
  • Japanese/Korean/Arabic users are likewise unable to trigger it

SKILL.md after the fix:

---
name: frontend-design
description: "Website design Website building Frontend development Web page creation UI interface Frontend design UI engineering 
Webデザイン フロントエンド 웹디자인 프론트엔드 تصميم واجهات واجهة المستخدم 
Production-quality frontend UI engineering with Next.js, Tailwind CSS, and modern component patterns"
---

Results after the fix:

  • Chinese trigger success rate: 0% → 95%+
  • Multilingual coverage: 1 language → 6 languages
  • Number of trigger words: ~8 words → ~25 words
  • Skill discovery rate improvement: 6x

🔧 Chapter 5: Automated Repair Technical Details

5.1 Language Keyword Matrix

Skill Curator maintains a six-language keyword mapping table. It is not a simple translation; instead, it selects based on the vocabulary users will actually use:

ConceptTraditional ChineseEnglishJapaneseKoreanArabic
Websitewebsite webpage site buildingwebsite web siteWebサイト ホームページ웹사이트 홈페이지موقع إلكتروني
Frontendfrontend frontend developmentfrontend UIフロントエンド프론트엔드واجهة أمامية
Designdesign interface layoutdesign layoutデザイン レイアウト디자인 레이아웃تصميم
Deploymentdeployment launch releasedeploy publishデプロイ 公開배포 게시نشر
Financefinance stocks investmentfinance stock investfinance stocks investment금융 주식 투자مالية أسهم استثمار

5.2 Keyword Injection Algorithm

1. Read current description
2. Identify semantic topics in description (via keyword matching)
3. For each missing language:
   a. Select matching phrases from keyword mapping table
   b. Ensure no duplicate injection of existing vocabulary
   c. Sort by frequency (high-frequency words first)
4. Append new keywords to end of description (preserve original English content)
5. Write to file (back up first)

5.3 Frontmatter Repair Strategies

Common corruption patterns and repair methods:

Corruption PatternFrequencyAutomatic Repair
YAML indentation error (using tab instead of space)2/6✅ Automatic conversion
Missing description: field3/6✅ Extracted from body text
Non-standard frontmatter field4/125✅ Retained but flagged
UTF-8 BOM residue1/6✅ Automatically removed
Empty description (description: "")1/6✅ Generated from skill name

🎯 Chapter 6: The Value Chain of Scenario Generation

Scenario generation is not just "adding a few sentences"; it is the bridge connecting user intent and skill capabilities.

The Three Tiers of Goals in Scenario Generation

🎯 First Tier

Reduce Cognitive Load

Users do not need to remember the names and functions of 125 skills. They only need to say "what I want to do," and the scenario is matched automatically.

🔄 Second Tier

Positive Trigger Loop

Trigger words in the scenario are synchronized to the description. The more users use it, the more precise the matching becomes, and the more they use it. This creates a self-reinforcing loop.

🔍 Third Tier

Expose Hidden Capabilities

Many high-value skills are never discovered because their names are too technical. Scenario generation translates their capabilities into everyday language.

Scenario Generation Example

Skill: DD-Meeting-Prep (Due Diligence Meeting Preparation)

## 🧪 Trigger Scenarios

| Scenario | User Says | Triggered Skill | Prerequisites |
|------|--------|---------|---------|
| Pre-M&A Preparation | "I have a meeting with the target company tomorrow, help me prepare questions" | DD-Meeting-Prep | Requires target company name |
| Investor Meeting | "List 20 questions to ask during due diligence" | DD-Meeting-Prep | Requires industry background |
| Risk Assessment | "What potential risks should I ask about for this company?" | DD-Meeting-Prep | Requires basic information |

📦 Chapter 7: One-Line Installation

Skill Curator's design philosophy is consistent throughout: zero-friction deployment.

mkdir -p skills/skill-curator && curl -sSL \
  https://raw.githubusercontent.com/Bryan-cmf/agentic-infrastructure/main/skill-curator/SKILL.md \
  -o skills/skill-curator/SKILL.md

Just one line. No npm install, no pip install, no configuration files. Download → takes effect.

This is the design principle of the Agentic Infrastructure seven-piece suite: each skill is a standalone SKILL.md file, with no external dependencies, and is installed in one line.

Experience After Installation

After installation, the Agent automatically gains the following capabilities:

  • Say "scan skill health" → triggers a full scan
  • Say "curate skills" → triggers the complete six-stage process
  • Say "repair skill keywords" → triggers automatic adaptation
  • Say "generate skill scenarios" → triggers scenario generation

No configuration needed. No API key needed. No learning needed.


📊 Chapter 8: Data, Quantitative Comparison Before and After Curation

Key Metrics

MetricBefore curationAfter curationChange
Overall health rate51.2%99.2%+48.0%
Fatal Skill61*-83.3%
Warning skills550-100%
Health skills64124+93.8%
Average language coverage1.25.8+383%
Non-English trigger success rate~20%~95%+375%
Total trigger word count~1,200~3,800+217%

* The remaining 1 fatal skill is "Ghost Skill". SKILL.md exists, but its content is an incorrect copy of another skill and cannot be automatically repaired; it has been flagged for manual handling.

Language Coverage Changes

{[ { lang: "🇬🇧 English", before: "100%", after: "100%", note: "Original" }, { lang: "🇹🇼 Traditional Chinese", before: "8%", after: "99%", note: "+91%" }, { lang: "🇨🇳 Simplified Chinese", before: "8%", after: "99%", note: "+91%" }, { lang: "🇯🇵 Japanese", before: "3%", after: "99%", note: "+96%" }, { lang: "🇰🇷 한국어", before: "2%", after: "99%", note: "+97%" }, { lang: "🇸🇦 العربية", before: "1%", after: "99%", note: "+98%" }, ].map((item, i) => (
{item.lang}
{item.before} {item.after}
{item.note}
))}

🔮 Chapter 9: Design Insights, What We Learned

Insight 1: Skill ≠ Capability. Curation ≈ Capability.

An uncurated skill is a potential capability, not an actual capability. It is like a closed book: you know it is there, but you will never read it.

Curation is the transformation process from "having" to "being usable." Without this process, your skill library is a digital tomb.

Insight 2: The Boundary of Automation Is the Most Important Design Decision

The greatest design highlight of Skill Curator is not how many skills it can repair, but that it clearly knows what should not be repaired automatically.

  • Keyword injection: automatic → reversible, risk-free
  • Uninstalling skills: manual approval → irreversible, risky

This boundary defines the collaboration model between the Agent and humans. The Agent is an executor, not a decision-maker.

Insight 3: Multilingualism Is Not an Add-On Feature, It Is Infrastructure

For 75% of internet users worldwide, their native language is not English. If a skill only has an English description, that means 75% of potential users cannot use it effectively.

Six-language injection is not a "nice-to-have"; it is a fundamental requirement that makes skills usable for global users.

Insight 4: Health Requires Continuous Monitoring

Curation is not a one-time event. As new skills are added, old skills are updated, and descriptions naturally drift, health will decline again.

Skill Curator is designed to be reusable. It is not a one-time fix, but a resident Agent skill that performs regular health checks.


🏁 Conclusion: Making Your Skills Truly Alive

"The moment a skill is downloaded, it merely exists. Only after it is curated does it come alive."

When we first ran Skill Curator and saw the health rate jump from 51.2% to 99.2%, it did not feel like "fixing a bug"; it felt like "waking a sleeping army."

The 125 skills had been there all along. They just needed someone to wake them up.

One-line install to awaken your skill library:

mkdir -p skills/skill-curator && curl -sSL \
  https://raw.githubusercontent.com/Bryan-cmf/agentic-infrastructure/main/skill-curator/SKILL.md \
  -o skills/skill-curator/SKILL.md

📚 Further reading: Skill Curator is the fourth layer of the Agentic Infrastructure seven-part suite, skill system deduplication and streamlining. Read the complete architecture article to learn about all seven layers of the design.