SkillOpt Deep Technical Breakdown: Training Agent Skill as a Neural Network
Core insight: Treat the Agent Skill Document as the weights of a neural network and train it, using rollout for forward propagation, reflect for backpropagation, bounded edits for gradient updates, and a held-out gate as the validation set.
Codebase size: ~16,000 LOC (core engine 11,215 + Sleep engine 5,114)
Experimental results: 52/52 wins, GPT-5.5 average +23.5 point improvement
1. Introduction: The Evolutionary Dilemma of Agent Skills
What makes today's AI Agents smarter? The answer is paradoxical: handwritten prompts.
We have spent countless hours fine-tuning model weights, yet we leave the Agent's behavioral rules to "one-shot LLM generation" or "uncontrolled self-modification." Worse still, these skill documents have almost no quality assurance: an edit may improve them, make them worse, or even cause them to collapse completely. No one implements a validation gate, no one controls the edit budget, and no one guarantees convergence.
SkillOpt's answer is simple: treat the skill document as the weights of a neural network and train it.
This framework, published by Microsoft Research, is the first known systematic and controllable text-space Agent Skill optimizer. It does not touch model weights (frozen agent); instead, it trains the skill document only in text space through a loop of rollout → reflect → bounded edit → held-out gate. At deployment, it has zero inference overhead and produces only a compact best_skill.md (300-2,000 tokens).
2. Core Insight: Deep Learning Analogy
The entire design of SkillOpt revolves around one core analogy: treating the skill document as the weights of a neural network and training it. This is not superficial rhetoric, but an architectural-level design decision. Once you understand this mapping, you know how to tune parameters, how to interpret convergence curves, and how to diagnose optimization failures.
| Deep Learning Concept | SkillOpt Mapping | Description |
|---|---|---|
| Model weights | Skill document (.md) | The object being optimized |
| Forward pass | Rollout | Target executes the task with the current skill |
| Loss function | Task evaluator | Scores the quality of task execution |
| Backpropagation | Reflect | Optimizer analyzes failures → generates edit patches |
| Gradients | Edit patches | Specific modification suggestions for the skill |
| Gradient aggregation | Patch aggregation | Merges semantically similar edits |
| Gradient clipping | Edit selection (clip) | Limits the maximum number of edits per step |
| Learning rate | learning_rate (edit budget) | Controls the magnitude of each update |
| LR Scheduler | lr_scheduler | Decay strategy: cosine / linear / constant |
| SGD step | Skill update | Applies the selected patches to the document |
| Validation set | Selection split | Gate checks whether there is improvement on the validation set |
| Early Stopping | Gate patience | Rejects updates that do not improve |
| Momentum | Slow update | Longitudinal comparison at epoch boundaries |
| Meta-learning | Meta skill | Optimizer strategy memory across epochs |
| Batch size | batch_size | Number of tasks sampled per step |
Experimentally validated transfer rules:
✅ Cosine schedule > constant, consistent with DL: cosine annealing helps convergence
✅ Moderate LR (4-16) > extremely high/extremely low, too few edits = learning is too slow, too many = noise
✅ Slow update is effective, longitudinal comparison prevents catastrophic forgetting across epochs
✅ Meta skill improves reflection, Optimizer benefits from cross-epoch strategy notes
⚠️ Batch size ≠ better, larger rollout batches have diminishing returns (limited by API cost)
⚠️ More epochs ≠ better, Skills converge faster than neural networks (2-4 epochs are usually sufficient)
3. Training Loop: Detailed Explanation of the Six Stages
SkillOpt's training loop is a two-level Epoch → Step structure:
for epoch in epochs:
for step in steps:
1. Rollout : Target executes task (forward propagation)
2. Reflect : Optimizer analyzes trajectory (backward propagation)
3. Aggregate : Merge semantically similar patches
4. Select : Sort + clip edits (gradient clipping)
5. Update : Apply patches to skill doc (parameter update)
6. Gate : Validation set check (validation)
Epoch Boundary:
Slow Update : Longitudinal comparison + prevent catastrophic forgetting
Meta Skill : Cross-epoch strategy memory
3.1 Rollout (Forward Pass)
The Target model uses the current skill document as a prompt to execute a batch of tasks. Each task produces a trajectory and a score.
# Analogy: neural network's forward pass
predictions = model(input, skill_document)
scores = evaluate(predictions, ground_truth)
Rollout supports three execution environments (harnesses):
- Direct chat: Direct API calls, suitable for QA benchmarks
- Codex CLI: Executed through the Codex agentic loop
- Claude Code CLI: Executed through the Claude Code agentic loop
Source code: skillopt/envs/<benchmark>/rollout.py
3.2 Reflect (Backward Pass), the Most Critical Innovation
It is not simply letting the LLM take a look at a failure and make arbitrary changes. SkillOpt's reflect has a three-layer design:
Layer 1: Minibatch trajectory analysis
Failed trajectories are analyzed in minibatches (size M) rather than one by one. This is similar to minibatch SGD vs per-sample SGD; batch analysis can capture systematic patterns rather than individual noise.
Layer 2: Dual analysis of Error + Success
It analyzes both failure and success cases at the same time. The Optimizer not only knows "what went wrong," but also knows "what approach was right."
Layer 3: Skill-aware reflection (v0.1+)
It distinguishes two types of failure:
- SKILL_DEFECT: A defect in the skill itself → modify the skill body
- EXECUTION_LAPSE: An execution lapse → only record a reminder in the appendix (do not modify the body)
# Analogy: neural network's backward pass
gradients = loss.backward() # → edit patches
Source code: skillopt/gradient/reflect.py, run_minibatch_reflect() · skillopt/optimizer/skill_aware.py
3.3 Aggregate (Gradient Aggregation)
Semantically similar edit patches are merged hierarchically to avoid duplicate edits wasting the edit budget.
Source code: skillopt/gradient/aggregate.py
3.4 Select: Textual Learning Rate and Gradient Clipping
This is the design that best embodies "textual deep learning."
Textual Learning Rate is not a floating-point number, but an integer N: at most N edits are applied per step. lr_scheduler controls how this N decays over the training process:
| Scheduler | Behavior |
|---|---|
| cosine | Early exploration (high LR), later convergence (low LR) |
| linear | Linear decay |
| constant | Fixed budget |
| autonomous | The Optimizer decides the next LR itself |
When the number of edits exceeds the budget, rank_and_select() calls the optimizer LLM to rank them by importance and keeps only the top-L edits.
# Analogy: gradient clipping
selected = top_k(edits, k=learning_rate)
Source code: skillopt/optimizer/clip.py, rank_and_select() · skillopt/optimizer/scheduler.py
3.5 Update (Parameter Update)
Selected edits are applied to the skill document. There are three operations:
| Edit Op | Description |
|---|---|
ADD | Append a rule/step at the end |
DELETE | Delete a line by exact match |
REPLACE | Replace a line by exact match |
Key constraint: bounded edits only. Adding entirely new sections is not allowed (to prevent catastrophic forgetting), and deleting entire paragraphs is not allowed.
A Skill document has a protected region (marked by <!-- SLOW_UPDATE -->). The Edit engine skips any modifications within the protected region.
Source code: skillopt/optimizer/skill.py, apply_edit()
3.6 Gate (Validation Gating), the Soul of SkillOpt
This is the most fundamental difference between SkillOpt and all other skill optimization methods.
A candidate skill must strictly improve its score on the held-out selection split before it is accepted. There is no accepting "close enough" or "possibly better"; it must be a strict improvement.
Three gate metrics:
| Metric | Description | Applicable scenario |
|---|---|---|
hard | Exact match accuracy | Tasks with clear answers |
soft | F1 / partial credit | Open-ended tasks |
mixed | (1-w)×hard + w×soft | Balances the two |
Rejected edits are not discarded; they are passed to the next round's optimizer as negative feedback (rejected-edit buffer).
Source code: skillopt/evaluation/gate.py
Epoch Boundary: Slow Update & Meta Skill
Slow Update (similar to Momentum):
- Simultaneously rollout the previous epoch's skill and the current skill on the same samples
- Classify them as: improved / regressed / persistent_fail / stable_success
- Generate high-level guidance and inject it into the protected region
- Prevent catastrophic forgetting such as "fixed A but broke B"
Meta Skill (similar to Meta-learning):
- Accumulate policy memory across epochs
- The optimizer gains cross-epoch context during future reflection
4. Source Code Architecture Overview
skillopt/ # Core engine (~11,215 LOC)
├── gradient/
│ ├── reflect.py # Backpropagation: error/success analysis → edit patches
│ └── aggregate.py # Gradient aggregation: hierarchical merging of similar patches
├── optimizer/
│ ├── clip.py # Gradient clipping: LLM ranking + top-L selection
│ ├── skill.py # Parameter update: apply_edit / protected region
│ ├── scheduler.py # LR scheduler: cosine/linear/constant/auto
│ ├── slow_update.py # Momentum: longitudinal comparison at epoch boundaries
│ ├── meta_skill.py # Meta-learning: cross-epoch strategy memory
│ ├── skill_aware.py # Skill-aware reflection
│ └── appendix.py # Protected appendix management
├── evaluation/
│ └── gate.py # Validation gating: accept/reject decision
├── model/
│ ├── azure_openai.py # Azure OpenAI backend
│ ├── claude_backend.py # Claude (Anthropic) backend
│ ├── codex_backend.py # Codex CLI backend
│ ├── qwen_backend.py # Qwen (vLLM) backend
│ ├── minimax_backend.py # MiniMax backend
│ └── router.py # Model routing
├── envs/ # 6 benchmarks
│ ├── searchqa/ # Search engine QA
│ ├── spreadsheetbench/ # Spreadsheet tasks
│ ├── livemathematicianbench/ # Mathematical reasoning
│ ├── alfworld/ # Embodied AI tasks
│ ├── officeqa/ # Office scenario QA
│ └── docvqa/ # Document visual QA
├── prompts/ # Analysis prompt templates
└── configs/ # YAML configuration (supports inheritance)
skillopt_sleep/ # Sleep engine (~5,114 LOC)
├── cycle.py # Main loop (harvest→mine→replay→consolidate)
├── consolidate.py # Core optimization logic
├── replay.py # Task replay
├── mine.py # Task mining
├── backend.py # Backend protocol
├── config.py # Configuration loading
├── types.py # TaskRecord / EditRecord / ReplayResult
└── experiments/ # Experiment scripts + benchmarks
Core Function Quick Reference:
| Function | File | Purpose |
|---|---|---|
load_config | config.py | Loads YAML configuration (supports inheritance + key=value override) |
run_minibatch_reflect | gradient/reflect.py | Runs error/success analysis on trajectory minibatch |
merge_patches | gradient/aggregate.py | Hierarchically merges semantically similar patches |
rank_and_select | optimizer/clip.py | LLM ranks edits + clips to LR budget |
build_scheduler | optimizer/scheduler.py | Builds LR scheduler (cosine/linear/constant/auto) |
apply_patch | optimizer/skill.py | Applies edits to skill document |
evaluate_gate | evaluation/gate.py | Pure accept/reject decision |
| run_slow_update | optimizer/slow_update.py | Generates longitudinal guidance at epoch boundaries |
5. In-Depth Analysis of Key Design Decisions
5.1 Why "Text Space" Instead of "Weight Space"?
Fine-tuning has several fundamental problems:
- Cost: Each improvement requires GPU hours + data preparation
- Not interpretable: It is unknown what the model has learned
- Non-transferable: The fine-tuning result of one model cannot be used by another model
- Non-iterative: It cannot be continuously improved after deployment
SkillOpt's text-space approach: zero inference overhead, fully interpretable (the output is Markdown), transferable across models, continuously iterable.
5.2 Why Is Optimizer/Target Separation Needed?
Use a strong model (optimizer) to write a skill for a weak model (target).
Experiments show that this is not merely "distillation." Even in the matched target-as-optimizer setting (using the same model as both target and optimizer), useful edits can be discovered under bounded/validated constraints. The role of the Optimizer is that of an analyst, not a replacement.
5.3 The Clever Use of the Rejected-edit Buffer
Edits rejected by the gate are not garbage; they are structured negative feedback. The next round's optimizer will see that "these directions have been tried but did not work", thereby avoiding repeating mistakes. This is similar to negative reward shaping in RL.
5.4 Protected Region Mechanism
The <!-- SLOW_UPDATE --> block in the Skill document is epoch-level guidance (similar to a momentum term) and will not be modified by step-level edits. This ensures that long-term learning signals are not drowned out by short-term noise.
6. SkillOpt-Sleep: The Deployment-Time Evolution Engine
SkillOpt solves "how to train a skill on a benchmark." But in reality, your coding agent uses its own skill every day to handle real tasks. SkillOpt-Sleep applies the same training discipline to your own daily usage.
6.1 A "Night" Workflow
harvest your session transcripts
→ mine repeated task patterns
→ replay offline (baseline vs candidate)
→ reflect on failures
→ bounded edits
→ GATE on held-out tasks (your real tasks)
→ stage proposal
→ (you) adopt
6.2 Three-Tier Data Split
| Split | Source | Purpose |
|---|---|---|
| train | Real tasks + optional "dreamed" variants | Optimizer learning |
| val (selection) | Real tasks only, held-out | Gate: strict validation |
| test | Real tasks only, never seen | Final performance report |
You can imagine additional training examples to learn rules robustly, but the rules are still judged on real, unseen tasks.
6.3 Controllable Switches
| Switch | Default | Description |
|---|---|---|
--gate on|off | on | Strict held-out gate vs greedy mode |
--rollouts-k K | 1 | Multi-trajectory contrastive reflection (K≥3 greatly improves signal quality) |
--preferences "..." | None | Your personal rules (as a prior to guide the optimizer) |
--optimizer-model | None | Use a strong model to optimize → weaker models benefit |
--budget-tokens N | None | Control nightly API spend |
--auto-adopt | off | Automatic application vs manual review |
6.4 Validated Results
- gbrain-evals
skillopt-v1: defective skill on held-out 0.00 → 1.00 (4 seeds, real tool-use loop) - Cross-model transfer: a skill optimized on one model can improve another model
- Cross-runtime transfer: skill transfer between Codex ↔ Claude Code is effective
- Gate correctly blocks regressions: harmful edits are systematically intercepted
7. OpenClaw Integration in Practice
SkillOpt-Sleep merged the OpenClaw plugin on 2026-06-15 (PR #59), and we were among the first users to install it.
7.1 Architecture
plugins/openclaw/
├── run_sleep.py # Main entry point
├── skillopt_sleep_openclaw.py # DeepSeek + Ollama backend
├── slash_sleep.py # /sleep slash command
├── run_sleep_cron.sh # cron wrapper
├── config.json # engine configuration
├── SKILL.md # OpenClaw skill manifest
└── tests/ # held-out test sets
├── research-cron-tasks.json
├── devops-tasks.json
└── wiki-tasks.json
7.2 Backend Design
OpenClawDeepSeekBackend implements three core methods:
attempt(task, skill, memory): Executes the task with DeepSeek V4 Pro and returns a responsejudge(task, response): Hard score (exact match) + Soft score (LLM judge scores according to rubric)reflect(failures, successes, skill, memory): Analyzes failure/success trajectories, then generates bounded edit proposals
7.3 Cost
Using DeepSeek V4 Pro as the optimizer, $0.02 per night ($0.59/month, ~$7.18/year).
7.4 Cron Configuration
# Automatically run daily at 3:00 AM HKT
0 3 * * * cd ~/.openclaw/workspace/skills/skillopt-sleep && \
bash run_sleep_cron.sh >> ~/.skillopt-sleep/logs/nightly.log 2>&1
8. Experimental Results Highlights
8.1 Main Results
Across 6 benchmarks × 7 target models × 3 execution harnesses = 52 cells, SkillOpt is best or tied for best, beating all competitors (human, one-shot LLM, Trace2Skill, TextGrad, GEPA, EvoSkill).
8.2 Improvements on GPT-5.5
| Harness | Improvement |
|---|---|
| Direct Chat | +23.5 points |
| Codex Agentic Loop | +24.8 points |
| Claude Code | +19.1 points |
8.3 Key Findings from Ablation Experiments
| Ablation | Impact |
|---|---|
| Removing Gate | Performance drops substantially (demonstrating the necessity of held-out validation) |
| Removing Rejected Buffer | Repeatedly makes the same mistakes |
| Removing Slow Update | Catastrophic forgetting; early progress disappears later |
| Using the Target itself as the Optimizer | Still effective (not merely distillation) |
| Cosine > Constant LR | Consistent with experience in DL |
8.4 Transfer Experiments
- Cross-model transfer: A skill trained on one model remains effective when transferred to another model
- Cross-harness transfer: Codex skill → Claude Code is effective, and vice versa
- Cross-benchmark transfer: A skill trained on one benchmark can transfer to a similar benchmark
9. Implications for Agent Developers
9.1 Skill ≠ Prompt
SkillOpt tells us an important fact: a good skill is not written, it is trained. Just as you would not hand-write the weights of a neural network, you should not hand-write complex agent skills. You need an optimization loop.
9.2 Validation Gate Is a Necessity
Self-modification without a gate = random walk. If you change the skill today, tomorrow you will not know whether it has improved or worsened. The gate provides the only quality guarantee.
9.3 Bounded Edits > Unbounded Rewrites
Allowing the optimizer to rewrite the entire skill document leads to catastrophic forgetting. Bounded edits (changing only a few specific locations each time) are key to stable learning.
9.4 Optimizer/Target Separation Is Valuable
Analyzing with a strong model and executing with a weak model are complementary. This means you do not need to train a skill for every model; a skill trained with a strong optimizer can help all target models.
10. Future Outlook
SkillOpt represents a new paradigm shift in Agent engineering:
- From "writing prompts" to "training prompts"
- From "one-shot generation" to "iterative optimization"
- From "no quality assurance" to "held-out validation"
However, there is still a long way to go:
- Multi-skill coordination: How can interactions among multiple skills be coordinated?
- Online learning: How can continuous online learning be performed during deployment?
- Skill composition: How can multiple skills be composed to solve complex tasks?
- Safety constraints: How can safety constraints be maintained during optimization?
For the OpenClaw ecosystem, SkillOpt-Sleep is the most immediate source of value. It gives our existing agent skill system continuous, quality-assured self-evolution capabilities.
This article is based on a complete analysis of the SkillOpt v0.1.0 source code (~16,000 LOC), as well as actual deployment tests in a Mac Studio + DeepSeek V4 Pro environment.
References: Paper arXiv 2605.23904 · GitHub microsoft/SkillOpt · Project page · PyPI: pip install skillopt
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built