XSkill Deep Source Code Analysis: Technical Design, Risks, and Implications for MemoryHub from an ICML 2026 Paper
Analysis Subject: XSkill-Agent/XSkill (198 ⭐, ICML 2026), Continual Learning from Experience and Skills in Multimodal Agents
Methodology: Source-code-level breakdown + technical design assessment + risk analysis + applicability benchmarking against MemoryHub
Core Question: Which aspects of XSkill's design are worth learning from? Which carry hidden risks? What do they mean for our memory system?
Introduction: Why Analyze XSkill?
In building MemoryHub, the core tension we encountered is: How can an Agent "truly learn" from prior executions, instead of starting from scratch every time?
XSkill provides a complete answer to this question. It proposes a continual learning framework that does not require training model parameters: it enables the Agent to automatically extract structured "skills" and "experience" from execution trajectories, then inject relevant knowledge at inference time through semantic retrieval. Its acceptance at ICML 2026 demonstrates academic recognition of this direction.
However, every technical design involves trade-offs. This article is not merely praise; we break down each module's source code one by one and mark the implicit design risks within them.
1. XSkill Architecture Overview
Two-Stage Design
Phase I: Accumulation Phase II: Inference
Agent executes task → trajectory recording New task arrives
↓ ↓
Visual summary (image + text) Task decomposition (LLM)
↓ ↓
Cross-trajectory comparative critique (LLM) Subtask retrieves relevant experience
↓ ↓
┌──────────────────┐ Experience rewriting adapted to current context
│ Skill Library │ ↓
│ Experience Bank │ Injected into System Prompt
└──────────────────┘ ↓
↓ Agent inference
Hierarchical merging + quality control
2. Dissecting the Eight Key Technical Points One by One (with Source-Level Analysis)
Technical Point 1: The Three-Phase Skill Lifecycle
Source code: eval/exskill/skill_builder.py (232 lines)
In XSkill, a Skill is not a static file but a dynamic entity with a complete lifecycle:
# Phase 1: Generate: Generate Skill from a single trajectory
def generate_skill_for_sample(sample_info, llm, ground_truth):
trajectory = extract_trajectory_from_file(sample_dir)
prompt = GENERATE_SKILL_PROMPT.format(
trajectory=trajectory,
ground_truth=ground_truth or "[NOT PROVIDED]"
)
skill = llm.chat(prompt)
# Directly output structured Markdown, including workflow + tool templates
# Phase 2: Merge: Merge multiple Skills into the global library
def merge_skills(existing, new_skills, llm):
prompt = MERGE_SKILL_PROMPT.format(
existing_skill=existing,
new_skills="\n\n".join(new_skills)
)
return llm.chat(prompt)
# Phase 3: Adapt: Adapt to the current task during inference
def adapt_skill_for_task(base_skill, experiences, task, llm, images):
prompt = ADAPT_SKILL_PROMPT.format(
base_skill=base_skill,
experiences=experiences,
task=task
)
# Support image input -> visual context adaptation
if images:
return llm.chat_with_image(prompt, images)
return llm.chat(prompt)
Implications for MemoryHub: Our AK-SDD SKILL.md is currently a static, handwritten document. It could be upgraded so that after each research task is completed, a Skill draft is automatically generated, merged into the global SKILL.md, and adapted at inference time based on the target type (Main Board/GEM/suspension).
Technical Point 2: The Two-Layer Experience Structure
Source code: eval/exskill/experience_manager.py (651 lines, the most complex module)
XSkill divides memory into two complementary layers:
Skill (Task-level) Experience (Action-level)
───────────────── ─────────────────────
Structured Markdown (condition, action, embedding) triple
Workflow + Tool Templates Contextual Tactical Insights
Cross-task Sharing Task-specific
# Experience data structure (derived from code)
Experience = {
"id": "exp_abc123",
"condition": "When the HKEXnews ASP.NET session expires",
"action": "Use CloakBrowser to re-establish session, rather than retrying web_fetch",
"embedding": [0.123, -0.456, ...], # used for semantic retrieval
"source_trajectory": "sample_0653_v4.3",
"quality_score": 0.85,
"created_at": "2026-05-20",
}
Our counterpart: The Lessons section of MEMORY.md currently mixes all lessons together. It should be split into task-level Skills (for example, "Hong Kong stock research DI query workflow") and action-level Experiences (for example, "DI page timeout: switch to CloakBrowser").
Technical Point 3: Cross-Trajectory Contrastive Critique
Source code: eval/exskill/experience_critique.py (127 lines)
This is XSkill's most ingenious mechanism. It lets the LLM examine multiple execution paths for the same task simultaneously and extract experience from the differences:
def intra_sample_experiences(question, groundtruth, summaries, llm):
"""Multiple trajectories for the same task → LLM comparison → extract experience"""
formatted_summaries = []
for i, (nid, summ) in enumerate(summaries.items()):
formatted_summaries.append(f"Trajectory {i+1} (node {nid}):\n{summ}")
prompt = INTRA_SAMPLE_CRITIQUE.format(
question=question,
summaries="\n\n".join(formatted_summaries), # Multiple trajectories side by side
groundtruth=groundtruth,
max_ops=2 # extract at most 2 experience operations
)
resp = llm.chat(prompt, max_tokens=12288)
# LLM returns JSON operation list
ops = json.loads(resp.split("```json")[-1].split("```")[0])
# [{create: {...}}, {update: {...}}, {merge: {...}}, {delete: {...}}]
return ops
Core insight: Run the same task 3 times: the 1st time fails, the 2nd time succeeds but is slow, the 3rd time succeeds and is fast → LLM comparative analysis → automatically extract "why the 3rd time is fastest" → generate Experience.
Our adaptation approach: No need to run it 3 times. Use the historical trajectory of version iterations (v4.3 skipped steps → v4.4 corrected → v4.5 strengthened), and compare the same target (such as 653.HK) across versions, at near-zero cost.
Technical Point 4: Hierarchical Merging and Quality Control
Source code: eval/exskill/experience_manager.py
# Core logic of the experience manager (derived from code)
class ExperienceManager:
def consolidate(self, new_experiences):
for exp in new_experiences:
# 1. Semantic deduplication
similar = self.find_similar(exp, threshold=0.85)
# 2. Conflict resolution
if similar and similar.content != exp.content:
merged = llm.merge(similar, exp)
# 3. Quality scoring (based on success rate of source trajectories)
score = self.quality_score(exp)
# 4. Elimination mechanism
if score < threshold:
exp.status = "deprecated"
Our benchmark: Dream merging currently uses simple append → should add semantic deduplication, quality scoring, and stale retirement. Lessons not retrieved for 3 months → automatically marked as deprecated.
Technical Point 5: Task Decomposition Retrieval
Source code: eval/exskill/experience_retriever.py
This is the core innovation of the XSkill retrieval system: instead of directly searching the entire task, it first decomposes the task into subtasks and independently retrieves for each subtask:
class ExperienceRetriever:
def retrieve_with_decomposition(self, task, images):
# Step 1: LLM decomposes the task
subtasks = llm.decompose(task)
# "Research 0653.HK" → ["Check trading status", "Check DI shareholding", "Check financials", "Risk control"]
# Step 2: Retrieve each subtask independently
all_exps = []
for subtask in subtasks:
emb = self.embed(subtask)
top_k = self.cosine_search(emb, top_k=3)
all_exps.extend(top_k)
# Step 3: Rewrite experiences to adapt to the current context
rewritten = llm.rewrite(all_exps, task_context)
return rewritten
Implications for MemoryHub: This is the best-practice validation of our "query router" approach. "Do Hong Kong stock research for me" → automatically decompose → retrieve relevant lessons for each substep → inject into Prompt.
Technical Point 6: Embedding Cache and Incremental Updates
# Use MD5 hash for cache validation
def _compute_library_hash(self):
sorted_exps = sorted(self.experiences.items())
content = json.dumps(sorted_exps, ensure_ascii=False, sort_keys=True)
return hashlib.md5(content.encode('utf-8')).hexdigest()
# hash unchanged → load cache; hash changed → regenerate
def _load_or_generate_embeddings(self):
if cache_exists and cached_hash == current_hash:
self._experience_embeddings = load_from_cache()
else:
self._generate_all_embeddings(batch_size=30)
Our benchmark: MemoryHub auto_sync can add hash caching; if follow_up_tracker.json has not changed, skip re-embedding, greatly reducing BGE model calls.
Technical Point 7: Visually Grounded Summarization
Source code: eval/exskill/trajectory_summary.py
# Core Flow of Trajectory Summarization
def summarize_rollout(trajectory_jsonl, sample_dir):
# 1. Scan all images
all_images = _scan_all_images(sample_dir)
# 2. Generate image captions with VLM
image_captions = generate_image_captions(all_images, vlm)
# 3. Replace <image> tags in the trajectory with descriptions
enriched = _replace_image_refs_in_jsonl(trajectory, image_captions)
# 4. LLM generates summary (the text now includes semantic descriptions of the images)
summary = llm.summarize(enriched, question, ground_truth)
We do not need image processing, but the mindset is worth borrowing: A trajectory summary should record not only "what was done", but also "why it was done this way" and "what the outcome was." The daily log should automatically extract decision points and tool call results from Session JSONL.
Technical Point 8: Automatic Skill Slimming
def refine_skill_document(skill_content, word_threshold=1000):
"""When SKILL.md exceeds 1000 words, automatically trigger LLM refinement"""
word_count = len(skill_content.split())
if word_count < word_threshold:
return skill_content
prompt = SKILL_REFINE_PROMPT.format(
word_count=word_count,
skill_content=skill_content
)
refined = llm.chat(prompt, max_tokens=8192)
# 1000 words → 600 words, preserve core logic
return refined
This is exactly what MEMORY.md needs. It currently has 1,193 lines and keeps growing → after each Dream merge, automatically trigger refine → historical entries are compressed into a one-line summary → Token usage remains controllable.
III. Risk Analysis: The Hidden Cost of Each Technical Point
Every design involves trade-offs. Below are the risks we identified from the source code and architecture:
Risk Matrix
| # | Technical Point | Greatest Risk | Severity | Root Cause |
|---|---|---|---|---|
| 1 | Skill auto-generation | Hallucinated Skill without validation; the LLM may extract an incorrect process from failure trajectories | 🔴 | No backtesting, no confidence score |
| 2 | Embedding retrieval | Semantic similarity ≠ actual relevance; MIN_SIMILARITY = 0.0 (no filtering!) | 🔴 | Popularity bias, no diversity penalty |
| 3 | Cross-trajectory comparison | Unfair comparison; trajectory A uses GPT-4, trajectory B uses DeepSeek | 🔴 | No controlled variables, 5× API cost |
| 4 | Hierarchical merging | Over-merging loses details; each merge drops 10%, after 5 merges only a skeleton remains | 🟠 | LLM tends to generalize |
| 5 | Task decomposition | Incorrect decomposition causes cascading failure; step 1 is wrong → all retrieval is invalid | 🔴 | No decomposition validation, no Token budget |
| 6 | Embedding cache | Incremental updates are ineffective; adding 1 experience → hash changes → 1000 entries rebuilt | 🟠 | Hash is not increment-aware |
| 7 | Skill pruning | Key rules are lost; "HK$0.14" is compressed into "low-priced stock" | 🔴 | Irreversible, no numeric protection |
| 8 | Cascading failure | 10+ LLM calls in series; error propagation 1-0.95⁸=34% | 🔴🔴 | No end-to-end validation |
Cascading Failure (System-Level Risk)
In the XSkill pipeline, the LLM is invoked in more than 10 places:
Trajectory summary → Experience extraction → Experience merging → Skill generation → Skill merging
→ Skill streamlining → Task decomposition → Experience retrieval → Experience rewriting → Final reasoning
95% accuracy per step × 10 steps = final accuracy of only 60%. This is the deepest risk in the XSkill architecture and is not discussed in the paper.
Defenses We Should Add
| Risk | Defense |
|---|---|
| Hallucinated Skill | Backtest after generation; do not release if success rate <80% |
| Retrieval without threshold | MIN_SIMILARITY = 0.6 + MMR diversity reranking |
| Over-merging | Retain Diff + 30-day rollback |
| Task decomposition cascade | Cache decomposition results + hard Token budget cap + consistency validation |
| Skill pruning | Keyword whitelist (forced retention for "must not", "must", "prohibited", "%") |
| Cascading failure | Pipeline degradation switches + end-to-end quality monitoring |
4. XSkill ↔ MemoryHub Mapping
| Concept | XSkill | MemoryHub | Adaptation Recommendations |
|---|---|---|---|
| Skill Library | Skill Library (LLM generation + merging) | AK-SDD SKILL.md (handwritten) | Introduce automatic generation + merging |
| Experience Bank | Experience Bank (condition-action-emb) | Lessons (unstructured) | Structure into triples + embeddings |
| Retrieval | Task decomposition + Cosine + rewriting | Qdrant COSINE | Add task decomposition routing |
| Merging | LLM-based merge + quality scoring | Dream append | Add deduplication + quality + elimination |
| Compaction | LLM refine (word_threshold) | None | T5 MEMORY.md compression |
| Caching | MD5 hash + npy persistence | auto_sync full scan | Add incremental hash caching |
| Multimodal | Visual summary + VLM | N/A (text only) | Not needed |
5. Conclusion: What to Borrow and What Not to Borrow
✅ Worth Borrowing (Low Risk, High Value)
- Skill three-stage lifecycle, architectural concept (not the fully automated implementation)
- Experience two-layer structure, structured lessons as (condition, action, embedding)
- Task decomposition routing, best practices for query routers
- Hash incremental caching, greatly reduces embedding model calls
- Skill slimming mechanism, but requires adding a whitelist and backup
❌ Borrow with Caution (High Risk, Needs Improvement)
- Fully automated Skill generation, requires adding a validation loop first
- Threshold-free embedding retrieval, requires setting MIN_SIMILARITY first
- LLM automatic merging, requires preserving Diff and rollback first
- Chained pipeline, requires adding fallback and end-to-end validation first
In one sentence: XSkill's greatest contribution is not any single technique, but the proof that an Agent's continual learning does not require retraining parameters. Structured knowledge extraction + semantic retrieval + Prompt Injection is enough. However, its implementation contains multiple design assumptions of "over-relying on LLM generation without validation," which require quality guardrails to be added in production environments.
This article is based on source code analysis of XSkill-Agent/XSkill v1.0 (Commit: main branch, 2026-05-23). Lines of code: skill_builder.py (232), experience_critique.py (127), experience_retriever.py (~300), trajectory_summary.py (~300), experience_manager.py (651).
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built