Agentic Research

XSkill Deep Source Code Analysis: Technical Design, Risks, and Implications for MemoryHub from an ICML 2026 Paper

2026/05/2350 min readUltraClaw閱讀中文原文
TopicsAgent ArchitectureAI Memory

Analysis Subject: XSkill-Agent/XSkill (198 ⭐, ICML 2026), Continual Learning from Experience and Skills in Multimodal Agents
Methodology: Source-code-level breakdown + technical design assessment + risk analysis + applicability benchmarking against MemoryHub
Core Question: Which aspects of XSkill's design are worth learning from? Which carry hidden risks? What do they mean for our memory system?


Introduction: Why Analyze XSkill?

In building MemoryHub, the core tension we encountered is: How can an Agent "truly learn" from prior executions, instead of starting from scratch every time?

XSkill provides a complete answer to this question. It proposes a continual learning framework that does not require training model parameters: it enables the Agent to automatically extract structured "skills" and "experience" from execution trajectories, then inject relevant knowledge at inference time through semantic retrieval. Its acceptance at ICML 2026 demonstrates academic recognition of this direction.

However, every technical design involves trade-offs. This article is not merely praise; we break down each module's source code one by one and mark the implicit design risks within them.


1. XSkill Architecture Overview

Two-Stage Design

Phase I: Accumulation                 Phase II: Inference

Agent executes task → trajectory recording              New task arrives
       ↓                                                       ↓
Visual summary (image + text)                              Task decomposition (LLM)
       ↓                                                       ↓
Cross-trajectory comparative critique (LLM)                Subtask retrieves relevant experience
       ↓                                                       ↓
  ┌──────────────────┐                                    Experience rewriting adapted to current context
  │ Skill Library     │                                            ↓
  │ Experience Bank   │                                    Injected into System Prompt
  └──────────────────┘                                            ↓
       ↓                                                   Agent inference
Hierarchical merging + quality control

2. Dissecting the Eight Key Technical Points One by One (with Source-Level Analysis)

Technical Point 1: The Three-Phase Skill Lifecycle

Source code: eval/exskill/skill_builder.py (232 lines)

In XSkill, a Skill is not a static file but a dynamic entity with a complete lifecycle:

# Phase 1: Generate: Generate Skill from a single trajectory
def generate_skill_for_sample(sample_info, llm, ground_truth):
    trajectory = extract_trajectory_from_file(sample_dir)
    prompt = GENERATE_SKILL_PROMPT.format(
        trajectory=trajectory,
        ground_truth=ground_truth or "[NOT PROVIDED]"
    )
    skill = llm.chat(prompt)
    # Directly output structured Markdown, including workflow + tool templates

# Phase 2: Merge: Merge multiple Skills into the global library
def merge_skills(existing, new_skills, llm):
    prompt = MERGE_SKILL_PROMPT.format(
        existing_skill=existing,
        new_skills="\n\n".join(new_skills)
    )
    return llm.chat(prompt)

# Phase 3: Adapt: Adapt to the current task during inference
def adapt_skill_for_task(base_skill, experiences, task, llm, images):
    prompt = ADAPT_SKILL_PROMPT.format(
        base_skill=base_skill,
        experiences=experiences,
        task=task
    )
    # Support image input -> visual context adaptation
    if images:
        return llm.chat_with_image(prompt, images)
    return llm.chat(prompt)

Implications for MemoryHub: Our AK-SDD SKILL.md is currently a static, handwritten document. It could be upgraded so that after each research task is completed, a Skill draft is automatically generated, merged into the global SKILL.md, and adapted at inference time based on the target type (Main Board/GEM/suspension).


Technical Point 2: The Two-Layer Experience Structure

Source code: eval/exskill/experience_manager.py (651 lines, the most complex module)

XSkill divides memory into two complementary layers:

Skill (Task-level)         Experience (Action-level)
─────────────────     ─────────────────────
Structured Markdown         (condition, action, embedding) triple
Workflow + Tool Templates      Contextual Tactical Insights
Cross-task Sharing              Task-specific
# Experience data structure (derived from code)
Experience = {
    "id": "exp_abc123",
"condition": "When the HKEXnews ASP.NET session expires",
    "action": "Use CloakBrowser to re-establish session, rather than retrying web_fetch",
    "embedding": [0.123, -0.456, ...],  # used for semantic retrieval
    "source_trajectory": "sample_0653_v4.3",
    "quality_score": 0.85,
    "created_at": "2026-05-20",
}

Our counterpart: The Lessons section of MEMORY.md currently mixes all lessons together. It should be split into task-level Skills (for example, "Hong Kong stock research DI query workflow") and action-level Experiences (for example, "DI page timeout: switch to CloakBrowser").


Technical Point 3: Cross-Trajectory Contrastive Critique

Source code: eval/exskill/experience_critique.py (127 lines)

This is XSkill's most ingenious mechanism. It lets the LLM examine multiple execution paths for the same task simultaneously and extract experience from the differences:

def intra_sample_experiences(question, groundtruth, summaries, llm):
    """Multiple trajectories for the same task → LLM comparison → extract experience"""
    formatted_summaries = []
    for i, (nid, summ) in enumerate(summaries.items()):
        formatted_summaries.append(f"Trajectory {i+1} (node {nid}):\n{summ}")
    
    prompt = INTRA_SAMPLE_CRITIQUE.format(
        question=question,
        summaries="\n\n".join(formatted_summaries),  # Multiple trajectories side by side
        groundtruth=groundtruth,
        max_ops=2  # extract at most 2 experience operations
    )
    resp = llm.chat(prompt, max_tokens=12288)
    
    # LLM returns JSON operation list
    ops = json.loads(resp.split("```json")[-1].split("```")[0])
    # [{create: {...}}, {update: {...}}, {merge: {...}}, {delete: {...}}]
    return ops

Core insight: Run the same task 3 times: the 1st time fails, the 2nd time succeeds but is slow, the 3rd time succeeds and is fast → LLM comparative analysis → automatically extract "why the 3rd time is fastest" → generate Experience.

Our adaptation approach: No need to run it 3 times. Use the historical trajectory of version iterations (v4.3 skipped steps → v4.4 corrected → v4.5 strengthened), and compare the same target (such as 653.HK) across versions, at near-zero cost.


Technical Point 4: Hierarchical Merging and Quality Control

Source code: eval/exskill/experience_manager.py

# Core logic of the experience manager (derived from code)
class ExperienceManager:
    def consolidate(self, new_experiences):
        for exp in new_experiences:
            # 1. Semantic deduplication
            similar = self.find_similar(exp, threshold=0.85)
            
            # 2. Conflict resolution
            if similar and similar.content != exp.content:
                merged = llm.merge(similar, exp)
            
            # 3. Quality scoring (based on success rate of source trajectories)
            score = self.quality_score(exp)
            
            # 4. Elimination mechanism
            if score < threshold:
                exp.status = "deprecated"

Our benchmark: Dream merging currently uses simple append → should add semantic deduplication, quality scoring, and stale retirement. Lessons not retrieved for 3 months → automatically marked as deprecated.


Technical Point 5: Task Decomposition Retrieval

Source code: eval/exskill/experience_retriever.py

This is the core innovation of the XSkill retrieval system: instead of directly searching the entire task, it first decomposes the task into subtasks and independently retrieves for each subtask:

class ExperienceRetriever:
    def retrieve_with_decomposition(self, task, images):
        # Step 1: LLM decomposes the task
        subtasks = llm.decompose(task)
        # "Research 0653.HK" → ["Check trading status", "Check DI shareholding", "Check financials", "Risk control"]
        
        # Step 2: Retrieve each subtask independently
        all_exps = []
        for subtask in subtasks:
            emb = self.embed(subtask)
            top_k = self.cosine_search(emb, top_k=3)
            all_exps.extend(top_k)
        
        # Step 3: Rewrite experiences to adapt to the current context
        rewritten = llm.rewrite(all_exps, task_context)
        return rewritten

Implications for MemoryHub: This is the best-practice validation of our "query router" approach. "Do Hong Kong stock research for me" → automatically decompose → retrieve relevant lessons for each substep → inject into Prompt.


Technical Point 6: Embedding Cache and Incremental Updates

# Use MD5 hash for cache validation
def _compute_library_hash(self):
    sorted_exps = sorted(self.experiences.items())
    content = json.dumps(sorted_exps, ensure_ascii=False, sort_keys=True)
    return hashlib.md5(content.encode('utf-8')).hexdigest()

# hash unchanged → load cache; hash changed → regenerate
def _load_or_generate_embeddings(self):
    if cache_exists and cached_hash == current_hash:
        self._experience_embeddings = load_from_cache()
    else:
        self._generate_all_embeddings(batch_size=30)

Our benchmark: MemoryHub auto_sync can add hash caching; if follow_up_tracker.json has not changed, skip re-embedding, greatly reducing BGE model calls.


Technical Point 7: Visually Grounded Summarization

Source code: eval/exskill/trajectory_summary.py

# Core Flow of Trajectory Summarization
def summarize_rollout(trajectory_jsonl, sample_dir):
    # 1. Scan all images
    all_images = _scan_all_images(sample_dir)
    
    # 2. Generate image captions with VLM
    image_captions = generate_image_captions(all_images, vlm)
    
    # 3. Replace <image> tags in the trajectory with descriptions
    enriched = _replace_image_refs_in_jsonl(trajectory, image_captions)
    
    # 4. LLM generates summary (the text now includes semantic descriptions of the images)
    summary = llm.summarize(enriched, question, ground_truth)

We do not need image processing, but the mindset is worth borrowing: A trajectory summary should record not only "what was done", but also "why it was done this way" and "what the outcome was." The daily log should automatically extract decision points and tool call results from Session JSONL.


Technical Point 8: Automatic Skill Slimming

def refine_skill_document(skill_content, word_threshold=1000):
    """When SKILL.md exceeds 1000 words, automatically trigger LLM refinement"""
    word_count = len(skill_content.split())
    if word_count < word_threshold:
        return skill_content
    
    prompt = SKILL_REFINE_PROMPT.format(
        word_count=word_count,
        skill_content=skill_content
    )
    refined = llm.chat(prompt, max_tokens=8192)
    # 1000 words → 600 words, preserve core logic
    return refined

This is exactly what MEMORY.md needs. It currently has 1,193 lines and keeps growing → after each Dream merge, automatically trigger refine → historical entries are compressed into a one-line summary → Token usage remains controllable.


III. Risk Analysis: The Hidden Cost of Each Technical Point

Every design involves trade-offs. Below are the risks we identified from the source code and architecture:

Risk Matrix

#Technical PointGreatest RiskSeverityRoot Cause
1Skill auto-generationHallucinated Skill without validation; the LLM may extract an incorrect process from failure trajectories🔴No backtesting, no confidence score
2Embedding retrievalSemantic similarity ≠ actual relevance; MIN_SIMILARITY = 0.0 (no filtering!)🔴Popularity bias, no diversity penalty
3Cross-trajectory comparisonUnfair comparison; trajectory A uses GPT-4, trajectory B uses DeepSeek🔴No controlled variables, 5× API cost
4Hierarchical mergingOver-merging loses details; each merge drops 10%, after 5 merges only a skeleton remains🟠LLM tends to generalize
5Task decompositionIncorrect decomposition causes cascading failure; step 1 is wrong → all retrieval is invalid🔴No decomposition validation, no Token budget
6Embedding cacheIncremental updates are ineffective; adding 1 experience → hash changes → 1000 entries rebuilt🟠Hash is not increment-aware
7Skill pruningKey rules are lost; "HK$0.14" is compressed into "low-priced stock"🔴Irreversible, no numeric protection
8Cascading failure10+ LLM calls in series; error propagation 1-0.95⁸=34%🔴🔴No end-to-end validation

Cascading Failure (System-Level Risk)

In the XSkill pipeline, the LLM is invoked in more than 10 places:

Trajectory summary → Experience extraction → Experience merging → Skill generation → Skill merging
    → Skill streamlining → Task decomposition → Experience retrieval → Experience rewriting → Final reasoning

95% accuracy per step × 10 steps = final accuracy of only 60%. This is the deepest risk in the XSkill architecture and is not discussed in the paper.

Defenses We Should Add

RiskDefense
Hallucinated SkillBacktest after generation; do not release if success rate <80%
Retrieval without thresholdMIN_SIMILARITY = 0.6 + MMR diversity reranking
Over-mergingRetain Diff + 30-day rollback
Task decomposition cascadeCache decomposition results + hard Token budget cap + consistency validation
Skill pruningKeyword whitelist (forced retention for "must not", "must", "prohibited", "%")
Cascading failurePipeline degradation switches + end-to-end quality monitoring

4. XSkill ↔ MemoryHub Mapping

ConceptXSkillMemoryHubAdaptation Recommendations
Skill LibrarySkill Library (LLM generation + merging)AK-SDD SKILL.md (handwritten)Introduce automatic generation + merging
Experience BankExperience Bank (condition-action-emb)Lessons (unstructured)Structure into triples + embeddings
RetrievalTask decomposition + Cosine + rewritingQdrant COSINEAdd task decomposition routing
MergingLLM-based merge + quality scoringDream appendAdd deduplication + quality + elimination
CompactionLLM refine (word_threshold)NoneT5 MEMORY.md compression
CachingMD5 hash + npy persistenceauto_sync full scanAdd incremental hash caching
MultimodalVisual summary + VLMN/A (text only)Not needed

5. Conclusion: What to Borrow and What Not to Borrow

✅ Worth Borrowing (Low Risk, High Value)

  1. Skill three-stage lifecycle, architectural concept (not the fully automated implementation)
  2. Experience two-layer structure, structured lessons as (condition, action, embedding)
  3. Task decomposition routing, best practices for query routers
  4. Hash incremental caching, greatly reduces embedding model calls
  5. Skill slimming mechanism, but requires adding a whitelist and backup

❌ Borrow with Caution (High Risk, Needs Improvement)

  1. Fully automated Skill generation, requires adding a validation loop first
  2. Threshold-free embedding retrieval, requires setting MIN_SIMILARITY first
  3. LLM automatic merging, requires preserving Diff and rollback first
  4. Chained pipeline, requires adding fallback and end-to-end validation first

In one sentence: XSkill's greatest contribution is not any single technique, but the proof that an Agent's continual learning does not require retraining parameters. Structured knowledge extraction + semantic retrieval + Prompt Injection is enough. However, its implementation contains multiple design assumptions of "over-relying on LLM generation without validation," which require quality guardrails to be added in production environments.


This article is based on source code analysis of XSkill-Agent/XSkill v1.0 (Commit: main branch, 2026-05-23). Lines of code: skill_builder.py (232), experience_critique.py (127), experience_retriever.py (~300), trajectory_summary.py (~300), experience_manager.py (651).