Agentic Research

Subagent Isolation Architecture: Reliability Lessons for AI Financial Applications, From the AK-SDD Data Contamination Incident to the Clean Context Design Pattern, A Complete Journey

2026/06/0882 min readUltraClaw閱讀中文原文
TopicsAgent ArchitectureHong Kong stocksOpenClaw

Core proposition: When an AI Agent consecutively analyzes multiple stocks in a single session and the context grows linearly to 50K+ tokens, financial data from earlier stocks remains in the LLM's attention window, and it cannot distinguish which data belongs to 00653 and which belongs to 00928. How does sub-agent isolation solve this problem?
Incident data: 00653's CR fund (HK$368 million) was mixed into the 00928 report · 00928 actually had no CR fund · remediated immediately after discovery
Methodology: Incident record → root cause tracing → isolation architecture design → sessions_spawn implementation → effectiveness validation → extraction of general principles


Preface: A Data Contamination That Should Never Have Happened

In June 2026, while using AK-SDD (Hong Kong stock information search and investigation system) to investigate two companies consecutively, we encountered a seemingly minor but highly destructive problem.

Research targets:

  1. 00653 Bonjour Holdings, beauty retail, property investment
  2. 00928 King International Investment, baijiu sales, health products, money lending

The two companies' businesses are completely different. However, when we received the research report for 00928, we found a puzzling description:

"The company holds CR Business Innovation Investment Fund, with a carrying value of approximately HK$368 million, of which approximately HK$154 million has been impaired..."

This does not belong to 00928. 00928's businesses are baijiu (Diwangchi series), health products (men's health products), and money lending. It has absolutely no CR Business Innovation Investment Fund.

This data came from 00653.

This is not an LLM hallucination. This is a systemic problem of context contamination.


I. Problem Scenario: Context Leakage in Sequential Multi-Stock Research

1.1 AK-SDD Workflow

AK-SDD is Junze Think Tank's Hong Kong stock information search and investigation system, with the following structure:

StageContentToken consumption
Company basic informationName, stock code, industry classification~500
Business descriptionMain business, revenue structure~2,000
Financial analysisRevenue, profit, assets and liabilities~5,000
Shareholding structureMajor shareholders, shareholding percentages~3,000
News and risksRecent announcements, litigation, regulatory matters~4,000-8,000
Total (single stock)~15,000-20,000 tokens

The problem is not with analyzing a single stock. The problem is that when you analyze two stocks consecutively in the same session, the first stock's 15,000-20,000 tokens all remain in the context.

1.2 The Exact Process of Reproducing the Incident

Step 1: Analyze 00653 Bonjour Holdings
  → Agent searches business, financials, property investment
  → Discovers CR Business Innovation Investment Fund
  → Analyzes fund size (HK$368 million), impairment (HK$154 million) in detail
  → Session tokens: ~18,000

Step 2: Start analyzing 00928 King International Investment
  → Session tokens: 18,000 + new searches ~15,000 = ~33,000
  → Two completely independent sets of financial data in the LLM's attention window
  → No structured mechanism marks "which data belongs to which stock"
  → Risk: the LLM may mix the data when generating the report

Step 3: Generate the 00928 report
  → The LLM's output includes a description of the CR fund
  → Cause: 00653's fund data is still in the context, and 00928's asset section does not have a strong enough signal to distinguish it
  → Result: data contamination

1.3 Why the LLM Cannot Tell Them Apart

This is not a matter of the LLM "not being smart enough"; it is an inherent limitation of the attention mechanism.

When the context contains multiple similar blocks of structured data (financial statements, fund descriptions, assets and liabilities), the LLM's attention weights are allocated based on semantic similarity rather than source separation. In plain language:

If you give an LLM the financial data of two companies in the same conversation, it may cite the first company's assets as the second company's assets, because they "look alike" in semantic space.

More critically, the LLM has no built-in concept of "data isolation". It treats the entire context as one continuous stream of information. Unless you explicitly mark the boundaries in the prompt (which we had not done at the time), it will not separate them automatically.


II. Specific Case: CR Fund Migration Path from 00653 to 00928

2.1 Actual Comparison of the Two Companies

Dimension00653 Bonjour Holdings00928 King International Investment
IndustryBeauty product retail + property investmentBaijiu sales + health products + money lending
Main AssetsRetail network, property portfolio, CR FundBaijiu inventory, loan receivables
CR Fund✅ Held, HK$368 million❌ None at all
Revenue StructureRetail store revenue + rental incomeBaijiu sales + interest income
Recent NewsShort-selling report controversyBaijiu brand promotion

The only similarity between the two companies is that both are small-cap Hong Kong-listed companies, and their names contain either "International" or "Holdings". In terms of business, assets, and revenue structure, they have no overlap whatsoever.

2.2 Specific Manifestations of Data Contamination

In the contaminated 00928 report, the description of the CR Fund was inserted into the "Asset Structure" section:

## Asset Structure

The company holds CR Business Innovation Investment Fund, with a carrying value of approximately HK$368 million.
The fund mainly invests in commercial properties in the Asia-Pacific region and has been affected by the recent adjustments in the commercial real estate market,
of which approximately HK$154 million has been impaired. Management stated that it will continue to monitor...

In addition, the company's main assets also include:
- Baijiu inventory (Emperor Pool Series)
- Loan receivables
- Health product inventory

This description is completely identical to 00653 in terms of data, but entirely wrong in context. It was seamlessly woven into 00928's report, with no obvious trace of "grafting."

2.3 Why This Incident Is Especially Dangerous

Traditional LLM hallucination is fabricating data out of thin air, and at least it looks fake. But errors produced by context contamination are citing real data while misattributing it; every number really exists, they simply belong to the wrong company.

The danger of this kind of error lies in:

  1. Extremely high surface credibility, the numbers all "match up"
  2. It requires domain knowledge to detect, one must know that 00928 does not have a CR fund
  3. Routine QA cannot catch it, the format is correct, the data is real, and the logic is internally consistent
  4. It can lead to serious consequences in financial decisions, investors may make judgments based on an incorrect asset structure

III. Root Cause Analysis: Blurred Data Boundaries in the LLM Context Window

3.1 OpenClaw's Session Context Model

OpenClaw's session uses a linear conversation model:

Session Context (tokens grow linearly with conversation length)
│
├─ Message 1: System prompt
├─ Message 2: User: "Analyze 00653"
├─ Message 3: Agent: [Search 00653 business]
├─ Message 4: Agent: [Retrieve 00653 financial data]
├─ Message 5: Agent: [Analyze CR fund, HK$368 million]
│   ... (15,000 tokens)
├─ Message 20: Agent: [00653 report completed]
├─ Message 21: User: "Analyze 00928"          ← At this point context ~18,000 tokens
├─ Message 22: Agent: [Search 00928 business]
├─ Message 23: Agent: [Retrieve 00928 financial data]
│   ... (added 15,000 tokens)
├─ Message 40: Agent: [Generate 00928 report]     ← Context ~33,000 tokens
│   ↑
│   All data for 00653 is still here
│   There is no mechanism to isolate it

Key issue: When Message 21 begins, all of 00653's research data (15,000+ tokens) is still in the context. When the LLM generates the 00928 report, this data is visible and can be referenced.

3.2 Why Simple Prompt Instructions Are Not Enough

We initially tried to solve this by adding an instruction to the prompt:

“Please use only the data of the company currently being queried and do not reference information from companies analyzed previously.”

This instruction works in most cases, but not 100% of the time. When the context becomes long enough (30K+ tokens), the LLM's instruction-following ability degrades. In our tests:

Context LengthPrompt Instruction Success Rate
< 10K tokens99%+
10K-20K tokens~95%
20K-35K tokens~85-90%
> 35K tokens~75-80%

For financial analysis, an 80% success rate means that one out of every five analyses may be wrong, which is completely unacceptable in financial scenarios that require precise data.

3.3 The Real Cause: Data Lacks Structured Attribution Markers

The deeper problem is that in traditional session models, all data is flat. There is no hierarchical structure, no namespace, and no data attribution markers.

Flat model (current state):
["00653 holds CR Fund HK$3.68 hundred million",    ← Attribution: Unmarked
 "00653 revenue HK$5.2 hundred million",             ← Attribution: Unmarked
 "00928 Baijiu sales revenue HK$1.2 hundred million",     ← Attribution: Unmarked
 "00928 money lending interest income HK$0.3 hundred million"]     ← Attribution: Unmarked

Isolation model (target):
[Namespace: 00653]                      [Namespace: 00928]
├─ CR Fund HK$3.68 hundred million                    ├─ Baijiu revenue HK$1.2 hundred million
├─ Revenue HK$5.2 hundred million                        ├─ Money lending interest HK$0.3 hundred million
└─ Beauty retail business                          └─ Health products business

In the first model, the LLM needs to infer data attribution on its own, and when the volume of data is large enough and the similarity is high enough, this inference may be wrong. In the second model, data attribution is structurally enforced, and the LLM cannot reference data across namespaces.


IV. Solution: Design and Implementation of the Subagent Isolation Architecture

4.1 Core Design Philosophy

The essence of the solution is very simple:

Do not analyze two stocks in the same conversation. Create a brand new, fully isolated session for each stock.

In technical implementation, use OpenClaw's sessions_spawn mechanism to create an independent subagent session for each stock research task:

Main Session (Coordination Layer)
│
├─ sessions_spawn: Analyze 00653
│  └─ Sub-agent Session 00653
│     ├─ Search 00653 business
│     ├─ Fetch 00653 financials
│     ├─ Generate 00653 report
│     └─ Return report to main session
│     ✅ Context: contains only 00653 data
│     ✅ No interference from other stock data
│
├─ sessions_spawn: Analyze 00928
│  └─ Sub-agent Session 00928
│     ├─ Search 00928 business
│     ├─ Fetch 00928 financials
│     ├─ Generate 00928 report
│     └─ Return report to main session
│     ✅ Context: contains only 00928 data
│     ✅ Fully isolated, 00653 data is not visible
│
└─ Summary report

4.2 Three-Layer Design of the Isolation Architecture

┌──────────────────────────────────────────────┐
│              Main Agent (Dispatch Layer)                 │
│  Responsibilities: task allocation, result aggregation, quality check              │
│  Context: contains only the summary result of each stock               │
└─────┬────────────────────┬───────────────────┘
      │                    │
      ▼                    ▼
┌──────────────┐   ┌──────────────┐
│ Subagent 00653 │   │ Subagent 00928 │
│ Independent Session  │   │ Independent Session
├──────────────┤   ├──────────────┤
│ Context:      │   │ Context:      │
│ • 00653 Business  │   │ • 00928 Business  │
│ • 00653 Finance  │   │ • 00928 Finance  │
│ • 00653 News  │   │ • 00928 News  │
│              │   │              │
│ ❌ Cannot see:   │   │ ❌ Cannot see:   │
│ any of 00928  │   │ any of 00653  │
│ Data          │   │ Data          │
└──────────────┘   └──────────────┘

First layer: Main Agent scheduling

  • Receive the user's research request
  • Create an independent subagent for each stock
  • Aggregate the reports returned by subagents
  • Perform the final quality check

Second layer: Subagent execution

  • Each subagent has a completely clean new session
  • Receives only one task description and the current stock's data
  • Is not affected by other stock data
  • Returns a structured report after completion

Third layer: Data boundary enforcement

  • The subagent's context is limited to the task description passed in
  • It cannot access the main session's historical messages
  • It cannot access other subagents' data
  • It is automatically cleaned up after the session ends

4.3 Implementation Code

In AK-SDD, the subagent call is encapsulated as a function:

async def analyze_stock_in_subagent(
    stock_code: str,
    stock_name: str,
    query_params: dict
) -> dict:
    """
    Analyze a single stock in an isolated subagent session.

    Each subagent has a completely clean context,
    and will not be polluted by data from other stocks.
    """
    task_prompt = f"""
    Conduct a Hong Kong stock information search investigation on {stock_code} {stock_name}.

    [Isolation Rules: Must Follow]
    1. You may only use data obtained from this search.
    2. Do not cite any external information or data from memory.
    3. Only analyze information related to {stock_code}.
    4. All numerical values must be annotated with source and date.

    [Investigation Scope]
    - Company basic information: industry, main business
    - Financial analysis: revenue, profit, assets and liabilities for the most recent three years
    - Equity structure: major shareholders and shareholding percentages
    - Risk factors: recent major events, litigation, regulatory matters
    - Asset details: properties, investments, intangible assets

    [Output Format]
    Return a structured JSON or Markdown report,
    including source citations for all investigation results.
    """

    # Create isolated subagent session
    result = await sessions_spawn(
        task=task_prompt,
        context="isolated",  # Key: completely isolated context
        tools=["web_search", "web_fetch", "tavily_search"],
        timeout=300  # 5-minute timeout
    )

    return result

For batch research, a parallel mode is used (multiple stocks proceed simultaneously):

async def batch_analyze_stocks(
    stocks: list[tuple[str, str]]  # [(code, name), ...]
) -> dict[str, dict]:
    """
    Analyze multiple stocks in parallel, each stock executed in an isolated subagent.

    Advantages of parallel mode:
    1. Independent session per stock -> no data contamination
    2. Multiple stocks execute simultaneously -> total elapsed time = slowest one
    3. Failure of any subagent does not affect others
    """
    tasks = [
        analyze_stock_in_subagent(code, name, params)
        for code, name in stocks
    ]

    # Execute all subagents in parallel
    results = await asyncio.gather(*tasks, return_exceptions=True)

    # Process results, isolating failures
    output = {}
    for (code, name), result in zip(stocks, results):
        if isinstance(result, Exception):
            output[code] = {
                "status": "error",
                "error": str(result),
                "stock_name": name
            }
        else:
            output[code] = {
                "status": "success",
                "report": result,
                "stock_name": name
            }

    return output

4.4 Technical Details of Isolation Guarantees

Subagent isolation is not "advisory"; it is mandatory:

Isolation DimensionImplementationGuarantee Level
Session ContextEach subagent creates an independent session, context="isolated"✅ Mandatory
Memory/PersistenceSubagents do not load MEMORY.md or daily notes✅ Mandatory
Vector MemorySubagents do not access Qdrant (or are restricted to read-only default slices)✅ Mandatory
File SystemSubagents are confined to a temporary working directory✅ Configurable
Tool AccessOnly search and extraction tools are granted, with no write permissions✅ Mandatory
LifecycleAutomatically cleaned up after the session ends, leaving no residue✅ Automatic

The most critical one is the first dimension, Session Context Isolation. This ensures that subagents cannot see data for other stocks at all, fundamentally eliminating the possibility of data contamination.

V. Effectiveness Validation: Data Accuracy After Sub-agent Isolation

5.1 Test Results After the Fix

We conducted rigorous testing on the fixed system, including continuous analysis of 20 different stock combinations:

Test GroupStock CombinationError Rate Before IsolationError Rate After Isolation
Group 100653 → 00928🟡 1 data contamination incident✅ 0 errors
Group 200005 → 00011🟢 0 errors✅ 0 errors
Group 300700 → 09988🟡 1 incident (PE data mix-up)✅ 0 errors
Group 400175 → 02333🟢 0 errors✅ 0 errors
Group 501928 → 06862🟡 1 incident (revenue data mix-up)✅ 0 errors
...(20 groups in total)3 contamination incidents in total0 contamination incidents in total

Before isolation: 3 of the 20 groups had data contamination (15% contamination rate)

After isolation: 0 of the 20 groups had data contamination (0% contamination rate)

5.2 A More Important Finding: Report Quality Actually Improved

Sub-agent isolation not only eliminated data contamination but also brought an unexpected quality improvement:

Before isolation (shared session):
├─ Average report length: ~2,500 characters
├─ Number of cited sources: 8-12
├─ Analysis depth: Moderate
└─ Reason: The context is crowded, so the LLM tends to simplify its output

After isolation (independent session):
├─ Average report length: ~4,000 characters
├─ Number of cited sources: 15-22
├─ Analysis depth: Significantly improved
└─ Reason: The context is clean, so the LLM can fully expand its analysis

Conclusion: Isolation not only prevented errors but also allowed the LLM to better "focus" on the current task.

5.3 Performance Advantage of Parallel Execution

Sub-agent isolation also brought another benefit: it naturally supports parallel execution:

Execution ModeTotal Time for 3 StocksDescription
Serial (shared session)~18 minutesExecuted sequentially, ~6 minutes per stock
Parallel (sub-agent isolation)~7 minutesThe three stocks execute simultaneously, waiting for the slowest one
Performance improvement2.6xThe more stocks, the more obvious the advantage

VI. General Lessons: Reliability Design Principles for Financial AI Applications

From this incident and the remediation process, we distilled five reliability design principles for financial AI applications:

🔴 Principle 1: Mandatory Data Attribution Tagging

Never let the LLM infer data attribution on its own.

❌ Wrong approach:
  List data for multiple entities in the same prompt,
  relying on the LLM to distinguish them itself

✅ Correct approach:
  Process each entity's data in a separate namespace/session,
  with data attribution enforced by architecture, not inferred by the LLM

Practical checkpoint: In your system, is there any case where the same session handles multiple independent entities? If so, refactor it immediately into an isolation pattern.

🔴 Principle 2: Session Context Minimization

A financial analysis session should not contain irrelevant historical data.

❌ Wrong approach:
  Session context grows linearly with the conversation,
  and all historical data remains in the attention window

✅ Correct approach:
  Use a clean session for each independent analysis task,
  passing summary results only when necessary

Practical checkpoint: Check your session token consumption. If more than 30% of tokens are unrelated to the current task, your context is too large.

🟡 Principle 3: Strong Isolation of Subtasks

Any subtask involving different data sources should be executed in isolation.

🔴 High-risk scenarios (must isolate):
  - Analyzing multiple stocks consecutively
  - Comparing financial data from multiple companies
  - Processing multiple legal documents
  - Reviewing multiple contracts

🟡 Medium-risk scenarios (isolation recommended):
  - Searching multiple topics in parallel
  - Multi-step data transformation
  - Cross-domain analysis

🟢 Low-risk scenarios (session can be shared):
  - In-depth discussion of a single topic
  - Multiple queries against the same data source
  - Interactive Q&A

🟡 Principle 4: Independent Verification Layer

Even when subagent isolation is used, an independent verification step is still required.

Verification strategy:
1. Check whether key data in the report matches the original source
2. Cross-comparison: whether the same data point appears in the correct company's report
3. Business logic check: whether the business description in the report matches the company's known business
4. Anomaly detection: whether data items that do not belong to the current company appear in the report

🟢 Principle 5: Failure Mode Design

Financial AI systems must assume that errors will occur and design failure modes accordingly.

Error scenario         → System response
───────────────────────────────────
Data contamination     → Isolate + regenerate
Hallucination (fabricated data) → Source verification + refuse output
Search failure         → Fall back to known data + mark as incomplete
Timeout                → Return partial results + mark as unfinished
API error              → Retry 3 times + final fallback

Key: Always return "flagged partial results" rather than "apparently complete but erroneous results." Incomplete reports can be identified and supplemented; erroneous reports can mislead decisions.


VII. Lessons and Reflections

7.1 What This Incident Taught Us

  1. An LLM is not a database. Do not expect it to maintain data isolation like a database. Any data in the context may be referenced, whether it should be or not.

  2. Prompt instructions are not a security boundary. Saying "do not reference previous data" in a prompt is a soft constraint, not hard isolation. When the context is large enough, soft constraints fail.

  3. Architectural isolation is more reliable than behavioral-layer instructions. Sub-agent session isolation is an architectural-layer solution; it fundamentally eliminates the possibility of data contamination rather than merely reducing its probability.

  4. Reliability requirements in financial scenarios are far higher than in general applications. In general chat, a 15% data contamination rate may just be "imperfect"; in financial analysis, it is "unacceptable".

7.2 Comparison with Claw Code's Smart Compaction

When we previously studied Claw Code, we conducted an in-depth analysis of its Smart Session Compaction mechanism, which controls context length by compressing old messages into structured summaries.

But this incident shows: Compression cannot solve the data attribution problem. Even if you compress all of 00653's data into a summary, this summary is still in the context and may still be referenced by the 00928 report.

Compression addresses the "token limit" problem; isolation addresses the "data contamination" problem. The two are complementary, but neither can replace the other.

7.3 Future Directions

Based on this experience, we plan to introduce the following improvements in AK-SDD and the broader financial analysis toolchain:

  1. Automatic isolation routing: the system automatically identifies scenarios that "require isolation", with no manual configuration required
  2. Cross-session result validation: after a sub-agent result is returned, an independent verification sub-agent is automatically triggered
  3. Tiered isolation levels: lightweight isolation (isolates context only), standard isolation (context + tools), strict isolation (context + tools + file system)
  4. Sub-agent execution audit logs: fully record each sub-agent's input, output, and tool calls for post-hoc auditing

VIII. Conclusion

This data contamination incident is a classic case of a problem that looks small on the surface but is actually very serious.

On the surface, it was just an AI output error that placed Company A's data in Company B's report. But digging deeper, it exposed a systemic flaw in LLM applications in financial scenarios: the lack of data attribution enforcement.

The solution is not to write better prompts, not to have more model parameters, but to enforce data isolation at the architectural level. The sub-agent isolation architecture, which creates a completely clean session for each independent analysis task, is a simple, effective, and reusable pattern.

In financial AI, reliability is not a feature but a foundational requirement. A 1% error rate can lead to a 100% loss of trust. The Clean Context design pattern is not an optional optimization; it is the minimum standard for financial AI applications.


This article is based on a real production incident and remediation process of the AK-SDD system on 2026-06-08.
Affected stocks: 00653.HK (Bonjour Holdings), 00928.HK (King International Investment)
Remediation: sessions_spawn sub-agent isolation architecture
Validation results: 20 test groups, 0% data contamination rate
Related systems: OpenClaw Agent Platform · AK-SDD Hong Kong Stock Information Search and Investigation System