Agentic Research

MemoryHub v2.0 System Architecture In-Depth Analysis: From Capture Daemon to MCP Real-Time Memory Capture

2026/05/2128 min readBryan Chan閱讀中文原文
TopicsMemoryHubAI MemoryVector DatabaseQdrantMCP

Introduction: Why Do We Need MemoryHub?

What is the biggest pain point for AI Agents? It is not that model capabilities are insufficient, nor that there are not enough tools; rather, it is that every time they wake up, everything is blank.

Whether OpenClaw, Claude Code, Hermes, or DeepSeek TUI, every AI platform faces the same fundamental problem: after a Session ends, all conversation context disappears with it. Even if there is a filesystem Session Log (JSONL), that is still raw data: it cannot perform semantic search, cannot associate across Sessions, and cannot integrate across platforms.

MemoryHub was designed precisely to solve this problem: a cross-platform persistent memory enhancement system.


1. System Overview: Dual-Mode Capture Engine

MemoryHub's core architecture is built around a unified Capture Daemon, supporting two complementary capture modes:

AI Platforms (OpenClaw / Hermes / DeepSeek / Claude Code)
        │                          │
        │ MODE B: File scan        │ MODE A: MCP real-time
│ (incremental scan every 5 minutes)       │ (real-time via /hook)
        ▼                          ▼
   Capture Daemon (:3872) ──── Qdrant Vector DB (:6333)
        │                          │
   Dashboard (:3872)          MCP Server (stdio)
   • Live feed                • mem_save / mem_search
   • 24h/7d charts            • mem_stats / capture_send
   • Global search             • 4 per-platform collections
   • Collection stats

Mode A-MCP Real-Time Capture

When an AI Agent calls an MCP tool (such as mem_save), the Capture Daemon receives captured data in real time through the /hook endpoint. This is a zero-latency capture path, suitable for important conversation snippets that need to be saved immediately.

Mode B: File System Incremental Scan

Every 5 minutes, the Daemon scans the Session directories of the four platforms and, through an offset mechanism, reads only newly added content to avoid duplicate processing. This is a passive but complete capture path, ensuring that no conversation is missed.

The Two Are Complementary

  • Mode A is fast but incomplete (it depends on the Agent actively calling MCP)
  • Mode B is slow but complete (it passively scans all Session files)
  • Combining the two = a complete and real-time memory capture system

2. Core Components in Detail

2.1 Capture Daemon

A unified daemon that integrates four major functions:

FunctionDescriptionTechnical Details
Mode A IngestionReal-time capture via POST /hookHTTP stdlib server, DAEMON_HOOK_URL
Mode B ScanningIncremental scan every 5 minutesFile offset tracking, capture_offsets.json
DashboardReal-time monitoring via Web UI6 API endpoints, 24h/7d charts
State ManagementIn-memory state + persistencecapture_daemon_state.json

2.2 Four-Platform Parser

Each AI platform uses a different conversation format, and MemoryHub includes a corresponding Parser:

PlatformParserFormatSpecial Handling
OpenClaw_scan_fileJSONLtype="message" + content blocks extraction
Hermes_scan_fileJSONLLine-by-line role/content parsing
DeepSeek TUI_scan_deepseek_checkpointJSONmessages array, tool_result/tool_use extraction
Claude Code_scan_fileJSONL6 directory scans, 42 tracked files

Key Design Decision: The tool_result and tool_use blocks in DeepSeek TUI are contained in the messages array and require special handling to extract content correctly. This was one of the core bugs fixed in v2.0.

2.3 MCP Server

MCP Server based on the JSON-RPC stdio protocol, exposing 6 tools externally:

ToolFunctionUse Case
mem_saveSave memory to Qdrant + notify daemon"Remember this customer preference"
mem_searchSemantic similarity search"Where is last month's DD report"
mem_statsCollection statistics"How many memories do we have"
mem_list_collectionsList all CollectionsBrowse memory categories
mem_deleteDelete memory entriesClean up outdated information
capture_sendExplicit conversation captureManually trigger capture

2.4 Dashboard Web UI

Built-in HTTP Server (port 3872), no additional frontend framework required:

AreaDisplayed Content
Stats BarToday's captures, MCP vs scan ratio, scan cycle, Qdrant points, uptime
Platform CardsCapture count per platform, tracked file count, progress bar
ChartsHourly trend chart over 24 hours or daily trend chart over 7 days
Live FeedReal-time stream of captured content
SearchFull-text search across all captured memories

3. Data Flow: From Conversation to Vector Memory

The complete data pipeline consists of 7 steps:

1. User converses with AI Agent
2. Agent Session file updated (JSONL/JSON)
3. Daemon Mode B scanner captures new lines (delay ≤5 minutes)
   or Agent invokes MCP tool → Mode A captures in real time
4. Content parsing, role extraction (user/assistant)
5. Save to ~/.memory-hub/captured/<platform>/YYYY/MM/DD.jsonl
6. Generate vector embeddings via BGE-m3 model (local inference)
7. Store in Qdrant Collection for semantic search

Storage Layout

~/.memory-hub/
├── capture_daemon_state.json   # Daemon process real-time state
├── capture_offsets.json         # Read offset for each file
├── captured/                    # All captured content (data source)
│   ├── openclaw/YYYY/MM/DD.jsonl
│   ├── hermes/YYYY/MM/DD.jsonl
│   ├── deepseek/YYYY/MM/DD.jsonl
│   └── claude/YYYY/MM/DD.jsonl
├── memories/                    # Memories saved by MCP
│   └── <uuid>.json
└── hooks/                       # Hook logs

4. Three-Layer Deduplication Mechanism

Duplicate capture is a common problem in memory systems. MemoryHub has designed a three-layer deduplication safeguard:

LayerMechanismDescription
Layer AUUID5 content hashGenerates a deterministic UUID based on content; identical content = identical ID
Layer BOffset incremental scanReads only new content after the last scan position
Layer CCross-session semantic similarityMemories with similarity >85% are marked as duplicates

5. Ten-Year Lifecycle Design

MemoryHub is not just "today's memory"; it has designed a complete time decay and compression strategy:

Daily raw logs → Weekly summaries → Monthly summaries → Annual retrospectives → Ten-year archive
     ↓              ↓                 ↓                  ↓                    ↓
  Never delete    Key events        Milestones         Annual review        Long-term trends

Core principle: Never delete, only compress. Raw data is retained permanently, and higher levels are simply more refined views.

Three-Layer Backup

LayerFrequencyStorage Location
hourlyHourly~/.memory-hub/backups/hourly/
dailyDaily~/.memory-hub/backups/daily/
weeklyWeekly~/.memory-hub/backups/weekly/

6. Technology Stack

ComponentTechnology ChoiceRationale
LanguagePython 3.9+Rich ecosystem, cross-platform
Vector DatabaseQdrant (Docker)High performance, Rust implementation, REST API
Embedding ModelBGE-m3 (local)Multilingual support, sentence-transformers ecosystem
Full-text IndexSQLiteZero configuration, built-in Python support
APIHTTP (stdlib)Zero dependencies, lightweight
MCPJSON-RPC stdioStandard protocol, multi-platform compatible
Auto-startlaunchd (macOS) / systemd (Linux)Native to the system

7. Relationship to Existing Memory Systems

MemoryHub is not meant to replace OpenClaw's agentmemory or Auto-Dream, but rather serves as a fourth-layer enhancement:

Layer 1: Session Log (JSONL)         ← automatically saved by the system
Layer 2: Daily Log (Markdown)        ← Structured log
Layer 3: MEMORY.md (long-term memory)        ← Auto-Dream integration
Layer 4: MemoryHub (Vector Semantics)        ← Cross-platform search and association  🆕

The fourth layer provides capabilities that the first three layers cannot achieve:

  • Semantic search: not grep keywords, but understanding intent
  • Cross-platform association: OpenClaw conversations can be associated with Claude Code conversations
  • Time travel: quickly locate "discussions about a project from three months ago"

Conclusion

MemoryHub v2.0 evolved from an internal script into a complete Python package, going through iterations from v1.2 (Hook upgrade) to v2.0 (CLI + packaging). It solves the most fundamental memory problem for AI Agents: from a blank state on every wake-up to persistent semantic memory across time and platforms.

Future evolution directions include: multi-user support, cloud synchronization, and deeper integration with agentmemory.


This article is based on the actual architecture of MemoryHub v2.0.0 (released 2026-05-20). GitHub: Bryan-cmf/memory-hub