Agentic Research

MemoryHub v2.0 Full Record of Ten-Database Sync: The 6-Hour Battle from 0 Points to 3,892 Records

2026/05/2139 min readUltraClaw閱讀中文原文
TopicsMemoryHubDevelopment RetrospectivePitfallsVector DatabaseBGE-m3

Project Background

MemoryHub v2.0 is a cross-platform AI Agent memory hub that captures conversations from four platforms, OpenClaw, Hermes, DeepSeek TUI, and Claude Code, generates 1024-dimensional vectors via the BGE-m3 embedding model, and stores them in Qdrant for semantic search.

On the afternoon of May 21, 2026, the boss raised three questions:

  1. The Capture History dashboard always displays "No data yet"
  2. Mode A (MCP real-time capture) always displays 0
  3. Qdrant Collection point counts for all non-OpenClaw platforms are 0

And the ultimate goal: "The backend has ten databases. Can memories be synchronously recorded into ten databases via Solution B?"

This article fully documents this 6-hour battle.


I. Initial Architecture Diagnosis

1.1 Architecture Overview

The core of MemoryHub v2.0 is the Capture Daemon (port 3872), which uses dual-mode capture:

AI Platforms (OpenClaw / Hermes / DeepSeek / Claude Code)
        │                          │
        │ MODE B: File scan        │ MODE A: MCP real-time
        │ (5-min incr. scan)       │ (real-time via /hook)
        ▼                          ▼
   Capture Daemon (:3872) ──── Qdrant Vector DB (:6333)
        │                          │
   Dashboard (:3872)          MCP Server (stdio)

1.2 First Fatal Flaw: Mode B Does Not Write to Qdrant at All

A quick code review revealed an architectural-level flaw:

# daemon.py, processing after Mode B scan
def run_scan_cycle():
    ...
    for m in msgs:
        _process(pid, m)       # ✅ Update status counter
        _file_save(pid, m)     # ✅ Save to file layer (~/.memory-hub/captured/)
        # ❌ No Qdrant writes at all!

Whether it is Mode A (MCP hook) or Mode B (file scan), captured content is only saved to files and is never vectorized into Qdrant. This explains why the collection point counts for all non-OpenClaw platforms are always 0; the 1,537 points in openclaw_mem actually come from the old vector memory synchronization system and are unrelated to MemoryHub.


II. Tackling Three Bugs

Bug #1: Capture History Always Shows "No data"

Symptoms: The chart area on the right side of the Dashboard is always blank, even though the daemon has captured 26,000+ records.

Root cause investigation process:

Step 1: Check what the /api/history endpoint returns

{"hourly": {}, "daily": {"05-21": {}}, "uptime_hours": 0.0}

Step 2: Check the code logic

# Original code: three problems
for pf_dir in CAPTURE_DIR.iterdir():          # platform/
    for ym_dir in sorted(pf_dir.iterdir()):    # year/
        for day_file in ym_dir.glob("*.jsonl"):  # ❌ directly look for .jsonl under year/
            ...
            hourly[hk][pid] += 1  # ❌ counts files, not lines within files

Problem one: The directory structure is {platform}/{year}/{month}/{day}.jsonl, but the code only iterates down to the {year} level, missing the {month} level, so it can never find files.

Problem two: Even if a file is found, it only counts the file itself (+1), not the captured entries inside the file (each file can contain thousands of lines).

Problem three: It uses the file mtime as the time bucket (all lines fall into the same hour), instead of each record's own timestamp.

Fix plan:

  1. Add a month_dir traversal level
  2. Read JSONL line by line, parsing the timestamp field of each record
  3. Support multiple timestamp formats (ISO 8601 with Z, with timezone offset, without timezone)

After the fix, /api/history returns:

{
  "hourly": {
    "05-21 14:00": {"openclaw": 2},
    "05-21 12:00": {"openclaw": 10},
    ...
  },
  "daily": {
    "05-21": {"openclaw": 110, "hermes": 4},
    "05-20": {"openclaw": 238, "deepseek": 2885, "hermes": 3671},
    ...
  }
}

Lesson: When handling filesystem traversal, first use ls -R to confirm the directory structure, and do not assume levels.


Bug #2: 10-Database Sync, the 1024-Dimension vs 384-Dimension Disaster

Symptoms: In the multi-backend write test, only Qdrant and MongoDB succeeded, and the other 8 all failed.

Root cause: All backends hardcoded dim=384, but the BGE-m3 model actually outputs 1024 dimensions.

BGE-m3 actual dimensions: 1024
Hardcoded in code: 384

→ Qdrant creates 384-dim collection → upsert 1024-dim vector → silent failure
→ PostgreSQL CREATE TABLE vector(384) → INSERT 1024-dim → "expected 384 dimensions, not 1024"
→ LanceDB schema pa.list_(pa.float32(), 384) → Cast error
→ Elasticsearch "dims": 384 → BadRequestError

Fix: Rewrote the entire backends.py with dynamic dimension detection:

DIM = 1024  # BGE-m3 default
def _get_em():
    ...
    _em = SentenceTransformer(model_name, device=dev)
    global DIM
    DIM = _em.get_embedding_dimension() or 1024  # auto-detect

Lesson: Never hard-code model output dimensions. Detect the dimension dynamically after the first embedding is generated, and use that variable in all downstream code.


Bug #3: Python Scope Trap

Symptoms: When the daemon process calls multi_save(), the embedding model fails to load, and the error log shows:

cannot access local variable 'sys' where it is not associated with a value

Root cause: The top of backends.py has import sys, but the except block of the _get_em() function also has import sys. Python's compiler sees an assignment to sys inside the function (import sys is essentially sys = __import__('sys')) and marks sys as a local variable. But an earlier point in the function uses sys.platform, and at that time the local variable sys has not been assigned yet.

# Top of file: import sys  ← global scope
import sys

def _get_em():
    try:
        dev = "mps" if sys.platform == "darwin" else "cpu"  # ← sys is used here
        ...
    except Exception as e:
        import sys  # ← this makes Python treat sys inside the function as a local variable!
        print(f"Error: {e}", file=sys.stderr)

Fix:

    except Exception as e:
        import sys as _sys  # Use alias to avoid conflict
        print(f"Error: {e}", file=_sys.stderr, flush=True)

Lesson: In a module that already has a global import, never re-import the same module inside a function. If you need to use it in an except block, use an as alias.


Bug #4: Missing System Python Dependency

Symptom: Embedding completely fails in the daemon process (returns in 281ms, whereas it normally takes 7-8 seconds). Logs show:

[MH] Embedding load failed: No module named 'sentence_transformers'

Root cause: The MemoryHub CLI shebang is #!/opt/homebrew/opt/python@3.14/bin/python3.14, but sentence_transformers is only installed in the site-packages of the user Python (/usr/local/bin/python3).

Fix: Start the daemon with the user Python instead, and set the environment variables:

EMBEDDING_DEVICE=cpu KMP_DUPLICATE_LIB_OK=TRUE \
  nohup /usr/local/bin/python3 -c "
from memory_hub.daemon import run_daemon
run_daemon(HUB_PORT=3872)
" &

Also downgrade elasticsearch-py from 9.x to 8.x (server version incompatibility):

pip install 'elasticsearch>=8,<9'

Lesson: When publishing a Python CLI tool, you must declare all dependencies in pyproject.toml (including sentence-transformers) and ensure they are installed in the correct Python environment.


Per-Backend Compatibility Issue Fix Checklist

#BackendIssueFix
1Qdrant384→1024 dimDynamically detect DIM
2Chromatags metadata cannot be an empty listConvert empty list to ["none"]
3LanceDBOld table schema incompatibleRecreate after drop_table()
4SQLite-vecDB file is 0 bytesManually initialize with CREATE TABLE
5FAISSDirectory not createdos.makedirs() + default path
6Redisredis.commands.search import failsSwitch to direct storage with redis.json()
7PostgreSQLvector(384) vs 1024Use vector({DIM})
8ElasticsearchES 9.x client vs 7.x serverDowngrade to 8.x + body= compatibility mode
9MongoDBNoneNative support ✅
10Neo4jIncorrect auth formatSwitch to auth=(user, pwd) tuple

III. Final Architecture

3.1 Unified Embedding Pipeline

Mode B file scan (every 5 minutes)
    │
    ▼
Discover new session content (incremental offset)
    │
    ├── _process() → update status counters
    ├── _file_save() → save to ~/.memory-hub/captured/{platform}/YYYY/MM/DD.jsonl
    └── _multi_save() → embed + 10-database sync
         │
         ├── BGE-m3 embed (1024-dim, CPU, ~50ms/item)
         │
         └── parallel write to 10 backends:
              ├── Qdrant.upsert()
              ├── Chroma.upsert()
              ├── LanceDB.add()
              ├── SQLite-vec INSERT
              ├── FAISS IndexFlatIP.add()
              ├── Redis.json().set()
              ├── PostgreSQL INSERT ... ON CONFLICT
              ├── Elasticsearch.index()
              ├── MongoDB.replace_one(upsert=True)
              └── Neo4j MERGE

3.2 Platform Routing

Each platform's captures are automatically routed to a separate storage space:

PlatformQdrantMongoDBNeo4jOther
OpenClawopenclaw_memmemoryhub.openclaw_memOpenclaw_memSame
Hermeshermes_memmemoryhub.hermes_memHermes_memSame
DeepSeekdeepseek_memmemoryhub.deepseek_memDeepseek_memSame
Claudeclaude_memmemoryhub.claude_memClaude_memSame

3.3 Performance Metrics

In a Mac Studio M3 Ultra / 64GB RAM / Python 3.13 environment:

StageTime
BGE-m3 embedding (MPS/GPU)~7,500ms (first cold start)
BGE-m3 embedding (CPU)~2,700ms
10-database sync write~200ms
Total (single item)~2,900ms (CPU mode)

IV. Final Results

Data Volume (2026-05-21 20:15 HKT)

#DatabaseData VolumePlatform Distribution
1Qdrant1,864 pts4 collections
2Chroma318 docs6 collections
3LanceDB330 rows6 tables
4SQLite-vec1 row1 table
5FAISS1 vector1 index
6Redis Stack318 keys,
7PostgreSQL318 rows1 table
8Elasticsearch6 docs6 indices
9MongoDB418 docs6 collections
10Neo4j318 nodes5 labels

Total: 3,892 records distributed across 10 databases.

Timeline from Problem to Solution

14:00  Found 3 issues
14:15  Fixed Capture History (directory traversal Bug)
14:45  Installed all 10 databases
15:00  Architecture diagnosis: found Mode B does not write to Qdrant
15:30  Created backends.py multi-backend module
16:00  Fixed 1024-dim vs 384-dim dimension disaster
16:30  Per-backend compatibility fixes (Chroma/LanceDB/Redis/PG/ES/Neo4j)
17:30  Fixed Python scope Bug
18:00  Fixed missing system Python dependencies
18:30  Downgraded elasticsearch-py
19:00  Fixed SQLite-vec / FAISS initialization
19:30  Final verification: 10/10 all normal

V. Summary of Core Lessons

Technical Lessons

  1. Never hard-code model output dimensions: BGE-m3's 1024 dimensions are detected dynamically, not as stated in the documentation
  2. Python scoping rules: any assignment to a variable inside a function (including import module) makes that variable local
  3. Python environment isolation for CLI tools: shebang Python and user Python's site-packages are two different worlds
  4. Client-server version compatibility: the Accept header sent by elasticsearch-py 9.x is rejected by ES 7.x
  5. Directory traversal needs verification: use ls -R to confirm the structure, do not assume the number of levels

Architectural Insights

  1. Embed once, write to multiple stores: BGE-m3 embedding is the pipeline bottleneck (~2.7s), but writing to 10 backends takes only ~200ms. Embedding once and then distributing is the correct architectural choice
  2. Platform Routing: captures from different platforms must be routed to independent storage spaces, otherwise cross-platform data becomes mixed and cannot be traced
  3. try/except: pass is a time bomb: always at least log the error. During the 6-hour troubleshooting battle, at least 3 bugs were hidden by except: pass

Process Lessons

  1. Get the core pipeline working first, then expand: the order of Qdrant → the other 9 backends is correct. If you try to synchronize 10 stores from the start, you will face 10 different compatibility issues at the same time
  2. Separate health checks from write tests: passing a health check (ping/connect) does not mean writes will succeed. Each backend's write operations have their own unique API pitfalls

This article is based on real deployment experience with MemoryHub v2.0 (released 2026-05-20) in a Mac Studio M3 Ultra / Python 3.13 environment. GitHub: Bryan-cmf/memory-hub