Agentic Research

A Lightweight Overhaul of MemoryHub in Practice: Cutting from 10 Databases to 5, and RAM from 6.2GB to 500MB

2026/05/2331 min readUltraClaw閱讀中文原文
TopicsMemoryHubBGE-m3Vector DatabaseAI Memory

Core proposition: If this AI memory system is to be offered to ordinary computers in millions of households (8-16GB RAM), which databases should stay? Which should go? Test environment: Mac Studio M3 Ultra · 512GB RAM → re-evaluated with an 8GB laptop as the target Methodology: performance dissection → value-for-money scoring → choosing 1 embedding model out of 5 → downgrade execution


Introduction: From a Mac Studio to an Ordinary Laptop

Our process of building MemoryHub on a Mac Studio (512GB RAM, 32-core M3 Ultra) went quite smoothly. All ten databases running? No problem. A 9.1GB BGE-m3 model permanently resident in memory? Imperceptible.

Until we asked a question: what if this system had to be installed on an ordinary 8GB RAM laptop?

The answer was unsettling. A total RAM footprint of 6.2GB is 1.2% of the system on a Mac Studio; on an 8GB laptop it is 78%, before even counting the OS, browser, or Docker Desktop.

We decided to do a thorough lightweight overhaul: starting with a performance dissection, judging the value for money of each database one by one, and finally cutting 10 databases down to 5 and RAM from 6.2GB to 500MB.


1. Performance Dissection: How Much Does Each Database Actually Consume?

Hardware Baseline

ItemSpecification
CPUApple M3 Ultra · 32 cores
RAM512 GB
Docker engineColima · single container limited to 3.814 GiB

Actual Resource Usage of the Ten Databases

DatabaseRAM (idle)Docker?StatusUnique CapabilityActually Used?
🧠 Qdrant17 MB✅🟢 NormalSemantic similarity search✅ Core function
🔍 FAISS~200 MB*❌ Embedded🟢 Normal3ms brute-force search🟡 Only used to store vectors
🗃️ SQLite-vec<50 MB*❌ Embedded🟢 NormalVector + SQL + full-text in one🟡 Only used to store vectors
📦 Chroma~300 MB*❌ Embedded🟢 NormalMetadata filtering🟡 Backup
🪶 LanceDB~300 MB*❌ Embedded🟢 NormalColumnar analytics🟡 Only used to store vectors
🔗 Neo4j507 MB✅🟢 NormalCypher graph queries🔴 Zero usage
🔎 Elasticsearch512 MB+✅🔴 Died from OOMFull-text search🔴 Never survived
🍃 MongoDB155 MB✅🟢 NormalDocument queries🔴 Only used to store vectors
⚡ Redis64 MB✅🟢 NormalSub-millisecond cache🟡 Query frequency too low
🐘 PostgreSQL67 MB✅🟢 NormalSQL JOIN🔴 Only used to store vectors

* Embedded databases share the daemon process's 5.4GB RSS (of which the BGE-m3 model accounts for 3-4GB)

The Resource Pyramid

5,400 MB  capture_daemon (including BGE-m3 model)
       /
  507 MB  Neo4j (JVM)
  155 MB  MongoDB
   67 MB  PostgreSQL
   64 MB  Redis
   17 MB  Qdrant
0 MB  ES (dead)
           ─────────────
~6,200 MB  Total (1.2% of Mac Studio, 78% of an 8GB laptop)

2. The Value-for-Money Trial: Who Stays, Who Goes

We used a simple formula to evaluate:

Cost-effectiveness = unique value provided ÷ resource consumption (RAM + CPU + disk)

🔴 Tier 3: Zero Value for Money (delete immediately)

Neo4j · Score 0/10

What did 507 MB of RAM buy? Zero Cypher queries, zero nodes imported, zero relationship edges. It idled for 45 hours.

Alternative: Python NetworkX (5MB) + a JSON edge list covers the functionality completely.

Elasticsearch · Score 0/10

Never ran successfully for more than 30 minutes. Insufficient JVM heap → terminated by the OOM Killer → repeat. On an 8GB laptop, starting the JVM alone consumes 512MB+, an instant death sentence.

Alternative: SQLite FTS5 (built into the Python standard library, 0MB extra) + jieba Chinese segmentation.

MongoDB · Score 1/10

155 MB RAM + 277 GB of disk writes (WiredTiger write amplification). But we only used it to store vectors: document queries and aggregation pipelines all sat idle. 277GB of writes is a hidden killer on a laptop with a 256GB SSD.

Alternative: JSON files or TinyDB (a Python library, <1MB).


🟡 Tier 2: Marginal Benefit (downgrade)

Redis · Score 3/10

64MB RAM for caching, but a memory system has an extremely low query frequency (1-2 times per minute), with a Cache Hit Rate of <5%.

Downgrade: Python functools.lru_cache (0MB extra).

PostgreSQL · Score 4/10

67MB, with functionality highly overlapping SQLite-vec (both do vector + SQL). Since SQLite-vec is already in-process, a second SQL engine is unnecessary.

Downgrade: merge into SQLite-vec.


🟢 Tier 1: High Value for Money (keep)

Qdrant · Score 8/10

17 MB. The most frugal of the ten databases. Native Rust, zero GC, COSINE search in 66ms. This is the heart of semantic memory, irreplaceable.

FAISS · Score 9/10

Embedded (zero separate process), pure C++, the 3ms brute-force search ceiling. No GPU needed.

SQLite-vec · Score 9/10

One library does vector search + SQL structured queries + FTS5 full-text search at the same time. Bundled with the Python standard library. It alone does the job of PostgreSQL + Elasticsearch + MongoDB combined.


3. Choosing 1 Embedding Model Out of 5: The Real Bottleneck Is Here

All of the discussion above assumed continuing to use BGE-m3. But BGE-m3 (9.1GB on disk, 3-4GB RAM) is disastrous for an ordinary computer.

We downloaded 4 lightweight alternatives and ran a complete comparison of performance and retrieval quality against BGE-m3:

Performance Comparison

ModelDiskRAM (est.)DimensionsSingle-item LatencyBatch Rate
BGE-m3 (current)9,130 MB3,500 MB1024619 ms169 items/s
bge-base-zh1,638 MB655 MB76852 ms1,078 items/s
all-mpnet-base877 MB351 MB76863 ms397 items/s
bge-small-zh 🏆193 MB77 MB51290 ms241 items/s
all-MiniLM-L6183 MB73 MB38432 ms308 items/s

Chinese Retrieval Quality (7-question field test)

Querybge-small-zhMiniLM-L6
What does the boss want me to do every day?✅ Top-1✅ Top-1
What cooperation does HKOW have?❌ Top-3✅ Top-1
How is HKEX disclosure data obtained?✅ Top-1❌
What are the rules for deploying a website?✅ Top-1❌
Why did Elasticsearch die?✅ Top-1✅ Top-1
Mandatory rules for memory writes?✅ Top-1❌
Cellfie Global financing?✅ Top-1✅ Top-1
Accuracy86% (6/7)57% (4/7)

The BGE-m3 GPU Trap

An overlooked finding: BGE-m3's single-item embedding latency is 619ms, but with a batch of 50 it drops to 6ms/item, a 100x difference. The root cause is that the GPU kernel launch on Apple Silicon MPS has a fixed 500-600ms overhead, and for single-item embedding this overhead accounts for 97% of total latency.

🔬 This means that if you use BGE-m3 for real-time memory queries (embedding 1 new piece of content at a time), the GPU spends 97% of its time "launching" rather than "computing". A lightweight model (bge-small-zh) has a single-item latency of only 90ms, and CPU inference has no GPU kernel overhead.


4. The Final Trimming Result

Cut (5)

❌ Neo4j         507 MB  → NetworkX + JSON (5 MB)
❌ Elasticsearch  512 MB  → SQLite FTS5 + jieba (<10 MB)
❌ MongoDB        155 MB  → JSON / TinyDB (<5 MB)
❌ Redis           64 MB  → functools.lru_cache (0 MB)
❌ PostgreSQL      67 MB  → Merged into SQLite-vec (0 MB added)

Kept (5)

✅ Qdrant      17 MB  Docker   Semantic search (core, irreplaceable)
✅ FAISS       28 MB  embedded   3ms fast recall (speed ceiling)
✅ SQLite-vec  <50 MB  embedded   vector+SQL+full-text all-in-one
✅ Chroma     ~200 MB  embedded   fallback semantic search
✅ LanceDB    ~200 MB  embedded   backup analysis

Model Switch

BGE-m3 (9.1 GB / 3.5 GB RAM / 1024d)
         ↓
bge-small-zh (193 MB / 77 MB RAM / 512d)

Before-and-After Resource Comparison

MetricBeforeAfterReduction
Number of databases105-50%
Docker containers61-83%
Docker RAM1,200 MB17 MB-99%
Daemon RSS5,400 MB~350 MB-94%
Embedding model RAM3,500 MB77 MB-98%
Total RAM~6,200 MB~500 MB-92%

Ordinary Computer Fit

ConfigurationBeforeAfter
Mac Studio 512GB1.2% (imperceptible)0.1%
16GB laptop39% (barely)3% (comfortable)
8GB laptop78% (impossible)6% (usable)
4GB Raspberry PiImpossiblePossible (only Qdrant+FAISS+SQLite, <200MB)

5. Insights and Lessons

Insight 1: Most "Essential" Components Are Actually Unnecessary

Neo4j, Elasticsearch, and MongoDB are all excellent databases, in the scenarios they excel at. But in the scenario of a personal AI memory system, their unique capabilities were never invoked; only their resource consumption was real.

"Choose tools by scenario, not by reputation."

Insight 2: Embedded > Docker Containers

FAISS, SQLite-vec, Chroma, and LanceDB are all embedded: zero extra processes, zero network latency, zero Docker overhead. The lightweight MemoryHub keeps only one Docker component, Qdrant (17MB).

Insight 3: The Size Gap Between Embedding Models Is Far Larger Than You Think

The retrieval quality gap between BGE-m3 and bge-small-zh is about 10%, but the resource consumption gap is 45x. For a personal memory system (a few thousand records), a 384-512 dimensional lightweight model is entirely sufficient.

Insight 4: A GPU Is Not Necessarily a Plus for Embedding Inference

BGE-m3's kernel launch overhead on Apple Silicon MPS (500-600ms) leaves single-item embedding with only 3% GPU utilization. A lightweight model on CPU inference is instead better suited to real-time scenarios.


6. Future Directions

  1. Switchable model design: users can choose a model based on their hardware (lightweight → standard → high-precision)
  2. Embedded-first: Qdrant is currently the only component requiring Docker; in the future, LanceDB/Chroma could fully replace it, achieving zero-Docker deployment
  3. A one-click install script: pip install memoryhub → automatically selects the best model and backend configuration

In one sentence: cutting from 10 databases to 5, swapping BGE-m3 for bge-small-zh, and dropping RAM from 6.2GB to 500MB (a 92% reduction), at a loss of only ~10% in retrieval quality. This is not a trade-off; it is engineering judgment.