Agentic Research

Complete Comparison of Top 10 Vector and Full-Text Databases: MemoryHub Backend Storage Selection Guide

2026/05/2147 min readBryan Chan閱讀中文原文
TopicsMemoryHubVector DatabaseQdrant

Introduction

The MemoryHub v2.0 installer provides 10 optional database backends, ranging from lightweight embedded solutions to enterprise-grade distributed clusters. Each database has its own unique design philosophy, performance characteristics, and suitable use cases.

This article is based on actual deployment experience with MemoryHub on Mac Studio (M3 Ultra) and provides a comprehensive evaluation of these 10 databases. We compare them across six dimensions: deployment difficulty, query performance, resource consumption, scalability, Chinese language support, and suitability for AI memory scenarios.


Evaluation Summary Table

#DatabaseTypeDeployment MethodSuitable Use CaseRating
1QdrantDedicated vector databaseDockerProduction workhorse⭐⭐⭐⭐⭐
2ChromaDedicated vector databasepip installRapid prototyping⭐⭐⭐⭐
3LanceDBEmbedded vector databasepip installMobile/edge⭐⭐⭐⭐
4SQLite-vecSQLite vector extensionpip installLightweight standalone⭐⭐⭐
5FAISSVector search librarypip installResearch/GPU⭐⭐⭐
6Redis StackIn-memory database + vectorDockerHigh concurrency⭐⭐⭐⭐
7PostgreSQL+pgvectorRelational + vectorDocker/pkgFull-featured⭐⭐⭐⭐⭐
8ElasticsearchFull-text + vectorDockerHybrid search⭐⭐⭐
9MongoDB AtlasDocument + vectorCloudCloud-first⭐⭐⭐
10Neo4jGraph + vectorDockerRelationship graph⭐⭐

1. Qdrant, Production-Grade Vector Database 🏆 Top Choice

Overview

A dedicated vector search engine written in Rust, and MemoryHub's default primary database.

Advantages

  • High performance: Implemented in Rust, millisecond-level query response (1 million vectors < 20ms)
  • Rich filtering: Supports combined queries of payload filtering and vector search
  • REST API: Native HTTP API, no driver required
  • Collection isolation: Multiple independent Collections, naturally suited for multi-platform memory separation
  • One-click Docker deployment: docker run -p 6333:6333 qdrant/qdrant

Disadvantages

  • Docker required: No native macOS package (requires Docker Desktop)
  • Moderate resource consumption: ~200MB RAM when idle, ~1-2GB with 1 million vectors loaded
  • No full-text search: Pure vector database; full-text search must be integrated separately (MemoryHub supplements it with SQLite)

MemoryHub Usage

docker run -d -p 6333:6333 -v qdrant_storage:/qdrant/storage qdrant/qdrant

Collection naming: mh_openclaw / mh_hermes / mh_deepseek / mh_claude

Applicable Scenarios

✅ Production environments, large-scale vector search, high-performance filtering requirements
❌ No Docker environment, pure embedded requirements


2. Chroma, the Easiest Vector Database to Get Started With

Overview

A Python-native vector database, ready to use with pip install, zero configuration.

Advantages

  • Simplest installation: pip install chromadb, no Docker dependency
  • Friendly API: Pythonic design, deep integration with LangChain/LlamaIndex
  • Embedded mode: Can be embedded directly into applications, no separate process required
  • Metadata filtering: Supports where condition filtering

Disadvantages

  • Moderate performance: Not as good as Qdrant for large-scale queries
  • Persistence reliability: Embedded mode may have issues under concurrent writes
  • Relatively new ecosystem: Smaller community compared with Qdrant
  • Python 3.14 compatibility: May sometimes require waiting for updates on the latest Python versions

Usage in MemoryHub

In the installer, it is available as a "No Docker" alternative, and selecting it automatically installs and configures it.

Use Cases

✅ Rapid prototyping, development and testing, and not wanting to install Docker
❌ High-concurrency production environments, massive vector queries

3. LanceDB, the Lightest Embedded Vector Database

Overview

An embedded vector database based on the Lance columnar storage format, with no server and no Docker required.

Advantages

  • Extremely lightweight: ready to use with pip install, no separate process
  • Columnar storage: based on Apache Arrow, with high query efficiency
  • Zero configuration: data is stored as local files and can be distributed with the project
  • Supports incremental writes: supports append-only data appends

Disadvantages

  • Fewer features: filtering and hybrid query capabilities are inferior to Qdrant
  • Young ecosystem: documentation and community are relatively small
  • No network API: purely embedded, does not support remote access

How MemoryHub Uses It

Available as an optional backend in the installer, suitable for edge devices or environments without Docker.

Use Cases

✅ Mobile, edge devices, and embedded scenarios
❌ Remote API access and complex filtering queries


4. SQLite-vec, When SQLite Meets Vectors

Overview

A vector extension for SQLite that enables the world's most popular embedded database to support vector search.

Advantages

  • SQLite ecosystem: leverages SQLite's mature ecosystem and toolchain
  • Zero configuration: like SQLite, no server installation required
  • Hybrid queries: a natural combination of SQL queries and vector search
  • Lightweight: single-file database, backup = copying the file

Disadvantages

  • Limited performance: large-scale vector search performance is inferior to dedicated vector libraries
  • Newer extension: a relatively new project, and its API is still evolving
  • Fewer indexing options: not as rich in indexing algorithms as Qdrant/FAISS

How MemoryHub Uses It

Serves as a vector enhancement option for the built-in SQLite full-text index.

Use Cases

✅ Small projects, single-user applications, SQL hybrid queries
❌ Million-scale or larger vector search and high-concurrency scenarios

V. FAISS-Meta: The Cornerstone of Vector Search

Overview

A vector similarity search library developed by Meta AI Research, one of the industry standards.

Advantages

  • GPU acceleration: Native CUDA support, extremely fast GPU queries
  • Rich index support: Supports 10+ index types (IVF, HNSW, PQ, etc.)
  • Academic standard: Widely cited, with abundant papers and documentation
  • Extreme performance: Highly optimized C++ core

Disadvantages

  • Not a database: Essentially a library rather than a database, so persistence must be managed yourself
  • No built-in API: Must be wrapped into a service yourself
  • Complex deployment: Installing faiss-cpu on macOS requires handling compilation dependencies
  • Memory management: All indexes are loaded into memory, causing high memory pressure in large-scale scenarios

MemoryHub Usage

It is not used as the primary database, but it can serve as an optional index acceleration layer during BGE-m3 embedding generation.

Use Cases

✅ Research experiments, GPU-accelerated indexing, scenarios requiring custom index strategies
❌ Scenarios requiring an out-of-the-box persistent database


6. Redis Stack, In-Memory Vector Search

Overview

An enhanced version of Redis with the built-in RediSearch module, supporting vector similarity search.

Advantages

  • Extremely low latency: pure in-memory operations, sub-millisecond response
  • Multiple data structures: supports Key-Value, Hash, List, Set, and others at the same time
  • Mature ecosystem: Redis is used by millions of applications worldwide
  • High concurrency: single-threaded event loop easily handles tens of thousands of QPS

Disadvantages

  • High memory consumption: all data is in memory, so costs are high in large-scale scenarios
  • Complex persistence: AOF/RDB persistence mechanisms require careful configuration
  • Relatively new vector features: RediSearch's vector features are still undergoing rapid iteration
  • Docker dependency: Redis Stack recommends Docker deployment

MemoryHub Usage

Suitable for high-frequency memory search scenarios that require sub-millisecond query response.

Use Cases

✅ High-concurrency real-time search, need for multiple data structures, cache plus vector in one
❌ Large-scale low-cost storage (high memory cost)


7. PostgreSQL + pgvector, the All-Rounder

Overview

The world's most popular open-source relational database, supporting vector search through the pgvector extension.

Advantages

  • One database, many uses: relational data + full-text search + vector search combined in one
  • Mature ecosystem: 30+ years of open-source history, with an extremely rich toolchain
  • ACID transactions: vector operations support transactional consistency
  • Hybrid queries: SQL WHERE + ORDER BY vector_distance combine naturally
  • Operations-friendly: mature backup, replication, and monitoring solutions

Disadvantages

  • Vector performance: large-scale ANN search is not as good as dedicated vector databases
  • Heavyweight deployment: PostgreSQL itself is heavier than Qdrant
  • Index build time: IVFFlat/HNSW index builds are slower on large datasets

MemoryHub Usage

Suitable for scenarios that require unified data storage, using one PostgreSQL instead of the Qdrant + SQLite combination.

Suitable Scenarios

✅ Projects already using PostgreSQL, scenarios requiring transaction support, hybrid queries
❌ Pure vector search, ultra-low latency requirements


8. Elasticsearch: The King of Full-Text Search, Plus Vectors

Overview

A distributed search engine based on Lucene; v8.0+ natively supports vector search.

Advantages

  • Hybrid full-text + vector: Hybrid ranking of BM25 keyword search + KNN vector search
  • Distributed architecture: Native support for clusters, shards, and replicas
  • Rich aggregations: Powerful data analysis and visualization capabilities
  • ELK ecosystem: Seamless integration with Logstash and Kibana

Disadvantages

  • High resource consumption: High JVM memory requirements; a single node needs at least 2-4GB RAM
  • Complex deployment: Configuration and tuning require expertise
  • Poor macOS support: Elasticsearch requires Rosetta or specific configuration on Apple Silicon
  • Operational cost: Cluster management and index optimization require ongoing investment

MemoryHub Usage

Suitable for large-scale deployments that require hybrid full-text + vector search and data analysis.

Use Cases

✅ Enterprise-grade full-text search + vector search, log analysis, large-scale clusters
❌ Lightweight deployments, resource-constrained environments


9. MongoDB Atlas, a Cloud-First Document Vector Database

Overview

MongoDB's cloud-hosted service. Atlas Vector Search enables document databases to support vector search.

Advantages

  • Cloud-hosted: Zero operations, automatic scaling
  • Document model: A natural combination of JSON documents and vector embeddings
  • Atlas Search: Integration with full-text search
  • Global distribution: Multi-region deployment, low-latency access

Disadvantages

  • Cloud-only: Vector Search is available only on Atlas; self-hosted MongoDB does not support it
  • Cost: Vector search requires a higher-tier Atlas cluster
  • Latency: Cloud API calls have higher latency than local Qdrant
  • Network dependency: Requires a stable internet connection

MemoryHub Usage

Suitable for cloud-first deployment architectures, especially team collaboration scenarios.

Use Cases

✅ Cloud-first architecture, team collaboration, existing MongoDB Atlas
❌ Local-first, offline scenarios, cost-sensitive


10. Neo4j, Vector Extension for Graph Databases

Overview

A leading graph database that supports hybrid graph and vector queries through a vector index plugin.

Advantages

  • Graph relationship modeling: Relationships between memories are naturally suited to graph structures.
  • Cypher queries: A powerful graph query language.
  • Hybrid reasoning: Combining graph traversal and vector similarity.
  • Visualization: Built-in graph visualization tools.

Disadvantages

  • Limited vector capabilities: Vector search is not a core feature, and its performance and functionality are inferior to dedicated vector databases.
  • Heavyweight deployment: Neo4j requires substantial memory and storage.
  • Learning curve: The graph database mental model differs from that of traditional databases.
  • Poor fit for AI memory scenarios: Most memory queries do not require graph traversal.

MemoryHub Usage

Retained as an optional backend, suitable for advanced scenarios that require building memory relationship graphs.

Applicable Scenarios

✅ Memory relationship graphs, knowledge graph construction, complex relationship reasoning
❌ General vector search, rapid prototyping


Comprehensive Recommendation Matrix

ScenarioPrimary ChoiceAlternativeAvoid
Individual developer, Mac StudioQdrant (Docker)Chroma (without Docker)Elasticsearch
Rapid prototypingChromaLanceDBNeo4j
Production environment, high concurrencyQdrantRedis StackFAISS (requires a self-built service)
Already using PostgreSQLpgvectorQdrantNone
Cloud-firstMongoDB AtlasQdrant CloudSelf-hosted
Full-text + vector hybridElasticsearchPostgreSQL+pgvectorChroma
Edge devicesLanceDBSQLite-vecElasticsearch
GPU accelerationFAISSNoneNone
Knowledge graphNeo4jNoneNone

MemoryHub's Actual Choice

In the actual deployment of MemoryHub v2.0, we chose the combination of Qdrant + SQLite:

  • Qdrant (primary database): handles all vector search, providing millisecond-level semantic queries
  • SQLite (auxiliary index): handles full-text search and metadata management

The performance of this combination on Mac Studio (M3 Ultra, 64GB):

  • Qdrant idle memory: ~200MB
  • Query latency: <15ms (at a scale of 100,000 vectors)
  • Embedding generation (BGE-m3): ~50ms/item (local inference)

Conclusion

There is no "best" database, only a database that is "best suited to the scenario." MemoryHub's design philosophy is to let users choose according to their own environment: if Docker is available, use Qdrant; if you do not want to install Docker, use Chroma or LanceDB; if PostgreSQL is already available, add the pgvector extension.

Ten databases, ten paths, leading to the same goal: making sure AI Agents no longer forget.


All evaluations are based on the Mac Studio M3 Ultra / macOS 15 / Python 3.12-3.14 environment. Database versions are as of May 2026.