RAG and Memory
How does an agent look up documents and store and retrieve what matters?
Build a minimal RAG, long-term memory, and contextual retrieval flow.
📌 Learning goals
- State the difference between RAG and memory in one sentence.
- Explain how data becomes chunks and embeddings, then gets retrieved.
- Build a minimal RAG pipeline whose answers include sources.
- Know what data is worth remembering and what should not be stored.
- Compare two approaches with a small test instead of relying on "it feels better."
Entry conditions
Finish the Stage 3 tool loop and one Stage 4 route; be able to run Python and install packages.
🧭 Lessons on this site
Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.- 01RAG In-Depth Principles and Practice: A Complete Guide from Chunking to Rerank
The complete tech stack for Retrieval-Augmented Generation (RAG): Embedding selection, Chunking strategies, vector database comparison, Rerank optimization, with practical code examples.
11 min - 02Three-Engine Search Strategy: DuckDuckGo + Tavily + Brave Complementary Configuration Guide
Role division, scenario selection, and cost comparison of three search engines in an Agent framework. DuckDuckGo is free and fast, Tavily provides AI summaries, Brave covers news and communities comprehensively.
9 min - 03Web Fetch and Web Scraping in Agent Applications
Configuration and scenario guide for using Web Fetch / Puppeteer / fetch_url to retrieve web content in a three-tier Agent framework.
6 min - 04Cross-Channel Memory Hub: A Full Record of the Memory System Architecture Design for OpenClaw Agent
From daily forgetting to automatic cross-channel integration: a complete record of the architectural evolution of our AI assistant memory system, Session isolation issues, time zone boundary bugs, the three-layer protection mechanism, and the Cron automation integration pipeline.
19 min - 05MemoryHub Practical Installation and Complete Usage Guide: From Zero to Four Platform Automatic Memory Capture
A step by step guide to installing MemoryHub v2.0: from Docker environment preparation, one click pip installation, four platform MCP configuration, Dashboard usage tips, to the complete process of daily maintenance and troubleshooting.
20 min - 06A Panorama of Agent Memory Systems: An In-Depth Comparison of Five Approaches in 2026
Memory is the first-principles problem for Agents: an Agent without memory is just an advanced chatbot. This article compares five approaches (agentmemory / PlugMem / Infini-Memory / verifiable-memory / Qdrant+MemoryHub) and provides a scenario selection matrix.
6 min - 07The Engineering of AI Memory Retrieval: A Full-Matrix Field Report on Four Access Paths × Ten Scenarios
When an AI assistant needs to "remember everything", how does it retrieve memory? Direct file reads, Qdrant vector search, the MemoryHub capture pipeline, and agentmemory semantic indexing: a field comparison of four paths across ten real scenarios, revealing the boundaries and sweet spots of each path.
27 min - 08Vector Memory Deep Dive: A Production-Grade Memory System That Ensures AI Agents Never Lose Memory Again
State amnesia is the #1 killer in production Agent environments. We built a four-layer vector memory system based on Qdrant + BGE-m3, with 9 retrieval modes, validated with 6,675 memories in real-world use, raising Chinese search accuracy from 0% to >78%. One line curl | bash, 100% local deployment, and data sovereignty stays in your hands.
25 min
📚 Required reading
- 1.LangChain Retrieval(components)⭐⭐⭐⭐⭐See how loaders, splitters, embeddings, vector stores, and retrievers cooperate.
- 2.LlamaIndex Concepts⭐⭐⭐⭐⭐Understand indexing and querying the document-centric way.
- 3.Chroma Getting Started⭐⭐⭐⭐See the minimal way to use a local vector database.
🎯 Curated resources
| Resource | Who it's for | Priority | Why |
|---|---|---|---|
RAG frameworks run-llama/llama_index | Beginners building document apps | ⭐⭐⭐⭐⭐ | Index, retriever, query engine; many packages, so start with the official starter. |
RAG frameworks deepset-ai/haystack | Comparing modular pipelines | ⭐⭐⭐⭐ | Components, pipelines, routing; Apache-2.0 — pick one framework to practice. |
RAG frameworks infiniflow/ragflow | Teams wanting a complete web product | ⭐⭐⭐⭐ | Document parsing, retrieval, UI; heavier to deploy than a teaching example. |
Vector databases chroma-core/chroma | First local vector search | ⭐⭐⭐⭐⭐ | Collection, add, query; practice and production setups differ. |
Vector databases qdrant/qdrant | Teams self-hosting or using a managed service | ⭐⭐⭐⭐⭐ | Dense, sparse, and hybrid queries; plan service and backups. |
Vector databases weaviate/weaviate | Need schema and hybrid search | ⭐⭐⭐⭐ | BM25 + vector search; BSD-3-Clause — start with a small baseline. |
Vector databases pgvector/pgvector | Teams already on PostgreSQL | ⭐⭐⭐⭐ | SQL and vectors in one database; still needs index and query tuning. |
Eval & full products vibrantlabsai/ragas | Teams building rerunnable evals | ⭐⭐⭐⭐⭐ | Datasets, metrics, experiments; metrics still need human calibration. |
Eval & full products onyx-dot-app/onyx | Reading a full AI assistant architecture | ⭐⭐⭐⭐ | Ingest, retrieval, chat, admin; the full product is large — an architecture reference, not a starter. |
Further reading LangGraph — Agentic RAG | After basic RAG | ⭐⭐⭐⭐ | See how an agent decides whether to search, rewrite the question, or search again. |
🛠 Hands-on practice (upstream)
Full exercises & starter codeSummaries from the upstream curriculum; full code, cost, and latency estimates live upstream.
- Exercise 1: turn two sentences into embeddings — watch similar sentences land closer in vector space.
- Exercise 2: put embeddings into a vector database — load them into Chroma and retrieve relevant fragments with one question.
- Exercise 3: compare three chunking methods — see what too-large, too-small, and too-much-overlap each do; read the document structure and test results instead of memorizing a standard size.
- Exercise 4: wire up full RAG — retrieve first, then answer, and show which source fragments were used.
- Exercise 5: remember one preference — store only necessary, permitted data, with a way to view, edit, and delete it.
- Recommended mini-project: an assistant that consults documents and remembers a preference — it says "I don't know" without evidence, reads the preference back after a restart, and you can delete the memory.
✅ Self-check
- I can say what retrieval, RAG, and memory each do.
- I can explain how chunks, embeddings, and a vector database connect.
- My RAG answers show sources and admit ignorance when no evidence is found.
- I compare before and after with a small question set, not one beautiful answer.
- Memory stores only necessary, permitted data that the user can view, edit, and delete.
Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.