RAG In-Depth Principles and Practice: A Complete Guide from Chunking to Rerank
What is RAG?
RAG (Retrieval-Augmented Generation) lets an LLM retrieve relevant documents from an external knowledge base before answering, then generate an answer based on the retrieved results. This solves two core problems of LLMs:
- Knowledge cutoff: LLMs only know information from their training data
- Hallucination: LLMs may fabricate facts that do not exist
User question → Embedding → Vector search → Top-K documents → LLM generation → Answer with sources
Core Component Breakdown
1. Embedding Model Selection
| Model | Dimensions | Language | Cost | Applicable Scenarios |
|---|---|---|---|---|
| text-embedding-3-small | 1536 | Multilingual | $0.02/1M tokens | General purpose, cost-sensitive |
| text-embedding-3-large | 3072 | Multilingual | $0.13/1M tokens | High-precision requirements |
| bge-large-zh-v1.5 | 1024 | Chinese | Free (local) | Preferred for Chinese scenarios |
| bge-m3 | 1024 | Multilingual | Free (local) | Mixed multilingual scenarios |
| jina-embeddings-v3 | 1024 | Multilingual | Free (API) | Long documents (8K tokens) |
Recommendation: For Chinese-first scenarios, use bge-large-zh-v1.5; for multilingual scenarios, use bge-m3; for API solutions, use text-embedding-3-small.
# Ollama Local Embedding
ollama pull nomic-embed-text
# or bge-m3
ollama pull bge-m3
2. Chunking Strategies
This is the most easily overlooked yet most critical step. Incorrect chunking directly leads to retrieval failure.
| Strategy | Suitable Scenarios | Pros | Cons |
|---|---|---|---|
| Fixed size (500 tokens) | General purpose | Simple, predictable | May truncate mid-sentence |
| Semantic splitting (by paragraph/section) | Structured documents | Preserves semantic integrity | Inconsistent chunk sizes |
| Recursive splitting (by \n\n → \n → 。) | Mixed documents | Adaptive | Complex to implement |
| Sentence window | Precise retrieval | Each sentence is independent | Loses context |
Best practices:
- Chunk size: 256-512 tokens (too small loses context, too large dilutes semantics)
- Overlap: 10-20% (ensures key information is not truncated)
- Use LangChain's
RecursiveCharacterTextSplitter
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=50,
separators=["\n\n", "\n", "。", ",", " ", ""]
)
chunks = splitter.split_text(document)
3. Vector Database Comparison
| Option | Deployment | Scale | Suitable Scenarios |
|---|---|---|---|
| Chroma | Local/Embedded | <100K docs | Prototyping/Personal projects |
| Qdrant | Local/Docker | Million-scale | Production |
| Milvus | Distributed | Hundred-million scale | Enterprise-grade |
| Pinecone | Cloud SaaS | Unlimited | No infrastructure management |
| pgvector | PostgreSQL extension | Million-scale | Projects already using PG |
Recommended path: Use Chroma for prototyping → use Qdrant Docker for production → use Milvus for very large scale.
# Qdrant Quick Start
docker run -p 6333:6333 qdrant/qdrant
4. Rerank (Reranking)
After initial vector retrieval, use a more precise Cross-encoder to rerank the Top-K results, significantly improving accuracy.
| Model | Language | Description |
|---|---|---|
| bge-reranker-v2-m3 | Multilingual | Best open source |
| Cohere Rerank | Multilingual | API, best performance |
| bge-reranker-large | Chinese/English | Recommended for Chinese scenarios |
Why is Rerank needed? Vector similarity ≠ semantic relevance. Rerank uses a Cross-encoder to compare query and document pairwise, making it more precise.
# Typical workflow
# 1. Vector retrieval → 20 candidates
candidates = vector_db.search(query, top_k=20)
# 2. Rerank → take top 5
from FlagEmbedding import FlagReranker
reranker = FlagReranker('BAAI/bge-reranker-v2-m3')
scores = reranker.compute_score([[query, doc] for doc in candidates])
top_5 = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True)[:5]
# 3. LLM generation
answer = llm.generate(query, context=top_5)
Complete RAG Pipeline Code
from langchain_community.embeddings import OllamaEmbeddings
from langchain_community.vectorstores import Chroma
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.llms import Ollama
# 1. Load document
with open("knowledge_base.txt") as f:
text = f.read()
# 2. Chunking
splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.split_text(text)
# 3. Embedding + store in vector store
embeddings = OllamaEmbeddings(model="bge-m3")
vectorstore = Chroma.from_texts(chunks, embeddings)
# 4. Retrieval
retriever = vectorstore.as_retriever(search_kwargs={"k": 5})
docs = retriever.get_relevant_documents("What is RAG?")
# 5. Generation
llm = Ollama(model="qwen2.5:7b")
context = "\n".join([d.page_content for d in docs])
answer = llm.invoke(f"Answer the question based on the following documents:\n{context}\n\nQuestion: What is RAG?")
Common Pitfalls
| Pitfall | Cause | Solution |
|---|---|---|
| Irrelevant retrieval results | Chunk size too large or too small | Tune to 256-512, add overlap |
| Slow retrieval | No index built | Build an HNSW index with Qdrant/Milvus |
| High embedding cost | Calls the API on every query | Deploy bge-m3 locally |
| Answer hallucination | LLM ignores retrieval results | Force source citations in the prompt |
| Mixed Chinese-English queries fail | Embedding model does not support them | Use the bge-m3 multilingual model |
Recommended Reading
More in Learn
- Complete LangChain Tutorial 2026: Building Enterprise-Grade LLM Applications from Scratch
- MemoryHub v2.0 System Architecture In-Depth Analysis: From Capture Daemon to MCP Real-Time Memory Capture
- May 2026 LLM API Pricing Landscape: Complete Comparison of DeepSeek, Qwen, GLM, Kimi, MiniMax, and Doubao
- Cross-Channel Memory Hub: A Full Record of the Memory System Architecture Design for OpenClaw Agent