Chunking
Also: 分段 · 切塊 · 文字分塊 · chunk size · chunking 策略
Before embedding, cut long documents into small pieces; a vector compresses the meaning of a whole passage into one point, and the longer the passage, the blurrier that point.
When you will meet it
This is RAG's quietest yet most critical knob. Chunks too big mix several topics into one passage, so the vector averages out and retrieval misses; chunks too small lose context, so a retrieved sentence makes no sense. When retrieval quality is mysteriously bad, chunking is usually the culprit.
An analogy
Like splitting a book into index cards for a library catalogue. Cards too big (a whole chapter) cover too much to match precisely; cards too small (a single word) lose all meaning. The right card is a passage that reads as a self-contained unit.
Minimal example
# 常見做法:固定長度切塊,並讓相鄰塊重疊一點,避免句子被切成兩半
def chunk(text, size=500, overlap=50):
out, i = [], 0
while i < len(text):
out.append(text[i:i + size])
i += size - overlap # 重疊:下一塊往回抓 50 個字
return out
# size / overlap 沒有萬能值,要照文件類型試:
# 技術文件、法規 → 偏小、貼著段落切
# 敘事、對話 → 偏大,保留上下文The point is overlap: it ensures a sentence cut at a boundary is at least whole inside some chunk. size and overlap directly change retrieval results yet have no standard answer — they are parameters you must tune on your own documents, not numbers to copy from someone else.
What people get wrong
- Cutting purely by fixed character count, ignoring sentence and paragraph boundaries. Half a sentence lands in one chunk and half in the next, distorting both vectors, so retrieval suffers. At minimum, cut at sentence breaks and newlines.
- Treating chunking as throwaway preprocessing and pouring effort into a pricier model instead. In practice, improving chunking usually lifts retrieval quality more — and cheaper — than switching models.