Agentic Research

RAG (Retrieval-Augmented Generation)

Also: 檢索增強生成 · 知識庫問答 · retrieval augmented generation

Look material up first, hand the retrieved passages to the model along with the question, and have it answer from them.

When you will meet it

This is the standard way to make AI answer from YOUR material rather than from what it remembers. Nearly all enterprise document Q&A is RAG.

An analogy

An open-book exam versus a closed-book one. Closed book tests memory (and fabricates); open book tests lookup and synthesis (and can cite the wrong page). RAG makes it open book.

Minimal example

沒有 RAG:
  問「我們公司的報銷上限是多少」→ 模型不知道,但會編一個數字

有 RAG:
  同一個問題 → 先去員工手冊的向量庫撈最相關的兩段 →
  把那兩段塞進 prompt → 模型照著回答,並能指出出自哪一段

The key difference is not correctness but traceability: when it is wrong you can tell whether retrieval missed, retrieved the wrong thing, or retrieved correctly and the model ignored it.

What people get wrong

  • Assuming RAG cures hallucination. It reduces it, but the model can still ignore the supplied passages and answer from memory.
  • Doing vector search with no reranking. Semantically near is not the same as correct; top-k results often mix in near-but-wrong passages.

Related terms

Next