Reranking
Also: 重新排序 · 重排 · reranker · cross-encoder · 第二輪排序
The second pass of retrieval: a cheap, fast method pulls a batch of candidates, then a smarter, pricier model reads the query together with each candidate to reorder them by true relevance.
When you will meet it
Pure vector search mistakes "reads like" for "answers correctly", so top-k often mixes in passages that are semantically near but off-topic. Reranking filters out those false positives and pushes the genuinely useful passage to the top — directly deciding the quality of context fed to the model.
An analogy
Like two-stage hiring. Stage one screens out obviously-unfit resumes by keyword (cheap, high volume); stage two has a manager read and interview each to rank who should really be hired (expensive, low volume). Vector search is stage one; reranking is stage two.
Minimal example
問題:「我們公司的報銷上限是多少?」
向量檢索 top-5(只看語義相近,順序不一定對):
1. 出差住宿的報銷流程 ⋯(像,但沒講上限)
2. 報銷上限是台幣三萬元 ← 正確答案,卻排在第 2
3. 報銷需要附收據 ⋯
4. 員工福利說明 ⋯
5. 財務系統登入方式 ⋯
rerank 之後(把問題和每段一起細看):
1. 報銷上限是台幣三萬元 ← 頂到最前
2. 出差住宿的報銷流程 ⋯
⋯Look at item 2: vector search did retrieve the right answer, but ranked it second, its similarity score edged out by other passages. If you feed only top-1 to the model, you miss it. Reranking's value is not "finding what was not found" but "reordering what was found and misplaced".
What people get wrong
- Assuming reranking can rescue bad retrieval. If vector search never pulled the right passage into the candidate set, reranking cannot conjure it — it only reorders what was already retrieved. Bad embedding and chunking upstream cannot be saved downstream.
- Running rerank over all hundreds of thousands of passages. Rerankers are slow and expensive; the right pattern is two-stage — vector search narrows to top-50, then rerank those 50 and pass the top few to the model.