Prefill and decode
Also: 預填充 · prefill · decode · 解碼階段 · TTFT
The two phases of a generation: prefill digests the entire prompt in one parallel pass (compute-bound, sets time-to-first-token); decode then emits one token at a time (memory-bandwidth-bound, sets tokens per second).
When you will meet it
You need this the first time you notice "a pause after sending, then smooth streaming", or see TTFT and tokens/s quoted side by side on a vendor's spec page. Without the two-phase split you misattribute latency: blaming a slow first token on a weak model (it is your long prompt) or slow streaming on bad network (it is bandwidth and model size) — and then optimize in the wrong direction.
An analogy
Like an interpreter: they absorb the whole statement first (prefill — one gulp, fully parallel), then render it sentence by sentence (decode — strictly one at a time, each requiring a rapid mental re-read of all notes so far). The longer the statement, the longer before they speak; the longer the translation, the longer the job — two entirely different bottlenecks.
Minimal example
送出一段長 prompt 之後,伺服器上發生的事:
[階段一 prefill]
整個 prompt 的所有 token 一次平行算完
順手把每一層的中間結果存成 KV 快取
→ 瓶頸:算力(大量矩陣運算)
→ 決定:你等多久才看到第一個字(TTFT)
[階段二 decode]
每生成一個 token:把整個模型的權重+已生成的
KV 快取從記憶體搬進計算單元,算出下一個 token
→ 瓶頸:記憶體頻寬(算得少、搬得多)
→ 決定:吐字速度(tokens/s)This split explains nearly every latency puzzle: a longer prompt makes prefill slower and the first token later (streaming speed unchanged); a longer output adds decode rounds and total time (first token unaffected). It is also why vendors price input and output tokens separately — the two phases have genuinely different cost structures.
What people get wrong
- Assuming slow decode means "not enough compute" and buying more FLOPs to speed up streaming. Each decode step does little arithmetic; nearly all the time goes to hauling weights and the KV cache — the bottleneck is memory bandwidth. This is why quantization (smaller weights) often directly speeds up tokens per second.
- Assuming either that earlier turns are always fully recomputed, or that they never are. In reality KV state can be reused within a session (prompt caching); when the cache hits, prefill gets dramatically cheaper — the origin of vendors' "cached input is cheaper" pricing.