Agentic Research

Prefill and decode

Also: 預填充 · prefill · decode · 解碼階段 · TTFT

The two phases of a generation: prefill digests the entire prompt in one parallel pass (compute-bound, sets time-to-first-token); decode then emits one token at a time (memory-bandwidth-bound, sets tokens per second).

When you will meet it

You need this the first time you notice "a pause after sending, then smooth streaming", or see TTFT and tokens/s quoted side by side on a vendor's spec page. Without the two-phase split you misattribute latency: blaming a slow first token on a weak model (it is your long prompt) or slow streaming on bad network (it is bandwidth and model size) — and then optimize in the wrong direction.

An analogy

Like an interpreter: they absorb the whole statement first (prefill — one gulp, fully parallel), then render it sentence by sentence (decode — strictly one at a time, each requiring a rapid mental re-read of all notes so far). The longer the statement, the longer before they speak; the longer the translation, the longer the job — two entirely different bottlenecks.

Minimal example

送出一段長 prompt 之後,伺服器上發生的事:

[階段一 prefill]
  整個 prompt 的所有 token 一次平行算完
  順手把每一層的中間結果存成 KV 快取
  → 瓶頸:算力(大量矩陣運算)
  → 決定:你等多久才看到第一個字(TTFT)

[階段二 decode]
  每生成一個 token:把整個模型的權重+已生成的
  KV 快取從記憶體搬進計算單元,算出下一個 token
  → 瓶頸:記憶體頻寬(算得少、搬得多)
  → 決定:吐字速度(tokens/s)

This split explains nearly every latency puzzle: a longer prompt makes prefill slower and the first token later (streaming speed unchanged); a longer output adds decode rounds and total time (first token unaffected). It is also why vendors price input and output tokens separately — the two phases have genuinely different cost structures.

What people get wrong

  • Assuming slow decode means "not enough compute" and buying more FLOPs to speed up streaming. Each decode step does little arithmetic; nearly all the time goes to hauling weights and the KV cache — the bottleneck is memory bandwidth. This is why quantization (smaller weights) often directly speeds up tokens per second.
  • Assuming either that earlier turns are always fully recomputed, or that they never are. In reality KV state can be reused within a session (prompt caching); when the cache hits, prefill gets dramatically cheaper — the origin of vendors' "cached input is cheaper" pricing.

Related terms

Next