Agentic Research

TTFT (time to first token)

Also: 首字延遲 · 第一個 token 的時間 · time to first token · 首 token 延遲 · 等待時間

The dead air between pressing send and the first character appearing. It is driven mainly by prompt length and prefill, and it is the number users actually experience as "lag".

When you will meet it

This is the real source of "feels slow", yet spec sheets almost never mention it because it does not market well. You meet it when building interactive tools: a system with a long TTFT feels unresponsive even if generation afterwards is fast. If you watch only tokens-per-second and ignore TTFT, you will ship something with pretty numbers and a poor feel.

An analogy

Like asking a friend a question and watching them think for three seconds before saying the first word. That silence (TTFT) decides whether they feel slow to respond; once they start talking, how fast they speak (tokens-per-second) is a separate matter. The longer the silence, the more you fidget.

Minimal example

同一次請求的兩個數字:
  TTFT(首字延遲)       ← prefill:把整段提示讀進去、算一遍
  生成速度 tokens/second  ← decode:之後一個個吐字

提示越長(塞了整份文件或超長對話)
  → prefill 要算的越多 → TTFT 越長
高並發下還要排隊 → TTFT 再被拉長

使用者「按下去卡住」的感受,幾乎全來自 TTFT,不是 tokens/second。

Prefill is compute-bound, so TTFT is sensitive to compute and prompt length; decode is memory-bound, so tokens-per-second is sensitive to bandwidth. These are two different bottlenecks — measure and optimize them separately.

LLM inference: the prefill and decode phasesFlow diagram: generation happens in two phases. Prefill processes the whole prompt in one parallel pass, storing each token's K and V in the KV cache; this phase is compute-bound and determines TTFT. Decode then produces one token at a time — read the cache, one forward pass, append the new token, repeat — until the model emits EOS or the output limit is hit; this phase is memory-bandwidth-bound and determines tokens per second. The timeline at the bottom shows the two bottlenecks differ, so shortening the prompt helps TTFT while making the model terser helps total time.Phase 1 · PREFILL (whole prompt, one parallel pass)explain this code to me plea… N tokens in totalone forward pass, all tokens at oncecompute-bound: limited by compute, so a longer prompt makes this slowerKV cacheK and V per token, computed once and keptprefill result storedPhase 2 · DECODE (one token at a time, until done)read cache + new tokenmemory-bandwidth-boundemit one tokenand append it to the cachethe token just produced becomes the next input — this is autoregressiononly two reasons to stop:the model emits EOS, or max output is hitwhat the user actually experiencesTTFT ← set by prompt lengthtokens/second ← set by output lengthDifferent bottlenecks, so different fixes: shorten the prompt (or use prefix caching) for TTFT; make the model terser for total time.Collapsing both into one "speed" number means optimising the wrong thing.
1/8Input: the whole prompt
System prompt, history, documents and this turn's question are concatenated into one token sequence.
Step 1 of 8 Input: the whole prompt

What people get wrong

  • Using tokens-per-second to judge responsiveness. The "lag" users feel is TTFT; however fast generation is, if the first character is slow to appear the experience is still poor.
  • Assuming a fuller prompt is smarter while ignoring the TTFT cost. Stuffing in a whole document or huge history raises the work prefill must do, and you pay that latency again on every interaction.
  • Watching only average TTFT, not tail latency (p95/p99). Under concurrency, queuing makes a few requests' TTFT spike, and the average hides that long tail.

Related terms

Next