TTFT (time to first token)
Also: 首字延遲 · 第一個 token 的時間 · time to first token · 首 token 延遲 · 等待時間
The dead air between pressing send and the first character appearing. It is driven mainly by prompt length and prefill, and it is the number users actually experience as "lag".
When you will meet it
This is the real source of "feels slow", yet spec sheets almost never mention it because it does not market well. You meet it when building interactive tools: a system with a long TTFT feels unresponsive even if generation afterwards is fast. If you watch only tokens-per-second and ignore TTFT, you will ship something with pretty numbers and a poor feel.
An analogy
Like asking a friend a question and watching them think for three seconds before saying the first word. That silence (TTFT) decides whether they feel slow to respond; once they start talking, how fast they speak (tokens-per-second) is a separate matter. The longer the silence, the more you fidget.
Minimal example
同一次請求的兩個數字:
TTFT(首字延遲) ← prefill:把整段提示讀進去、算一遍
生成速度 tokens/second ← decode:之後一個個吐字
提示越長(塞了整份文件或超長對話)
→ prefill 要算的越多 → TTFT 越長
高並發下還要排隊 → TTFT 再被拉長
使用者「按下去卡住」的感受,幾乎全來自 TTFT,不是 tokens/second。Prefill is compute-bound, so TTFT is sensitive to compute and prompt length; decode is memory-bound, so tokens-per-second is sensitive to bandwidth. These are two different bottlenecks — measure and optimize them separately.
What people get wrong
- Using tokens-per-second to judge responsiveness. The "lag" users feel is TTFT; however fast generation is, if the first character is slow to appear the experience is still poor.
- Assuming a fuller prompt is smarter while ignoring the TTFT cost. Stuffing in a whole document or huge history raises the work prefill must do, and you pay that latency again on every interaction.
- Watching only average TTFT, not tail latency (p95/p99). Under concurrency, queuing makes a few requests' TTFT spike, and the average hides that long tail.