Agentic Research

Tokens per second (generation throughput)

Also: 每秒 token · 生成速度 · tokens/sec · tok/s · 輸出速度

How many tokens a model emits per second. During decode it is set mainly by memory bandwidth, not compute; the vendor's figure and your measured figure often differ by a wide margin.

When you will meet it

This is the number spec sheets love to show and the easiest to mislead with. You meet it when reading benchmarks and comparing hardware. The trap: a vendor's tokens-per-second is usually measured under ideal conditions (short prompt, low concurrency, a specific quantization) far from how you actually use it. More fundamentally, decode is memory-bandwidth-bound, so simply adding compute will not raise this number. Miss that and you will buy the wrong hardware for a pretty figure.

An analogy

Like the flow rate of a tap. What sets it is not how fancy the faucet is (compute) but how wide the pipe is (memory bandwidth). Swap in a pricier faucet without widening the pipe and the flow stays the same. Decode works this way: the bottleneck is the pipe that carries data, not the faucet that computes.

Minimal example

為什麼「廠商數字」≠「你的數字」:
  廠商可能在測:短提示、單請求、特定量化、暖機後最佳狀態
  你在用:      長提示、多並發、不同量化、已跑很久(降頻)

而且 tokens/second 有兩個層次,別混:
  單串流速度  一個使用者感受到的吐字速度
  總吞吐量    整台機器每秒服務所有使用者的 token 總和
  → 提高並發會讓總吞吐量上升,但單串流速度可能下降

For any tokens-per-second figure, first ask three things: at what prompt length, what concurrency, what quantization? A number stripped of its conditions means nothing. Also separate "how fast for one user" from "how fast in total for the whole machine" — two goals that often trade off against each other.

What people get wrong

  • Assuming more compute (higher FLOPS) raises generation speed. Decode is memory-bandwidth-bound; when compute is already in surplus, adding more barely helps.
  • Taking the vendor's tokens-per-second as the speed you will get. It is usually a best case under ideal conditions; your long-prompt, high-concurrency reality will be much lower.
  • Treating "single-stream speed" and "total throughput" as one number. A machine can have high total throughput (serving many people) while each user's single-stream speed is only average.

Related terms

Next