FLOPS (floating-point operations per second)
Also: 浮點運算 · 算力 · TFLOPS · 每秒浮點運算
A raw compute number: how many floating-point operations per second. It looks the most technical, yet it predicts the speed you actually feel the worst.
When you will meet it
Spec sheets and marketing pages love to bombard you with FLOPS because the numbers are big and impressive. But the speed you actually feel (how long a token takes) is mostly set not by FLOPS but by memory bandwidth. If you treat FLOPS as the answer to "is it fast", you will buy a high-compute card, find generation no faster, and be unable to explain why.
An analogy
FLOPS is like how fast a chef can chop. But if the ingredients (data) must be carried in from a distant warehouse one trip at a time, the fastest knife still waits. During generation most of the time goes to carrying data, not chopping — so the carrying speed (memory bandwidth) is the bottleneck.
Minimal example
兩個階段,瓶頸不同:
prefill(讀你的提示) → 計算密集 → FLOPS 比較有意義
decode(一個個吐 token) → 記憶體密集 → 頻寬才是關鍵,FLOPS 幫不上忙
你「感受到的慢」多半在 decode,
而那一段恰恰是 FLOPS 最無用的地方。So what you should look at is memory bandwidth and measured tokens-per-second and TTFT, not FLOPS. FLOPS is somewhat relevant to prefill and almost irrelevant to the generation speed you wait on every day.
What people get wrong
- Assuming more FLOPS means faster token generation. Generation (decode) is memory-bandwidth-bound; adding compute barely speeds it up.
- Comparing FLOPS figures across architectures and precisions. The same card reports very different FLOPS at different precisions (say FP16 versus FP8), so numbers on different bases cannot be compared.
- Confusing FLOPS (per second, a rate) with FLOPs (a total count, an amount of work). The capital-versus-lowercase s distinction is one marketing copy often blurs, deliberately or not.