Concurrency (how many requests at once)
Also: 並發 · 並發數 · 同時請求數 · 併發
How many requests a system handles at the same time. How fast one user feels it is, and how many people one machine can serve at once, are two completely different problems.
When you will meet it
When your AI goes from "just me" to "my team or my users", you hit this word. Everything is fast in solo testing, then slows or collapses the moment several people use it at once — because you kept measuring single-stream speed and never measured throughput and latency under concurrency. Without this distinction your capacity planning will be completely off.
An analogy
Like a restaurant. One guest is served quickly (single-stream speed), but whether the kitchen holds up when twenty tables order at once is a different matter (concurrency). When it cannot hold up, it is not one dish that slows — every table waits. And to maximize total covers served, the kitchen may deliberately make some tables wait a moment (batching), which is the throughput-versus-latency trade-off again.
Minimal example
正確的容量測試是做「並發掃描」:
同時 1 個請求 → 記錄:吞吐量、每個請求的延遲
同時 5 個請求 → 記錄:吞吐量、延遲
同時 20 個請求 → 記錄:吞吐量、延遲
…逐步加壓,直到延遲或錯誤率不可接受
你會看到一條取捨曲線:
並發越高 → 總吞吐量越高,但每個請求的延遲也越高
「最佳工作點」取決於你的業務能忍受多少延遲Looking only at "how fast for one request" badly overestimates capacity. Always look at latency percentiles (p95/p99), not just the average — as concurrency rises, the worst-hit are the tail requests that wait longest.
What people get wrong
- Estimating multi-user capacity from single-user speed. Under concurrency, resources are shared and requests queue, so the number you can actually serve is far below the "speed × people" intuition.
- Assuming total throughput keeps rising with concurrency. At some point memory (the KV cache) or compute is exhausted; beyond it, more concurrency only spikes latency or starts throwing errors.
- Watching only average latency. The pain of concurrency is in the tail; the average completely hides "a few requests wait forever".