Agentic Research

Batching (serving many requests together)

Also: 批次處理 · 連續批次 · continuous batching · dynamic batching · 批量推理

Running several requests together to amortize the cost of hauling weights into the compute units. It raises the machine's total throughput but can lengthen any single request's latency.

When you will meet it

This is the core trick of every production inference engine (like vLLM) and the answer to "why can one machine serving many people reach far higher total throughput than handling them one by one". You meet it when building a multi-user service. Without batching you cannot see why "raise throughput" and "cut single-request latency" so often conflict — the central trade-off of capacity planning.

An analogy

Like a washing machine. Washing one item versus gathering ten and washing them together takes about the same time and water — because the fixed cost of "starting the machine" is amortized across ten. Batching is gathering requests up and washing them together. The cost: the one item you need to wear now has to wait for the rest of the load (single-request latency rises).

Minimal example

為什麼批次有效:
  decode 每一步都要把整個模型權重搬進計算單元
  搬一次權重只服務 1 個請求   → 浪費
  搬一次權重同時服務 N 個請求 → 搬運成本被 N 個請求分攤
  → 總吞吐量大幅提升

連續批次(continuous batching):
  不等整批都做完才換下一批
  哪個請求先生成完就立刻釋放、補進新請求
  → 讓 GPU 盡量不空轉

Note the trade-off: the bigger the batch, the higher the total throughput, but each request shares resources with more others and may wait longer (latency rises). Engine parameters like max-num-seqs are exactly how you tune this operating point.

What people get wrong

  • Assuming batching makes "each request faster". It raises total throughput; a single request's latency can actually grow from sharing and queuing.
  • Assuming bigger batches are always better. Past a point the KV cache drains VRAM, or single-request latency exceeds what is acceptable, and the gain backfires.
  • Treating static batching (wait for the whole batch to finish) as the modern approach. Continuous/dynamic batching refills with new requests on the fly so the GPU never idles; it is standard in production engines.

Related terms

Next