KV cache
Also: KV cache · 鍵值快取 · KV 快取 · key-value cache · 注意力快取
Storing each processed token's Key and Value so the next token need not recompute them. The cost: this memory grows with sequence length, which is exactly why long context is so RAM-hungry.
When you will meet it
This is the hidden hinge connecting context length, VRAM and generation speed. You meet it when tuning an inference engine's memory settings, or puzzling over "why does VRAM balloon and concurrency plummet once I enable long context". Without the KV cache you cannot see why long conversations eat more and more memory, or why high concurrency and long context cannot both be had — the most concrete limits in production deployment.
An analogy
Like copying each page's key points onto a growing stack of sticky notes as you read. To continue, you flip the notes (the cache) instead of re-reading the whole book, so it is fast. But the thicker the book, the more notes you stick, until the whole desk (VRAM) is covered — that is the real reason long context eats memory.
Minimal example
沒有 KV cache:
生成下一個 token → 把前面所有 token 全部重算一遍 → 極慢
有 KV cache:
前面所有 token 的 Key/Value 已存好 → 只算新的這個 → 快
但記憶體代價:
KV cache 大小 ∝ 序列長度 × 並發請求數 ×(模型層數與維度)
· 上下文越長 → 每個請求的 KV cache 越大
· 並發越高 → 同時存在的 KV cache 越多
→ 這就是為什麼「長上下文」和「高並發」會互搶 VRAMInference engines (like vLLM's PagedAttention) exist to manage this cache cleverly: split it into blocks, allocate on demand, share prefixes, and fit more concurrent requests into the same VRAM. Understanding how the KV cache grows is the prerequisite for tuning those parameters.
What people get wrong
- Assuming the KV cache is optional acceleration. It is the very precondition for fast decode; without it every token recomputes the whole sequence, too slow to use.
- Assuming its memory is fixed. It grows with sequence length; in long conversations or long context the KV cache can eat more VRAM than the model weights.
- Assuming maxing out context is free. The longer the context, the more VRAM is reserved for the KV cache and the fewer concurrent requests you can serve — a direct trade-off.