Agentic Research

VRAM (GPU memory)

Also: 顯存 · GPU 記憶體 · 視訊記憶體 · video memory · GPU RAM

The GPU's own memory. Model weights and the KV cache must fit here first; if they do not, the model does not run.

When you will meet it

This is the first and hardest wall you hit when running a model locally. It is not a performance metric; it is the on/off switch for whether the model runs at all. Without it you stare helplessly at "load failed even though there is a GPU" — because the problem is not compute, it is that the memory cannot hold the model.

An analogy

VRAM is a table, not a warehouse. Anything you need must be laid out on the tabletop to be usable; the table is only so big, and what does not fit simply does not fit. Piling the overflow onto the floor nearby (system memory) means you cannot use it, and walking to the floor each time is so slow it is unusable.

Minimal example

載入一個模型時,VRAM 至少要吃下:
  模型權重      ← 參數量 × 每個參數的位元組(量化越少越大)
  KV cache      ← 隨上下文長度成長,長對話越來越大
  執行時開銷    ← 框架、暫存張量
  ─────────────
  加總 > VRAM 容量  →  不是變慢,是直接報錯、載入失敗

The key is the last line: exceeding capacity is a hard failure. Many beginners assume it will just slow down; in reality most frameworks throw an out-of-memory error and abort.

What people get wrong

  • Assuming too little VRAM just "slows things down". It is a wall: if the model does not fit you get an out-of-memory failure, not a graceful slowdown.
  • Assuming it is enough that the weights fit. Forgetting that the KV cache grows with context length, and in long conversations or high concurrency it can eat more VRAM than the weights.
  • Treating "lots of system RAM" as "lots of VRAM". They are different pools; on a discrete GPU, the part that spills from VRAM into system memory is so slow it is nearly unusable.

Related terms

Next