Agentic Research

Unified memory (CPU and GPU share one pool)

Also: 統一記憶體 · UMA · unified memory · Apple Silicon 記憶體 · 共享記憶體

CPU and GPU share one block of memory, with no copying between two pools. Apple Silicon is the classic example.

When you will meet it

When you ask "why can a Mac laptop load a model that a big-discrete-GPU desktop cannot", the answer is here. It explains an apparent paradox: a machine with weaker compute can run a bigger model. Without this trade-off you will shop by GPU core count alone and miss the memory capacity and architecture that actually decide how much you can load.

An analogy

A traditional layout is two departments each with their own warehouse, goods shuttled back and forth (copied). Unified memory is one big shared warehouse: whoever needs stock takes it on the spot, no shuttling. The cost: everyone crowds the same warehouse and the same aisle (bandwidth), and they block each other when it is busy.

Minimal example

離散 GPU(傳統):
  權重必須先放進 VRAM
  VRAM 裝不下 → 溢出到系統記憶體 → 每步都要跨池搬 → 慢到不可用

統一記憶體(Apple Silicon 等):
  CPU/GPU 看的是同一個大池子
  池子夠大 → 大模型整個放得下 → 不必跨池複製
  但頻寬是共享的 → 生成速度(decode)受限於這條共享頻寬

The trade is clear: it removes the fatal "copy between pools" bottleneck and replaces it with the milder limit of shared bandwidth. That is why a big-memory Mac can run a huge model, yet its generation speed does not necessarily beat a high-bandwidth discrete GPU.

What people get wrong

  • Assuming unified memory has "no bottleneck". It swaps the copy bottleneck for a shared-bandwidth one — bandwidth is split among CPU, GPU and other programs, which caps decode speed.
  • Assuming unified memory means "a strong GPU". What lets a big model load is memory capacity; how fast it generates is a separate matter, depending on bandwidth and GPU cores.
  • Assuming a big pool means you can set an arbitrarily long context. The model and the KV cache share the same pool; too long a context still drains it and fails midway through generation.

Related terms

Next