GPU (graphics processing unit)
Also: 顯示卡 · 顯卡 · graphics card · GPU 是什麼 · CUDA · 獨立顯卡
A chip built to do many simple calculations at once, which is exactly the shape of work a large model needs.
When you will meet it
You meet it at the very first step of picking a machine and reading a spec sheet. But "has a GPU" carries almost no information on its own — whether a model runs is usually decided by how much memory (VRAM) the card has, not how many cores. If all you know is whether a GPU is present, you will hit a wall the moment you try to load a model and be unable to explain why.
An analogy
A CPU is a handful of PhDs: each handles complex, branching work, but there are few of them. A GPU is a factory floor of schoolchildren: each can only add and multiply, but there are so many that they finish a wall of numbers in the same second. A model's forward pass is exactly "one simple operation, repeated millions of times" — the schoolchildren's specialty.
Minimal example
兩台機器都「有 GPU」:
A:入門顯卡,VRAM 很小 → 大模型載不進去,直接失敗
B:資料中心顯卡,VRAM 很大 → 同一個模型輕鬆載入
「有 GPU」回答不了你能跑什麼。
要接著問:VRAM 多大?記憶體頻寬多高?哪一代、哪個架構?Spec sheets love to print core counts and FLOPS in big type, but whether a model runs at all is usually decided by VRAM, and how fast it runs by memory bandwidth. Look at those two first and core count last.
What people get wrong
- Assuming "it has a GPU, so it can run a big model". Whether it can is decided first by whether VRAM holds the model; if it does not fit that is a hard failure, not a slowdown.
- Comparing cards by core count or FLOPS alone and calling it. A "core" is not the same thing across architectures and generations, so the numbers are not directly comparable.
- Treating integrated graphics as a discrete GPU. Integrated graphics usually share system memory and have far less compute; the model sizes they can run differ enormously.