Agentic Research

GPU (graphics processing unit)

Also: 顯示卡 · 顯卡 · graphics card · GPU 是什麼 · CUDA · 獨立顯卡

A chip built to do many simple calculations at once, which is exactly the shape of work a large model needs.

When you will meet it

You meet it at the very first step of picking a machine and reading a spec sheet. But "has a GPU" carries almost no information on its own — whether a model runs is usually decided by how much memory (VRAM) the card has, not how many cores. If all you know is whether a GPU is present, you will hit a wall the moment you try to load a model and be unable to explain why.

An analogy

A CPU is a handful of PhDs: each handles complex, branching work, but there are few of them. A GPU is a factory floor of schoolchildren: each can only add and multiply, but there are so many that they finish a wall of numbers in the same second. A model's forward pass is exactly "one simple operation, repeated millions of times" — the schoolchildren's specialty.

Minimal example

兩台機器都「有 GPU」:
  A:入門顯卡,VRAM 很小      → 大模型載不進去,直接失敗
  B:資料中心顯卡,VRAM 很大   → 同一個模型輕鬆載入

「有 GPU」回答不了你能跑什麼。
要接著問:VRAM 多大?記憶體頻寬多高?哪一代、哪個架構?

Spec sheets love to print core counts and FLOPS in big type, but whether a model runs at all is usually decided by VRAM, and how fast it runs by memory bandwidth. Look at those two first and core count last.

What people get wrong

  • Assuming "it has a GPU, so it can run a big model". Whether it can is decided first by whether VRAM holds the model; if it does not fit that is a hard failure, not a slowdown.
  • Comparing cards by core count or FLOPS alone and calling it. A "core" is not the same thing across architectures and generations, so the numbers are not directly comparable.
  • Treating integrated graphics as a discrete GPU. Integrated graphics usually share system memory and have far less compute; the model sizes they can run differ enormously.

Related terms

Next