CPU inference (running a model without a GPU)
Also: CPU 推理 · 純 CPU 跑模型 · 沒有顯卡能跑嗎 · CPU only · 用處理器跑 AI
Running a model on the CPU alone, with no GPU. For small models, low volume and offline batch work it is genuinely usable; for large models or interactive use it is essentially hopeless.
When you will meet it
This is the answer to "I have no GPU — can I still do local AI". Without knowing its limits you fall into one of two extremes: assuming nothing runs without a GPU (small models do), or assuming a CPU can run big models too (and each token takes an age, so you think the program is broken). Knowing the line stops you from over-buying hardware or waiting for nothing.
An analogy
CPU inference is like moving a mountain of stones with one versatile but short-handed machine: it can do anything, but only a little at a time. A small pile (small model, short output) gets moved eventually; a mountain (big model, long output) takes forever.
Minimal example
CPU 推理「堪用」的場景:
· 小模型 + 短輸出(分類、抽取、改寫幾句話)
· 不趕時間的離線批次(半夜跑一批文件)
· 只想驗證流程、對速度無所謂
CPU 推理「無望」的場景:
· 大模型(權重遠超快取能容納的量,每步都要從慢速記憶體搬)
· 互動式聊天(你不會想每句話都等上半天)
· 長輸出(生成越長,CPU 的劣勢被放大得越厲害)The watershed is that decode keeps hauling weights into the compute units. A CPU's memory bandwidth and parallelism are far below a GPU's, and the bigger the model and longer the output, the more fatal that gap. Small tasks are tolerable; large ones are not.
What people get wrong
- Assuming "no GPU means no local models at all". Small and quantized models are usable on CPU alone — just slow.
- Assuming a CPU is only a little slower and you can live with it. For big models the gap is not "a little" — it is slow enough to exhaust your patience, and the interactive experience drops to zero.
- Running long outputs or high concurrency on a CPU and expecting GPU-like throughput. A CPU is weakest exactly at generation and parallelism, which is what these scenarios need most.