Quantization
Also: 量化 · 模型量化 · 權重量化 · Q4 · Q8 · 4-bit
Storing model weights in lower-precision numbers, trading tolerable quality loss for dramatically smaller size and memory footprint.
When you will meet it
You must understand this the first time you run a model locally and see the same model offered as Q8, Q6 and Q4 downloads — otherwise you are guessing from filenames. Without it you make two symmetric mistakes: grab the biggest file and blow past your memory, or grab the most aggressive one and wonder why the model got stupid.
An analogy
Like saving a photo as JPG: compress a little and the eye sees no difference at half the file size; compress more and it still holds; past a threshold the image suddenly turns to mush. Quality loss is not linear — the first steps feel free, the last one is a cliff.
Minimal example
同一個模型的不同精度版本(示意,實際比例依模型而異):
FP16/BF16 基準大小 幾乎無損,伺服器端常用
8-bit(Q8) 約一半 多數任務與基準難以分辨
4-bit(Q4) 約四分之一 開始量得到品質下降,本機最常見
更激進 更小 風險陡增,除非別無選擇,別碰
注意兩件事:
· 縮小的是「每個權重佔的位元數」,不是刪掉權重
· 好的量化方案會把特別敏感的層留在較高精度Halve the bits and file size and memory roughly halve — that part is arithmetic. Quality loss is not proportional: moderate quantization is often imperceptible, while aggressive quantization can collapse specific capabilities suddenly and unevenly (multi-step reasoning usually breaks before small talk does).
What people get wrong
- Assuming quantization only makes files smaller and downloads faster. What it actually unlocks is fit: the weights must sit entirely in memory (VRAM or unified memory) to run at all, and quantization is often the exact line between "runs locally" and "does not". A side effect — less data to haul — can even speed up decode.
- Assuming quality loss scales linearly with compression, so "Q8 is imperceptible" implies "more aggressive grades are only slightly worse". The curve steepens as you push, and different capabilities break in different orders — other people's anecdotes are hints, not evidence. Validate aggressive settings on your own eval set.