Parameter count
Also: 模型大小 · 7B · 70B · 參數量 · 參數數量 · billions of parameters
How many adjustable numbers a model holds. It is often used as a label for "how big and smart" a model is, but that label misleads more often than you think.
When you will meet it
The first number you see when picking a model is usually this one (7B, 70B and so on). The intuition is "bigger is stronger", so beginners chase the largest parameter count, then hit the VRAM wall or find the big model is slower and no better on their simple task. Understanding that "more parameters" does not mean "better for your task" is what keeps you from choosing the wrong tool.
An analogy
Parameter count is like a book's word count. A longer book usually holds more, but when you need only one phone number a thick encyclopedia is harder to use than a sticky note — what matters is not the word count but whether it has what you need and how easily you can reach it.
Minimal example
參數量 → 記憶體(不是效能保證):
參數量 × 每個參數的位元組 ≈ 權重大小
· 量化位元越高 → 權重越大 → 越吃 VRAM
· 同一個參數量,量化後大小可以差好幾倍
但「總參數量」對 MoE 模型會失真:
總量很大,實際每次推理只激活其中一部分
→ 用總參數量會高估它的算力/記憶體需求(詳見 moe 詞條)Two points: parameter count mainly predicts how much memory a model takes, not directly how smart or how fast it is; and for MoE models the total is a misleading number — cross-reference the moe entry rather than restating its mechanism here.
What people get wrong
- Assuming more parameters is always better for your task. On simple tasks a bigger model is just slower and hungrier for memory, not necessarily higher quality.
- Treating parameter count as a direct guarantee of speed or VRAM. You must multiply by bytes-per-parameter (quantization) to know the real size.
- Using total parameter count to reason about MoE models. MoE activates only a subset per step, so the total makes you overestimate its compute and memory needs.