MoE (Mixture of Experts)
Also: 混合專家 · Mixture of Experts · 專家混合 · 稀疏模型 · sparse model · MoE 是什麼
An architecture that keeps many "expert" sub-networks inside but routes each token to only a few of them — huge total parameters, much smaller compute per token.
When you will meet it
When a spec sheet says "large total parameters, only a fraction activated per token", or someone claims a huge model "runs fine locally", MoE is usually why. Without it you use the wrong two rulers: estimating compute cost from total parameters (overstating the speed bottleneck), or estimating memory from activated parameters (understating it so badly the model never loads).
An analogy
Like a big hospital: the triage desk (router) reads your symptoms and sends you to the two or three relevant specialists, never a full-staff consultation. How many doctors the hospital employs (total parameters) decides how big the building must be (memory); what each patient actually consumes is only the on-duty doctors' time (compute).
Minimal example
示意數字,只為講清比例關係:
密集模型: 100 份參數,每個 token 都用到全部 100 份
MoE 模型: 100 份參數分成很多組專家,
每個 token 只經過其中 10 份
後果:
· 每 token 的計算量 ≈ 密集模型的十分之一 → 生成快、成本低
· 但所有專家都得常駐記憶體 → 記憶體需求仍按 100 份算
· 「知識容量」大(參數多)而「每步算力」小(啟用少),
正是 MoE 想兩邊都拿的設計The one takeaway is the asymmetry: compute per token follows activated parameters; memory follows total parameters. When a spec sheet gives only one of the two, derive the other — get either direction wrong and your local-deployment feasibility estimate becomes fiction.
What people get wrong
- Seeing a huge total parameter count and concluding "unrunnable, necessarily expensive". Per-token compute tracks only the activated slice, so a large MoE is often far cheaper and faster than a dense model of similar total size — the intuition "big parameters = slow = expensive" misfires on MoE.
- Conversely, seeing "few activated parameters" and assuming a small model any machine can run. Every expert must be loaded and stay resident — memory demand follows the total count and is not small at all. MoE is compute-light, memory-heavy.