Agentic Research

MoE (Mixture of Experts)

Also: 混合專家 · Mixture of Experts · 專家混合 · 稀疏模型 · sparse model · MoE 是什麼

An architecture that keeps many "expert" sub-networks inside but routes each token to only a few of them — huge total parameters, much smaller compute per token.

When you will meet it

When a spec sheet says "large total parameters, only a fraction activated per token", or someone claims a huge model "runs fine locally", MoE is usually why. Without it you use the wrong two rulers: estimating compute cost from total parameters (overstating the speed bottleneck), or estimating memory from activated parameters (understating it so badly the model never loads).

An analogy

Like a big hospital: the triage desk (router) reads your symptoms and sends you to the two or three relevant specialists, never a full-staff consultation. How many doctors the hospital employs (total parameters) decides how big the building must be (memory); what each patient actually consumes is only the on-duty doctors' time (compute).

Minimal example

示意數字,只為講清比例關係:

  密集模型:  100 份參數,每個 token 都用到全部 100 份
  MoE 模型:  100 份參數分成很多組專家,
             每個 token 只經過其中 10 份

後果:
  · 每 token 的計算量 ≈ 密集模型的十分之一 → 生成快、成本低
  · 但所有專家都得常駐記憶體 → 記憶體需求仍按 100 份算
  · 「知識容量」大(參數多)而「每步算力」小(啟用少),
    正是 MoE 想兩邊都拿的設計

The one takeaway is the asymmetry: compute per token follows activated parameters; memory follows total parameters. When a spec sheet gives only one of the two, derive the other — get either direction wrong and your local-deployment feasibility estimate becomes fiction.

What people get wrong

  • Seeing a huge total parameter count and concluding "unrunnable, necessarily expensive". Per-token compute tracks only the activated slice, so a large MoE is often far cheaper and faster than a dense model of similar total size — the intuition "big parameters = slow = expensive" misfires on MoE.
  • Conversely, seeing "few activated parameters" and assuming a small model any machine can run. Every expert must be loaded and stay resident — memory demand follows the total count and is not small at all. MoE is compute-light, memory-heavy.

Related terms

Next