Inference
Also: 推理 · 模型推理 · model inference · 推理成本 · inference cost
Running an already-trained model to produce output — the counterpart of training, and the entirety of the computation that happens when you use a model.
When you will meet it
Every API call you make and every local chat you have is inference. Without the training/inference split you cannot make sense of three things: why APIs bill per use (each call burns compute), why the model does not "remember" what you taught it (weights are frozen), and why "cannot run the big model locally" is a memory-and-bandwidth problem, not a disk-space problem.
An analogy
Training is culinary school: long, expensive, and it changes the chef. Inference is ordering delivery: pay per order, on demand — but the chef is not transformed because you ordered the same dish three times. The vendor already paid for culinary school; you pay for each meal.
Minimal example
訓練(廠家做的事,一次性):
海量語料 → 反覆調整全部權重 → 得到模型檔案
推理(你做的事,每一次):
prompt → 載入權重(凍結,不會被你的輸入改變)→
一個 token 一個 token 產生輸出 → 按用量計費
你送給 API 的資料只影響這一次的輸出,
不會變成模型的一部分 —— 除非你另外花錢做微調。The key phrase is "weights frozen": inference is forward computation only, no learning. This is also why the bill attaches to each use — the lab paid for training once and amortized it into pricing, while inference cost accrues exactly as much as you consume.
What people get wrong
- Assuming the model "remembered" what you taught it in chat. Inference does not touch weights; in the next conversation it knows nothing about you. What feels like memory is the harness re-stuffing old turns into the prompt.
- Assuming inference is light because training is heavy, so running a model locally should be easy. Per request it is indeed vastly cheaper than training, but every token still pushes the whole model forward — bigger models make real, hard demands on memory and bandwidth.