Agentic Research

Inference

Also: 推理 · 模型推理 · model inference · 推理成本 · inference cost

Running an already-trained model to produce output — the counterpart of training, and the entirety of the computation that happens when you use a model.

When you will meet it

Every API call you make and every local chat you have is inference. Without the training/inference split you cannot make sense of three things: why APIs bill per use (each call burns compute), why the model does not "remember" what you taught it (weights are frozen), and why "cannot run the big model locally" is a memory-and-bandwidth problem, not a disk-space problem.

An analogy

Training is culinary school: long, expensive, and it changes the chef. Inference is ordering delivery: pay per order, on demand — but the chef is not transformed because you ordered the same dish three times. The vendor already paid for culinary school; you pay for each meal.

Minimal example

訓練(廠家做的事,一次性):
  海量語料 → 反覆調整全部權重 → 得到模型檔案

推理(你做的事,每一次):
  prompt → 載入權重(凍結,不會被你的輸入改變)→
  一個 token 一個 token 產生輸出 → 按用量計費

你送給 API 的資料只影響這一次的輸出,
不會變成模型的一部分 —— 除非你另外花錢做微調。

The key phrase is "weights frozen": inference is forward computation only, no learning. This is also why the bill attaches to each use — the lab paid for training once and amortized it into pricing, while inference cost accrues exactly as much as you consume.

LLM inference: the prefill and decode phasesFlow diagram: generation happens in two phases. Prefill processes the whole prompt in one parallel pass, storing each token's K and V in the KV cache; this phase is compute-bound and determines TTFT. Decode then produces one token at a time — read the cache, one forward pass, append the new token, repeat — until the model emits EOS or the output limit is hit; this phase is memory-bandwidth-bound and determines tokens per second. The timeline at the bottom shows the two bottlenecks differ, so shortening the prompt helps TTFT while making the model terser helps total time.Phase 1 · PREFILL (whole prompt, one parallel pass)explain this code to me plea… N tokens in totalone forward pass, all tokens at oncecompute-bound: limited by compute, so a longer prompt makes this slowerKV cacheK and V per token, computed once and keptprefill result storedPhase 2 · DECODE (one token at a time, until done)read cache + new tokenmemory-bandwidth-boundemit one tokenand append it to the cachethe token just produced becomes the next input — this is autoregressiononly two reasons to stop:the model emits EOS, or max output is hitwhat the user actually experiencesTTFT ← set by prompt lengthtokens/second ← set by output lengthDifferent bottlenecks, so different fixes: shorten the prompt (or use prefix caching) for TTFT; make the model terser for total time.Collapsing both into one "speed" number means optimising the wrong thing.
1/8Input: the whole prompt
System prompt, history, documents and this turn's question are concatenated into one token sequence.
Step 1 of 8 Input: the whole prompt

What people get wrong

  • Assuming the model "remembered" what you taught it in chat. Inference does not touch weights; in the next conversation it knows nothing about you. What feels like memory is the harness re-stuffing old turns into the prompt.
  • Assuming inference is light because training is heavy, so running a model locally should be easy. Per request it is indeed vastly cheaper than training, but every token still pushes the whole model forward — bigger models make real, hard demands on memory and bandwidth.

Related terms

Next