Agentic Research

LLM (Large Language Model)

Also: 大模型 · 大型語言模型 · foundation model · 大語言模型

A neural network trained to predict the next token. It can only output text.

When you will meet it

It is the core of every AI tool. Understanding "it is only predicting the next token" explains half of why AI gets things wrong.

An analogy

Like someone who has read an enormous amount of text and is extremely good at continuing a sentence. Whether it sounds right and whether it is right are two different questions.

Minimal example

輸入:「臺灣最高峰是」
模型內部:對詞彙表裡每個 token 算一個機率
輸出:機率最高的那個 → 「玉山」

然後把「玉山」接回輸入,再算下一個 → 一直重複到結束

That loop is the whole of generation. There is no "look it up" step inside the model — unless outside code does the looking.

LLM inference: the prefill and decode phasesFlow diagram: generation happens in two phases. Prefill processes the whole prompt in one parallel pass, storing each token's K and V in the KV cache; this phase is compute-bound and determines TTFT. Decode then produces one token at a time — read the cache, one forward pass, append the new token, repeat — until the model emits EOS or the output limit is hit; this phase is memory-bandwidth-bound and determines tokens per second. The timeline at the bottom shows the two bottlenecks differ, so shortening the prompt helps TTFT while making the model terser helps total time.Phase 1 · PREFILL (whole prompt, one parallel pass)explain this code to me plea… N tokens in totalone forward pass, all tokens at oncecompute-bound: limited by compute, so a longer prompt makes this slowerKV cacheK and V per token, computed once and keptprefill result storedPhase 2 · DECODE (one token at a time, until done)read cache + new tokenmemory-bandwidth-boundemit one tokenand append it to the cachethe token just produced becomes the next input — this is autoregressiononly two reasons to stop:the model emits EOS, or max output is hitwhat the user actually experiencesTTFT ← set by prompt lengthtokens/second ← set by output lengthDifferent bottlenecks, so different fixes: shorten the prompt (or use prefix caching) for TTFT; make the model terser for total time.Collapsing both into one "speed" number means optimising the wrong thing.
1/8Input: the whole prompt
System prompt, history, documents and this turn's question are concatenated into one token sequence.
Step 1 of 8 Input: the whole prompt

What people get wrong

  • Assuming the model searches the web. A bare LLM does not; it can only do so because tools were attached to it.
  • Assuming it knows when it is unsure. It emits the most answer-like text, not an answer tagged with confidence.

Related terms

Next