LLM (Large Language Model)
Also: 大模型 · 大型語言模型 · foundation model · 大語言模型
A neural network trained to predict the next token. It can only output text.
When you will meet it
It is the core of every AI tool. Understanding "it is only predicting the next token" explains half of why AI gets things wrong.
An analogy
Like someone who has read an enormous amount of text and is extremely good at continuing a sentence. Whether it sounds right and whether it is right are two different questions.
Minimal example
輸入:「臺灣最高峰是」
模型內部:對詞彙表裡每個 token 算一個機率
輸出:機率最高的那個 → 「玉山」
然後把「玉山」接回輸入,再算下一個 → 一直重複到結束That loop is the whole of generation. There is no "look it up" step inside the model — unless outside code does the looking.
1/8Input: the whole prompt
System prompt, history, documents and this turn's question are concatenated into one token sequence.
Step 1 of 8 Input: the whole prompt
What people get wrong
- Assuming the model searches the web. A bare LLM does not; it can only do so because tools were attached to it.
- Assuming it knows when it is unsure. It emits the most answer-like text, not an answer tagged with confidence.
Related terms
Next
- AI Agent 是什麼:從 ChatGPT 到會自己做事的 AI12 minChinese only
- Understanding the Transformer Architecture: Attention Is All You Need7 min
Read further
Transformer 架構理解 — Attention Is All You Need