Transformer
Also: Transformer 架構 · 注意力機制 · attention · self-attention · 自注意力
A neural-network architecture built around attention, introduced in 2017; virtually every modern LLM is built on it.
When you will meet it
The first paragraph of nearly every model write-up says "based on the Transformer architecture". You do not need the math, but you do need to know what attention does and why it beat the older designs — otherwise everything downstream (prefill and decode, the KV cache, why long context is slow and pricey) reads like noise.
An analogy
The recurrent networks before it were a game of telephone: the message passes position by position, degrading with distance, and nothing can run concurrently. A Transformer is a round-table meeting: any two words can address each other directly and everyone deliberates at once — fast, and distance stops mattering.
Minimal example
句子:「那隻貓沒有跳上桌子,因為牠太累了。」
問題:「牠」指的是誰?
注意力機制:處理「牠」這個 token 時,模型對句子裡每個詞
各算一個注意力權重 ——「貓」拿到很高的權重,「桌子」很低。
「牠」的內部表示因此被「貓」的資訊大幅改寫。
這樣的注意力層堆疊很多層,每一層學不同的關注模式。The point: every token can directly see every other token, and this is computed for the whole sequence in one pass. That parallelizability is why Transformers replaced recurrence and why models could grow by piling on hardware.
What people get wrong
- Assuming attention equals understanding. Attention computes relevance weights learned from data, not hardcoded grammar rules — it usually lands on the right things, but with no guarantee, which is one of the mechanisms behind confidently wrong outputs.
- Assuming attention is cheap. Every token is weighed against every other token and the cost climbs steeply with sequence length — the root reason longer context is slower and pricier.