Agentic Research

Transformer

Also: Transformer 架構 · 注意力機制 · attention · self-attention · 自注意力

A neural-network architecture built around attention, introduced in 2017; virtually every modern LLM is built on it.

When you will meet it

The first paragraph of nearly every model write-up says "based on the Transformer architecture". You do not need the math, but you do need to know what attention does and why it beat the older designs — otherwise everything downstream (prefill and decode, the KV cache, why long context is slow and pricey) reads like noise.

An analogy

The recurrent networks before it were a game of telephone: the message passes position by position, degrading with distance, and nothing can run concurrently. A Transformer is a round-table meeting: any two words can address each other directly and everyone deliberates at once — fast, and distance stops mattering.

Minimal example

句子:「那隻貓沒有跳上桌子,因為牠太累了。」

問題:「牠」指的是誰?

注意力機制:處理「牠」這個 token 時,模型對句子裡每個詞
各算一個注意力權重 ——「貓」拿到很高的權重,「桌子」很低。
「牠」的內部表示因此被「貓」的資訊大幅改寫。

這樣的注意力層堆疊很多層,每一層學不同的關注模式。

The point: every token can directly see every other token, and this is computed for the whole sequence in one pass. That parallelizability is why Transformers replaced recurrence and why models could grow by piling on hardware.

Inside one Transformer layer: how attention is computedFlow diagram: four tokens (I / like / to / read) each become an embedding vector, then are projected into Q, K and V. Each token's Q is dotted with every token's K to give an n×n grid of raw scores; scaling by √d_k and applying softmax turns them into attention weights that sum to 1 per row; those weights take a weighted sum of V. The result passes through a residual add and LayerNorm, a feed-forward network, then another residual add and LayerNorm to produce the layer output. Because every token scores against every other token the grid is n×n, so doubling sequence length quadruples the work. Weights shown are illustrative.Attention (the core of the layer)Iliketoreadeach token → an embedding vectorQqueryKkeyVvaluesame vectors, three projectionsIliketoreadIliketoread6111252111351243This grid is the point:every token scores itselfagainst every other token→ hence *self*-attentionweights × V → weighted sumrest of the layerresidual add + LayerNormoriginal input added backfeed-forward (FFN)each token independentlyresidual add + LayerNormthis layer's outputmany such layers stacked;one layer's output is the next's inputthe grid is n×n: double the sequence, 4× the workthis is why long context is expensive
1/9Input: four tokens
The model does not see words, it sees tokens. Here, a four-token sentence.
Step 1 of 9 Input: four tokens

What people get wrong

  • Assuming attention equals understanding. Attention computes relevance weights learned from data, not hardcoded grammar rules — it usually lands on the right things, but with no guarantee, which is one of the mechanisms behind confidently wrong outputs.
  • Assuming attention is cheap. Every token is weighed against every other token and the cost climbs steeply with sequence length — the root reason longer context is slower and pricier.

Related terms

Next