Agentic Research

Tokenization

Also: 分詞 · 切詞 · tokenizer · BPE · subword · 子詞 · token 化

The rule set that chops text into tokens — and every model family chops differently.

When you will meet it

The token entry told you what the model eats; this one explains how the food gets cut. When the same sentence costs different token counts in different models, or your Chinese prompt bills more than expected, the answer lives in the cutting rules. Without them, token counts are guesswork and your estimates never reconcile.

An analogy

Like an IME's phrase memory: the more frequent a phrase, the likelier it is stored whole and typed in one go, while rare words must be spelled out piece by piece. BPE-style algorithms follow exactly this logic — the more frequent a string in the training corpus, the likelier it becomes one whole token.

Minimal example

BPE 的基本思路(示意):
  從單一字元出發,反覆統計語料裡「哪兩個相鄰單位最常一起出現」,
  就把它們合併成一個新單位,直到詞彙表到達設定的大小。

於是常見的切法像這樣(示意,實際依模型詞彙表而異):
  "unhappiness"  → ["un", "happi", "ness"]
  罕見的人名      → 往往被切成更多碎片
  常用短詞        → 多半整詞就是一個 token

Two key consequences. First, the vocabulary is built from frequency statistics over the training corpus, which for most models is English-heavy — so the same meaning usually costs more tokens in Chinese. Second, a space is just another character, and BPE merges the space before a word into that word's first token — which is why you see leading spaces inside token strings.

What people get wrong

  • Assuming there is one standard tokenization. Every model family ships its own tokenizer and vocabulary, so the same text yields different token counts, bills and context usage. For exact numbers, only that model's own tokenizer will tell you.
  • Estimating tokens by multiplying a word or character count by a factor, then wondering why it disagrees with the API's reported usage. The factor is an average that breaks badly on code, JSON and mixed-language text.

Related terms

Next