Agentic Research

Top-p (nucleus sampling)

Also: 核取樣 · nucleus sampling · top_p · 累積機率取樣

The other randomness knob: rank candidate tokens by probability, keep the smallest set whose cumulative probability reaches p (say 0.9), discard the tail entirely, then sample from what remains.

When you will meet it

It shows up in the same API docs and the same snippets as temperature, so beginners tune both at once and get behaviour that is even harder to predict. Knowing the division of labour — one reshapes the distribution, the other truncates the candidate set — tells you which to touch, and why touching one is usually enough.

An analogy

A lottery only front-runners may enter: rank by odds, close the gate once cumulative odds hit ninety percent, disqualify the long tail, then draw among the qualifiers. The weather (temperature) changes everyone's odds; the gate decides who qualifies — different mechanisms that interact.

Minimal example

下一個 token 的候選與機率(示意):

  「好」    0.50   累積 0.50
  「不錯」  0.25   累積 0.75
  「行」    0.15   累積 0.90   ← top_p=0.9 的界線落在這裡
  「妙」    0.07   (淘汰)
  「絕妙」  0.03   (淘汰)

本輪只從前三個裡抽。若模型這一步非常確定
(例如某個選項機率 0.98),核裡可能只剩一個候選。

Notice the nucleus breathes: the more certain the model, the fewer candidates survive; the more torn it is, the more it keeps. top_p=1 means no truncation at all, which is what most APIs default to.

What people get wrong

  • Tuning temperature and top-p together to "find a feel". Temperature moves every candidate's probability, which changes who makes the nucleus — the two are coupled, so co-tuning is guesswork on two knobs that shift each other. Convention: move one at a time — either temperature with top-p left at 1, or a fixed temperature and only top-p.
  • Reading top_p=0.9 as "keep the top 90% of candidates by count". It is 90% of cumulative probability, not of count — the vocabulary may hold tens of thousands of candidates, yet the nucleus usually keeps only a handful.

Related terms

Next