Temperature
Also: 溫度 · 隨機性 · sampling temperature · temperature 參數
The knob for sampling randomness: lower values push the model toward the highest-probability token; higher values give unlikely candidates a real chance.
When you will meet it
The first API snippet you copy probably contains a temperature parameter, and most people paste it for a long time without knowing what it tunes. Without it you cannot explain two things: why the same question gets a different answer each time, and why tasks like classification, extraction or test runs need it turned down before "reproducible" means anything.
An analogy
Like an actor's improv dial: same script, dial at minimum and they deliver every line verbatim — steady but stiff; dial up and they riff, which can sparkle or derail. Note that the dial does not change the actor's skill — temperature does not make the model smarter either.
Minimal example
from openai import OpenAI
client = OpenAI()
# 抽取任務:要穩定、可重複 → 低溫度
client.chat.completions.create(
model="...",
messages=[{"role": "user", "content": "從這段文字抽出所有日期,只輸出 JSON"}],
temperature=0,
)
# 發想任務:要多樣 → 提高溫度
client.chat.completions.create(
model="...",
messages=[{"role": "user", "content": "幫這個產品想十個不一樣的名字"}],
temperature=1.0,
)Mechanically, temperature divides the logits before softmax: below 1 sharpens the distribution (top choices dominate), above 1 flattens it (tail choices rise), and 0 usually means taking the single most probable token (greedy decoding).
What people get wrong
- Assuming temperature=0 guarantees identical outputs. It only means "as greedy as possible": server-side batch composition, floating-point addition order, GPU and inference-stack versions, or a provider silently re-routing model versions can each change the result. Zero is one ingredient of reproducibility, never a guarantee.
- Treating temperature as a quality dial to jiggle when output is meh. It changes how tokens are drawn from the same distribution, not the distribution itself — whatever the model cannot do, it cannot do at any temperature. For real improvement, change the prompt, change the model, or add retrieval.
Related terms
Next
- 第一次呼叫 LLM API:Token、計費與常見錯誤21 minChinese only