Frontier model
Also: 前沿模型 · 旗艦模型 · frontier model · SOTA · 最先進模型
The small set of most capable models at a given moment — a moving label, not a fixed technical specification.
When you will meet it
Reading any benchmark, pricing page or model-selection piece, you will keep meeting "frontier", "flagship" and "SOTA". Without understanding the label's nature you treat marketing as engineering: either reflexively buying the priciest tier (many tasks never use frontier capability, so the money is wasted), or anxiously re-platforming at every launch — when the actual discipline is re-evaluating on your own tasks periodically.
An analogy
Like a "flagship phone": genuinely the strongest and priciest on launch day, then surpassed within months, repriced and slotted into the mid-range. Flagship is a shelf position, not a permanent property of the device — the same holds for frontier models, with even faster turnover.
Minimal example
供應商的模型選單通常長這樣(示意):
前沿/旗艦層 能力最強、單價最高、通常也最慢;
新能力(更長的上下文、思考模式、多模態、
更強的 Agent 工具使用)先落在這一層
中階層 上一代的旗艦,或刻意平衡成本的新模型
輕量層 最快最便宜,能力上限明顯較低
幾季之後再看同一張選單:層級還在,成員已經換過一輪。Notice "frontier" encodes a relative position: it always means "the current top tier", so the word itself carries zero absolute capability information. Last year's frontier can underperform this year's mid-tier — whenever you see the label, ask: "frontier as of when?"
What people get wrong
- Assuming "frontier" means "best for your task". Frontier models lead on broad public evaluations; your actual workload may be narrow and concrete, where a mid-tier or small model matches or wins at a fraction of the cost. Without your own eval set nobody knows the answer — not even the vendor.
- Comparing vendor-published benchmark scores as if frontier-tier models rank against each other. Eval-set selection, harness and even contamination differ; scores are comparable only within one methodology — and vendor-published numbers almost never share one.