Context window
Also: 上下文長度 · context length · 上下文視窗
The total tokens a model can see at once — and input and output share that budget.
When you will meet it
It is the direct reason an AI "forgets". It is also the line that decides whether you need RAG or context compression.
An analogy
Like a desk of fixed size. You can spread documents out, but the desk is only so big; when new material arrives, older material is moved off or covered. The model will not tell you what it covered.
Minimal example
128K 上下文大致可以放:
系統提示 + 工具定義 約 5K~20K
對話歷史 越聊越多
檢索到的文件 看你塞多少
─────────────────────────
剩下來的,才是模型能寫答案的空間Fill the context and the model has no room left to reason or answer. This is a leading cause of sudden quality drops in long conversations.
What people get wrong
- Assuming bigger context is strictly better and that filling it is harmless. The fuller it is, the more likely the model misses the key passage in the middle.
- Assuming overflow raises an error. Usually it is silently truncated: no warning, just a dumber answer.