Agentic Research

Context window

Also: 上下文長度 · context length · 上下文視窗

The total tokens a model can see at once — and input and output share that budget.

When you will meet it

It is the direct reason an AI "forgets". It is also the line that decides whether you need RAG or context compression.

An analogy

Like a desk of fixed size. You can spread documents out, but the desk is only so big; when new material arrives, older material is moved off or covered. The model will not tell you what it covered.

Minimal example

128K 上下文大致可以放:
  系統提示 + 工具定義      約 5K~20K
  對話歷史                 越聊越多
  檢索到的文件             看你塞多少
  ─────────────────────────
  剩下來的,才是模型能寫答案的空間

Fill the context and the model has no room left to reason or answer. This is a leading cause of sudden quality drops in long conversations.

What people get wrong

  • Assuming bigger context is strictly better and that filling it is harmless. The fuller it is, the more likely the model misses the key passage in the middle.
  • Assuming overflow raises an error. Usually it is silently truncated: no warning, just a dumber answer.

Related terms

Next