Agentic Research

Checkpoint

Also: 檢查點 · 狀態保存 · resume · 斷點恢復

A saved snapshot of a task half-finished, so work resumes from here after an interruption instead of restarting from zero.

When you will meet it

Any task longer than a few tens of minutes — deep research, batch refactors, deploy pipelines — will meet it: the network drops, the token budget dies, you press stop, or the flow needs human approval. Without checkpoints every interruption means starting over: time, tokens and already-correct steps all wasted together.

An analogy

Like a save point in a game: save before the boss, and a loss restarts you there instead of at level one. Nobody accepts replaying the whole game after every death, yet people routinely let multi-hour, token-burning agent runs go with no saves at all.

Minimal example

一個調研任務的中斷與恢復(示意):

  步驟 1 收集財報   ✅ 完成 → 結果寫入 work/step1.json
  步驟 2 股權分析   ✅ 完成 → 結果寫入 work/step2.json
  步驟 3 生成報告   ⛔ 中斷(API 超時)

  恢復(resume): 讀 work/ 下已完成的步驟 → 只重跑步驟 3
  重來(restart):三步全部從零開始,步驟 1、2 的 token 白燒

The key is that state lives outside the model: the model remembers no progress, so a checkpoint must land in files or a database and be re-assembled into context on resume. Same idea as agent memory, different payload — one stores progress, the other knowledge.

What people get wrong

  • Assuming the chat transcript is a checkpoint. A transcript records what was said, not how far the task got or where the intermediate artefacts live. Resuming needs executable state, not a replay.
  • Saving too sparsely. One save at the end is no save — an interruption between save points still loses that whole stretch. Save frequency should track the cost of redoing the work.

Related terms

Next