Checkpoint
Also: 檢查點 · 狀態保存 · resume · 斷點恢復
A saved snapshot of a task half-finished, so work resumes from here after an interruption instead of restarting from zero.
When you will meet it
Any task longer than a few tens of minutes — deep research, batch refactors, deploy pipelines — will meet it: the network drops, the token budget dies, you press stop, or the flow needs human approval. Without checkpoints every interruption means starting over: time, tokens and already-correct steps all wasted together.
An analogy
Like a save point in a game: save before the boss, and a loss restarts you there instead of at level one. Nobody accepts replaying the whole game after every death, yet people routinely let multi-hour, token-burning agent runs go with no saves at all.
Minimal example
一個調研任務的中斷與恢復(示意):
步驟 1 收集財報 ✅ 完成 → 結果寫入 work/step1.json
步驟 2 股權分析 ✅ 完成 → 結果寫入 work/step2.json
步驟 3 生成報告 ⛔ 中斷(API 超時)
恢復(resume): 讀 work/ 下已完成的步驟 → 只重跑步驟 3
重來(restart):三步全部從零開始,步驟 1、2 的 token 白燒The key is that state lives outside the model: the model remembers no progress, so a checkpoint must land in files or a database and be re-assembled into context on resume. Same idea as agent memory, different payload — one stores progress, the other knowledge.
What people get wrong
- Assuming the chat transcript is a checkpoint. A transcript records what was said, not how far the task got or where the intermediate artefacts live. Resuming needs executable state, not a replay.
- Saving too sparsely. One save at the end is no save — an interruption between save points still loses that whole stretch. Save frequency should track the cost of redoing the work.