Agentic Research

Cost per token (how API pricing actually works)

Also: token 計費 · API 定價 · 每 token 價格 · cost per token · 輸入輸出價格

APIs price input tokens and output tokens separately, output is usually dearer, and cached input is often discounted. An agent's bill is frightening because each turn it resends a longer and longer context.

When you will meet it

This is the first thing that makes your heart race when an AI project moves from demo to production. Without understanding the pricing structure you will make a costly mistake: assuming cost equals "the length of the answer", when the real bill comes from "the context resent every turn". The longer an agent talks, the more cost grows superlinearly — this wrong intuition will make your budget estimate absurdly low.

An analogy

Like a print shop that charges differently for the manuscript you hand over and the finished copies it prints, with each printed copy dearer. The agent's counterintuitive quirk: every time it asks, it re-hands the entire previous stack plus the new question (resending context). So the "delivery" on turn 50 costs far more than turn 1 — even if the question itself is one sentence.

Minimal example

一次 API 呼叫的帳單長這樣:
  輸入 token × 輸入單價     ← 你送進去的:系統提示 + 歷史 + 文件 + 這次問題
  輸出 token × 輸出單價     ← 模型吐出來的:通常單價比輸入高
  命中快取的輸入 × 折扣價   ← 重複的前綴(例如同一段系統提示)常有優惠

Agent 的成本陷阱(示意,非實際數字):
  第 1 輪送:系統提示 + 問題1
  第 2 輪送:系統提示 + 問題1 + 答1 + 問題2
  第 N 輪送:系統提示 + 前面所有問答 + 問題N
  → 每一輪都把「越來越長的歷史」重送一次
  → 成本隨對話長度「超線性」成長,不是線性

Two practical conclusions: output usually costs more than input, so making the model "cut the chatter and give the result" saves money; and an agent's real cost is the resent context, so context compression, prefix caching and starting fresh conversations are where the big savings are.

What people get wrong

  • Assuming cost comes mainly from "the model's answer". For a multi-turn agent the bulk of the bill is the input context resent every turn, not the output.
  • Assuming cost grows linearly with conversation length. Because every turn resends the whole history, it grows superlinearly (roughly quadratically), and long conversations suddenly get very expensive.
  • Picking a provider by comparing "price per token" alone. What matters is the total bill under your real input/output ratio and cache-hit rate; a plan with a lower unit price but worse cache discounts can actually cost more.

Related terms

Next