Agentic Research

Rate limit

Also: 速率限制 · 頻率限制 · too many requests · rate limit · quota

The cap on how many requests or tokens a service lets you send in a given window; go over it and you are throttled with a 429.

When you will meet it

You will rarely hit it by hand, but an agent that loops and fires calls in parallel burns through the quota in seconds. Understanding rate limits is why agents build in backoff, queuing and "do not send everything at once".

An analogy

Like a highway ramp meter: when traffic is heavy the light releases cars one at a time — not banning you, just pacing the flow so everyone keeps moving. Force it and you are held back (429) until the next green (Retry-After).

Minimal example

送第 1 次  → 200 OK
送第 2 次  → 200 OK
⋮
送第 61 次 → 429 Too Many Requests
             Retry-After: 30        ← 30 秒後再試

正確重試(指數退避):等 30 秒 → 若仍失敗等 60 秒 → 120 秒 ⋯
錯誤重試:立刻瘋狂重送 → 一直 429,甚至可能被暫時封鎖

Two numbers matter: requests-per-minute (RPM) and tokens-per-minute (TPM). A long prompt can trip TPM after only a few calls. Retry-After is the server telling you exactly how many seconds to wait — better than guessing.

What people get wrong

  • Retrying instantly and faster on a 429. That only keeps you blocked, and some services temporarily ban you for it. The right move is to wait and back off, not to speed up.
  • Assuming a rate limit is a total-usage cap. It is usually a per-unit-time rate; your monthly quota may be fine, but sending too densely in this minute still yields a 429.

Related terms

Next