Claude Code 429 "Request rejected (429)": How Long It Retries, and Why retry-after Over 60 Seconds Fails Instantly (Tested)
Short answers first
- What
API Error: Request rejected (429)means: the backend (Anthropic, or your gateway / relay) says you are sending too much. The part after·is the backend's own message — that tells you who is rate-limiting you.- Should you retry by hand? Usually no. Claude Code has already retried up to 10 times over about 3 minutes; seeing the error means you were limited the whole time.
- It failed instantly without waiting: the backend sent a
retry-afterabove 60 seconds, and Claude Code gave up (the threshold measured exactly between 60 and 61 seconds). Wait that long and send again, or ask your gateway to raise the limit.- Interactive mode keeps showing
Retrying in Ns · attempt N/10: normal backoff, not a hang. PressEscto interrupt.- Monthly spend-cap 429s: per the official docs these carry no
retry-afterand keep failing until access resumes. Retrying will not help; check your usage in the Console.
Background
On long tasks, or behind a third-party gateway, Claude Code sometimes prints:
API Error: Request rejected (429) · Number of request tokens has exceeded your per-minute rate limit
People searching for this want to know who is limiting them, whether Claude Code retries on its own, and how long to wait. The docs say the SDK retries with exponential backoff and honors retry-after, but not how many times Claude Code retries, how long it waits, or what happens when retry-after is large.
Our previous article measured the basic 429 retry curve but left gaps: recovery after a 429, the interactive display, large retry-after values, and what gateway 429s look like. This one fills them in.
Where 429s come from
All from primary sources:
| Source | What the 429 looks like | Reference |
|---|---|---|
| Anthropic API | {"type":"error","error":{"type":"rate_limit_error","message":"…"}}, usually with retry-after |
API errors |
| Anthropic monthly spend cap | Also rate_limit_error, but no retry-after, keeps failing until access resumes |
Same page: A tier spend-cap 429 has no retry-after header and keeps failing until access resumes |
| new-api per-model limit | OpenAI-style: {"error":{"message":"您已达到请求数限制:N分钟内最多请求M次 (request id: …)","type":"new_api_error","code":""}} |
new-api source middleware/model-rate-limit.go, middleware/utils.go (commit 1a4166d8e8) |
| new-api global limit | Empty body, only a Retry-After header set to the full window in seconds |
writeRateLimited in middleware/rate-limit.go |
nginx limit_req |
HTML error page <title>429 Too Many Requests</title> |
nginx default error page |
| DeepSeek | Docs list 429 - Rate Limit Reached without a body format |
DeepSeek error codes |
The 2.1.285 binary also contains header names like anthropic-ratelimit-unified-status and anthropic-ratelimit-unified-reset, plus strings like Usage limit reached. These belong to subscription (Pro / Max) usage limits; as shown below, they have no effect when you connect with an API key.
Method
| Approach | Verdict | Why |
|---|---|---|
| Local stub backend | ✅ used | Exact control of status, body and retry-after; every request logged with ms timestamps; no cost, no effect on a real account |
| Hit real rate limits with a real account | ❌ | Hard to trigger reliably, retry-after can't be controlled, burns real quota, and other services on the same account may get limited too |
| Read strings from the binary only | ❌ not alone | The UI text is compiled into fragments; it shows a string exists, not when it is displayed |
To be clear: the server responses are simulated; Claude Code's reaction (retries, waits, display) is real. Bodies follow the primary sources above.
Isolation, same as last time:
- Each run uses
env -iand a fresh emptyCLAUDE_CONFIG_DIR, so the machine's login and session history are untouched. ANTHROPIC_BASE_URLpoints to a stub on127.0.0.1, with a fakeANTHROPIC_API_KEY.HTTPS_PROXYalso points to the stub, which logs and refuses everyCONNECT. Outbound targets seen this round — api.anthropic.com, github.com, raw.githubusercontent.com, downloads.claude.ai, registry.npmmirror.com — were all blocked; nothing reached a real service.-pruns useclaude -p 'reply with exactly OK' --model haiku < /dev/null; interactive runs start the real TUI in tmux and capture the screen once per second.
Results
33 claude -p runs plus 3 interactive sessions. Every -p case ran at least twice with identical messages, exit codes and request counts.
1. What different 429s look like in Claude Code
With CLAUDE_CODE_MAX_RETRIES=0 (stdout verbatim, exit 1, stderr 0 bytes in every case):
| Backend returns | Claude Code shows |
|---|---|
Anthropic-style rate_limit_error |
API Error: Request rejected (429) · Number of request tokens has exceeded your per-minute rate limit |
| new-api per-model limit | API Error: Request rejected (429) · 您已达到请求数限制:1分钟内最多请求10次 (request id: 2026…) |
| new-api global limit (empty body) | API Error: Request rejected (429) · status code (no body) |
| nginx HTML page | API Error: Request rejected (429) · Too Many Requests |
Anthropic-style + anthropic-ratelimit-unified-status: rejected |
Identical to the first row; the subscription header is ignored |
2. How long it retries
Default settings, the 429 never clears (simulating a spend-cap 429 with no retry-after):
| Case | Requests | Wall time | Gaps between requests (ms) |
|---|---|---|---|
No retry-after, run 1 |
11 | 182.30 s | 618, 1118, 2129, 4920, 9818, 16225, 35616, 35118, 36829, 39009 |
No retry-after, run 2 |
11 | 176.41 s | 536, 1093, 2371, 4760, 9361, 16311, 32934, 34650, 38356, 35746 |
| With subscription headers × 2 | 11 / 11 | 175.94 / 176.52 s | same curve |
One request plus 10 retries, starting at 0.5 s and doubling, capped at 32–40 s from the 7th retry, about 3 minutes in total — the same curve we measured on 2.1.280.
3. Recovery is invisible
Three 429s, then a normal response:
| Case | Exit | Wall time | Gaps (ms) | stdout |
|---|---|---|---|---|
429 with retry-after: 1 ×3 → OK |
0 / 0 | 5.06 / 4.45 s | 1024, 1015, 2187 / 1011, 1061, 2110 | OK |
429 without retry-after ×3 → OK |
0 / 0 | 4.73 / 4.23 s | 513, 1225, 2129 / 573, 1260, 2108 | OK |
In -p mode, stdout is just OK, stderr is 0 bytes, exit 0. A script cannot tell it was rate-limited; it is only a few seconds slower.
The third gap in the first row is 2.1 s, not 1 s: with retry-after: 1 the wait appears to be the larger of retry-after and the exponential backoff. That is inferred from the numbers, not confirmed in code.
4. The 60-second retry-after cap
The key finding. A 429 with different retry-after values, then a normal response:
retry-after |
Result | Requests | Wall time |
|---|---|---|---|
| 10 | ✅ waited, recovered, exit 0 | 2 / 2 | 10.88 / 10.31 s (gaps 10018 / 10012 ms) |
| 60 | ✅ waited, recovered, exit 0 | 2 / 2 | 60.87 / 60.31 s (gaps 60022 / 60012 ms) |
| 61 | ❌ no retry, immediate error, exit 1 | 1 / 1 / 1 | 0.55 / 0.27 / 0.27 s |
| 90, 120, 180, 300, 301, 360 | ❌ same | 1 each | 0.53–0.59 s |
| 600 | ❌ same | 1 / 1 | 0.88 / 0.45 s |
retry-after ≤ 60 s is honored in full; ≥ 61 s is not retried at all. Interactive mode behaves the same: with retry-after: 61 the final error appeared after 0 s (Churned for 0s).
This matters behind gateways: new-api's global limiter puts the full window length in Retry-After, so any window longer than a minute makes Claude Code fail instantly — it looks like it "errored without retrying".
Searching the binary for environment variables containing RETRY, the only retry-related one is CLAUDE_CODE_MAX_RETRIES (the count). I found no setting that changes the 60-second cap.
5. What interactive mode shows
A real interactive session in tmux, backend always returning 429 (retry-after: 1), one frame per second. The status line over time:
✻ API error · Retrying in 1s · attempt 2/10
✻ 429 Number of request tokens has exceeded your per-minute rate limit · Retrying in 3s · attempt 3/10
✻ 429 Number of request tokens has exceeded your per-minute rate limit · Retrying in 9s · attempt 5/10
✻ 429 Number of request tokens has exceeded your per-minute rate limit · Retrying in 39s · attempt 7/10
…
⏺ API Error: Request rejected (429) · Number of request tokens has exceeded your per-minute rate limit
✻ Brewed for 3m 3s
- Attempt 2 shows a generic
API error; from attempt 3 on it shows429and the backend's message. Retrying in Nscounts down every second, so a moving screen is not a hang.- After the final error the session stays open; you can just send again.
In a second session where the 429s cleared, the screen only flashed API error · Retrying in 2s · attempt 2/10 and then printed OK — no trace of the error in the transcript.
6. Interactive mode sends 2 requests per prompt
The stub log showed that each interactive prompt sends two concurrent /v1/messages requests: the main one with 28 tool definitions, and a side request with 0 tools whose content starts with <session>. Each retries independently.
| Time (s) | Main | Side |
|---|---|---|
| 6.08 / 6.09 | 429 | 429 |
| 8.09 / 8.10 | 429 | 429 |
| 10.10 / 10.10 | OK | OK |
For a gateway that limits requests per minute, interactive mode uses twice the requests you might expect, and both retry when limited. -p mode sends one (all -p request counts above match).
What to do
1. Read the part after ·. English rate limit text: Anthropic or an Anthropic-compatible gateway. Chinese "您已达到请求数限制": a new-api style relay — raise the limit or switch groups there. status code (no body): a gateway's global limiter. Too Many Requests: an nginx in front.
2. Instant failure with no waiting usually means retry-after > 60. Check what the backend actually sends (base URL and key from your Claude Code environment; use a model your backend supports):
curl -s -o /dev/null -D - "$ANTHROPIC_BASE_URL/v1/messages" \
-H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-haiku-4-5","max_tokens":8,"messages":[{"role":"user","content":"hi"}]}' \
| grep -iE '^HTTP|retry-after'
This uses one request of your quota. If retry-after is above 60, wait that long, or ask the gateway to shorten its window.
3. Scripts / CI: if the 429 clears within ~3 minutes, -p output and exit code are completely normal. For fast failure (health checks), use CLAUDE_CODE_MAX_RETRIES=0; don't use 0 for long jobs.
4. Gateway operators: keep Retry-After at 60 or below if you want Claude Code to wait; above that it fails immediately. Count 2 requests per interactive prompt when limiting by request count.
5. Monthly spend cap: these 429s have no retry-after, so Claude Code retries for 3 minutes and then fails. Retrying won't fix it; check usage or spend limits in the Console.
Every request in this test went to the local stub; real API cost was zero.
Pitfalls along the way
- A watchdog
sleepheld the pipe open. The runner's timeout watchdog was a background subshell; killing the subshell orphaned itssleep 900, which kept stdout open. Fine when writing to a file, but piping togrephung for 15 minutes. Fixed withexec >/dev/nullin the subshell and atrapthat kills thesleep. - The two interactive requests split the 429 sequence. The first recovery run was planned as "4 × 429, then OK" and finished twice as fast as expected; the log showed the main and side requests had each taken 2. That is how section 6 was found.
- Binary strings are fragments.
Usage limit reached,Retrying inandtoo far out to wait forare all there, but no full sentences; the captured screens are the source of truth.
Not verified: the real Anthropic 429 body text (this article reuses the message from the previous one); DeepSeek's 429 body (not documented); the usage-limit UI for OAuth subscription logins (Pro / Max) — only API-key mode was tested, where the subscription headers are ignored; boundary jitter for retry-after between 59 and 60; retry-after given as an HTTP date instead of seconds. Each -p case ran only 2–3 times.