Magic Tools
Pitfall NotesBy CooconOctober 4, 202633 views8 min read

Claude Code 429 "Request rejected (429)": How Long It Retries, and Why retry-after Over 60 Seconds Fails Instantly (Tested)

Short answers first

  • What API Error: Request rejected (429) means: the backend (Anthropic, or your gateway / relay) says you are sending too much. The part after · is the backend's own message — that tells you who is rate-limiting you.
  • Should you retry by hand? Usually no. Claude Code has already retried up to 10 times over about 3 minutes; seeing the error means you were limited the whole time.
  • It failed instantly without waiting: the backend sent a retry-after above 60 seconds, and Claude Code gave up (the threshold measured exactly between 60 and 61 seconds). Wait that long and send again, or ask your gateway to raise the limit.
  • Interactive mode keeps showing Retrying in Ns · attempt N/10: normal backoff, not a hang. Press Esc to interrupt.
  • Monthly spend-cap 429s: per the official docs these carry no retry-after and keep failing until access resumes. Retrying will not help; check your usage in the Console.

Background

On long tasks, or behind a third-party gateway, Claude Code sometimes prints:

API Error: Request rejected (429) · Number of request tokens has exceeded your per-minute rate limit

People searching for this want to know who is limiting them, whether Claude Code retries on its own, and how long to wait. The docs say the SDK retries with exponential backoff and honors retry-after, but not how many times Claude Code retries, how long it waits, or what happens when retry-after is large.

Our previous article measured the basic 429 retry curve but left gaps: recovery after a 429, the interactive display, large retry-after values, and what gateway 429s look like. This one fills them in.

Where 429s come from

All from primary sources:

Source What the 429 looks like Reference
Anthropic API {"type":"error","error":{"type":"rate_limit_error","message":"…"}}, usually with retry-after API errors
Anthropic monthly spend cap Also rate_limit_error, but no retry-after, keeps failing until access resumes Same page: A tier spend-cap 429 has no retry-after header and keeps failing until access resumes
new-api per-model limit OpenAI-style: {"error":{"message":"您已达到请求数限制:N分钟内最多请求M次 (request id: …)","type":"new_api_error","code":""}} new-api source middleware/model-rate-limit.go, middleware/utils.go (commit 1a4166d8e8)
new-api global limit Empty body, only a Retry-After header set to the full window in seconds writeRateLimited in middleware/rate-limit.go
nginx limit_req HTML error page <title>429 Too Many Requests</title> nginx default error page
DeepSeek Docs list 429 - Rate Limit Reached without a body format DeepSeek error codes

The 2.1.285 binary also contains header names like anthropic-ratelimit-unified-status and anthropic-ratelimit-unified-reset, plus strings like Usage limit reached. These belong to subscription (Pro / Max) usage limits; as shown below, they have no effect when you connect with an API key.

Method

Approach Verdict Why
Local stub backend ✅ used Exact control of status, body and retry-after; every request logged with ms timestamps; no cost, no effect on a real account
Hit real rate limits with a real account ❌ Hard to trigger reliably, retry-after can't be controlled, burns real quota, and other services on the same account may get limited too
Read strings from the binary only ❌ not alone The UI text is compiled into fragments; it shows a string exists, not when it is displayed

To be clear: the server responses are simulated; Claude Code's reaction (retries, waits, display) is real. Bodies follow the primary sources above.

Isolation, same as last time:

  • Each run uses env -i and a fresh empty CLAUDE_CONFIG_DIR, so the machine's login and session history are untouched.
  • ANTHROPIC_BASE_URL points to a stub on 127.0.0.1, with a fake ANTHROPIC_API_KEY.
  • HTTPS_PROXY also points to the stub, which logs and refuses every CONNECT. Outbound targets seen this round — api.anthropic.com, github.com, raw.githubusercontent.com, downloads.claude.ai, registry.npmmirror.com — were all blocked; nothing reached a real service.
  • -p runs use claude -p 'reply with exactly OK' --model haiku < /dev/null; interactive runs start the real TUI in tmux and capture the screen once per second.

Results

33 claude -p runs plus 3 interactive sessions. Every -p case ran at least twice with identical messages, exit codes and request counts.

1. What different 429s look like in Claude Code

With CLAUDE_CODE_MAX_RETRIES=0 (stdout verbatim, exit 1, stderr 0 bytes in every case):

Backend returns Claude Code shows
Anthropic-style rate_limit_error API Error: Request rejected (429) · Number of request tokens has exceeded your per-minute rate limit
new-api per-model limit API Error: Request rejected (429) · 您已达到请求数限制:1分钟内最多请求10次 (request id: 2026…)
new-api global limit (empty body) API Error: Request rejected (429) · status code (no body)
nginx HTML page API Error: Request rejected (429) · Too Many Requests
Anthropic-style + anthropic-ratelimit-unified-status: rejected Identical to the first row; the subscription header is ignored

2. How long it retries

Default settings, the 429 never clears (simulating a spend-cap 429 with no retry-after):

Case Requests Wall time Gaps between requests (ms)
No retry-after, run 1 11 182.30 s 618, 1118, 2129, 4920, 9818, 16225, 35616, 35118, 36829, 39009
No retry-after, run 2 11 176.41 s 536, 1093, 2371, 4760, 9361, 16311, 32934, 34650, 38356, 35746
With subscription headers × 2 11 / 11 175.94 / 176.52 s same curve

One request plus 10 retries, starting at 0.5 s and doubling, capped at 32–40 s from the 7th retry, about 3 minutes in total — the same curve we measured on 2.1.280.

3. Recovery is invisible

Three 429s, then a normal response:

Case Exit Wall time Gaps (ms) stdout
429 with retry-after: 1 ×3 → OK 0 / 0 5.06 / 4.45 s 1024, 1015, 2187 / 1011, 1061, 2110 OK
429 without retry-after ×3 → OK 0 / 0 4.73 / 4.23 s 513, 1225, 2129 / 573, 1260, 2108 OK

In -p mode, stdout is just OK, stderr is 0 bytes, exit 0. A script cannot tell it was rate-limited; it is only a few seconds slower.

The third gap in the first row is 2.1 s, not 1 s: with retry-after: 1 the wait appears to be the larger of retry-after and the exponential backoff. That is inferred from the numbers, not confirmed in code.

4. The 60-second retry-after cap

The key finding. A 429 with different retry-after values, then a normal response:

retry-after Result Requests Wall time
10 ✅ waited, recovered, exit 0 2 / 2 10.88 / 10.31 s (gaps 10018 / 10012 ms)
60 ✅ waited, recovered, exit 0 2 / 2 60.87 / 60.31 s (gaps 60022 / 60012 ms)
61 ❌ no retry, immediate error, exit 1 1 / 1 / 1 0.55 / 0.27 / 0.27 s
90, 120, 180, 300, 301, 360 ❌ same 1 each 0.53–0.59 s
600 ❌ same 1 / 1 0.88 / 0.45 s

retry-after ≤ 60 s is honored in full; ≥ 61 s is not retried at all. Interactive mode behaves the same: with retry-after: 61 the final error appeared after 0 s (Churned for 0s).

This matters behind gateways: new-api's global limiter puts the full window length in Retry-After, so any window longer than a minute makes Claude Code fail instantly — it looks like it "errored without retrying".

Searching the binary for environment variables containing RETRY, the only retry-related one is CLAUDE_CODE_MAX_RETRIES (the count). I found no setting that changes the 60-second cap.

5. What interactive mode shows

A real interactive session in tmux, backend always returning 429 (retry-after: 1), one frame per second. The status line over time:

✻ API error · Retrying in 1s · attempt 2/10
✻ 429 Number of request tokens has exceeded your per-minute rate limit · Retrying in 3s · attempt 3/10
✻ 429 Number of request tokens has exceeded your per-minute rate limit · Retrying in 9s · attempt 5/10
✻ 429 Number of request tokens has exceeded your per-minute rate limit · Retrying in 39s · attempt 7/10
…
⏺ API Error: Request rejected (429) · Number of request tokens has exceeded your per-minute rate limit
✻ Brewed for 3m 3s
  • Attempt 2 shows a generic API error; from attempt 3 on it shows 429 and the backend's message.
  • Retrying in Ns counts down every second, so a moving screen is not a hang.
  • After the final error the session stays open; you can just send again.

In a second session where the 429s cleared, the screen only flashed API error · Retrying in 2s · attempt 2/10 and then printed OK — no trace of the error in the transcript.

6. Interactive mode sends 2 requests per prompt

The stub log showed that each interactive prompt sends two concurrent /v1/messages requests: the main one with 28 tool definitions, and a side request with 0 tools whose content starts with <session>. Each retries independently.

Time (s) Main Side
6.08 / 6.09 429 429
8.09 / 8.10 429 429
10.10 / 10.10 OK OK

For a gateway that limits requests per minute, interactive mode uses twice the requests you might expect, and both retry when limited. -p mode sends one (all -p request counts above match).

What to do

1. Read the part after ·. English rate limit text: Anthropic or an Anthropic-compatible gateway. Chinese "您已达到请求数限制": a new-api style relay — raise the limit or switch groups there. status code (no body): a gateway's global limiter. Too Many Requests: an nginx in front.

2. Instant failure with no waiting usually means retry-after > 60. Check what the backend actually sends (base URL and key from your Claude Code environment; use a model your backend supports):

curl -s -o /dev/null -D - "$ANTHROPIC_BASE_URL/v1/messages" \
  -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-haiku-4-5","max_tokens":8,"messages":[{"role":"user","content":"hi"}]}' \
  | grep -iE '^HTTP|retry-after'

This uses one request of your quota. If retry-after is above 60, wait that long, or ask the gateway to shorten its window.

3. Scripts / CI: if the 429 clears within ~3 minutes, -p output and exit code are completely normal. For fast failure (health checks), use CLAUDE_CODE_MAX_RETRIES=0; don't use 0 for long jobs.

4. Gateway operators: keep Retry-After at 60 or below if you want Claude Code to wait; above that it fails immediately. Count 2 requests per interactive prompt when limiting by request count.

5. Monthly spend cap: these 429s have no retry-after, so Claude Code retries for 3 minutes and then fails. Retrying won't fix it; check usage or spend limits in the Console.

Every request in this test went to the local stub; real API cost was zero.

Pitfalls along the way

  • A watchdog sleep held the pipe open. The runner's timeout watchdog was a background subshell; killing the subshell orphaned its sleep 900, which kept stdout open. Fine when writing to a file, but piping to grep hung for 15 minutes. Fixed with exec >/dev/null in the subshell and a trap that kills the sleep.
  • The two interactive requests split the 429 sequence. The first recovery run was planned as "4 × 429, then OK" and finished twice as fast as expected; the log showed the main and side requests had each taken 2. That is how section 6 was found.
  • Binary strings are fragments. Usage limit reached, Retrying in and too far out to wait for are all there, but no full sentences; the captured screens are the source of truth.

Not verified: the real Anthropic 429 body text (this article reuses the message from the previous one); DeepSeek's 429 body (not documented); the usage-limit UI for OAuth subscription logins (Pro / Max) — only API-key mode was tested, where the subscription headers are ignored; boundary jitter for retry-after between 59 and 60; retry-after given as an HTTP date instead of seconds. Each -p case ran only 2–3 times.

Get field notes like this every Saturday

Subscribe to Dev Breakfast: daily AI coding picks at 8:00, plus a Saturday roundup of this week's hands-on tests with Claude Code / Codex / local models. Written in Chinese.

Related Articles

Claude Code "Prompt is too long" vs "maximum context length": Why Auto-Compact Works for One and Not the Other (Tested)

Claude Code decides a request was too long by matching the error text. If the backend says prompt is too long or input is too long for requested model, it auto-compacts the conversation and retries — invisible to you. DeepSeek and OpenAI-style gateways say This model's maximum context length is …, which it doesn't recognize: you get API Error: 400 on every turn. Manual /compact works; the better fix is to declare the real window with CLAUDE_CODE_MAX_CONTEXT_TOKENS (non-claude- model IDs) or CLAUDE_CODE_AUTO_COMPACT_WINDOW (claude- IDs) so it compacts before hitting the limit. Tested on Claude Code 2.1.285 against a local stub.

claude-codedeepseek+6
pitfallsOct 5, 20267 min
47

Claude Code Stuck on a Spinner With No Response: It Waits 6 Minutes per Attempt, and 10 Retries Can Hang It for Over an Hour (Tested)

When Claude Code spins without output or sits at Retrying in 0s, the backend is usually not sending anything. Tested on 2.1.285 (API key + ANTHROPIC_BASE_URL): with no response headers it waits 6 minutes (360 s) per attempt before timing out — setting API_TIMEOUT_MS to 600000 or 900000 doesn't change that; only lower values work. With the default 10 retries it can hang for over an hour. A gateway that buffers the reply until generation finishes will never deliver a reply that takes over 6 minutes. Press Esc to interrupt. All tested against a local stub.

claude-codetroubleshooting+5
pitfallsOct 5, 20269 min
34

Clash Verge TUN Mode Breaks All Internet Access: Hysteria2 Traffic Loops Back Into the TUN, and Tailscale Hijacks DNS

Clash Verge Rev 2.5.6 on macOS works fine in system-proxy mode, but as soon as TUN (virtual network adapter) mode is on, not even Baidu loads. Debugging through the mihomo core's API turned up two independent root causes stacked together. First, Hysteria2's outbound UDP isn't bound to the physical interface, so the TUN routes pull it back in and it loops. Second, Tailscale MagicDNS (100.100.100.100) has taken over system DNS, so queries go out through Tailscale's utun where Clash's dns-hijack can't see them, and come back poisoned. The fix is two Merge overrides: route-exclude-address to keep node IPs out of the TUN, and sniffer to recover domains from SNI. Every step comes with the commands and real output.

troubleshootingtailscale+7
pitfallsOct 4, 20266 min
51
Claude Code MCP Shows Connected but 0 Tools: "Invalid result for tools/list" (ttlMs / cacheScope), Reproduced and Fixed

Claude Code MCP Shows Connected but 0 Tools: "Invalid result for tools/list" (ttlMs / cacheScope), Reproduced and Fixed

An MCP server shows as connected but exposes 0 tools, and the log says Invalid result for tools/list with ttlMs and cacheScope failing validation. Reproduced with a stub server on Claude Code 2.1.280, 2.1.285 and 2.1.288: the cause is not extra fields being rejected. The server negotiated MCP 2026-07-28 and then left out fields that revision requires (resultType, ttlMs, cacheScope); adding unknown fields works fine. Whether stdio uses the new protocol is decided by a remote flag that is off by default, so the same version breaks for some users and not others, and 2.1.280 fails the same way once negotiation is on. MCP_PROTOCOL_NEGOTIATION=legacy (also works in settings.json env) restores the tools; servers fix it by adding the three fields.

mcpclaude-code+4
pitfallsOct 3, 20268 min
37

Published by Magic Tools