Magic Tools
Pitfall NotesBy CooconOctober 5, 20268 views9 min read

Claude Code Stuck on a Spinner With No Response: It Waits 6 Minutes per Attempt, and 10 Retries Can Hang It for Over an Hour (Tested)

Short answers first

  • Spinner with a growing timer (Clauding… (4m 34s)): Claude Code is waiting for the backend, which hasn't sent a byte. With an API key and ANTHROPIC_BASE_URL (gateway / relay), it waits a full 6 minutes per attempt before timing out.
  • Stuck at Retrying in 0s · attempt 1/10: the previous attempt timed out; the retry is now waiting another 6 minutes. It is not frozen.
  • How long can it hang? With the default 10 retries and a backend that never answers, it took about 69 minutes (4147.78 s) to finally show Request timed out.
  • What to do now: press Esc and send again. If it keeps happening, the problem is the backend or network — test your base URL directly with curl (see "What to do", item 1).
  • Fail sooner: set API_TIMEOUT_MS=60000 (only lowering works; raising doesn't) and CLAUDE_CODE_MAX_RETRIES=2.
  • Only long tasks hang, short questions are fine: your gateway probably buffers the reply until generation finishes. A reply that takes over 6 minutes is cut off and retried every time — it never arrives. Use a gateway that streams, or split the task.
  • Stuck on a command (the UI shows a Bash call, not a spinner): different problem — see Command timed out after 2m 0s.

Background

"Claude Code is stuck" usually means a spinner verb and a timer, with no error:

❯ refactor this function
✢ Clauding… (4m 34s)

After a while it may change to this and then stop moving again:

✻ API error · Retrying in 0s · attempt 1/10

You can't tell whether it is working, waiting on the network, or dead. The docs list several timers under Streaming idle watchdogs and No response from API, each with different conditions. So we reproduced each kind of stall and measured how long Claude Code waits and what it shows.

What "stuck" can mean on the wire

Where it stalls What the backend does Timer in the docs
① No response headers Accepts the connection but never sends HTTP headers (upstream queueing, a buffering gateway waiting for generation to finish) The first-byte deadline (docs: does not run when ANTHROPIC_BASE_URL is set), otherwise API_TIMEOUT_MS (docs: default 10 minutes)
② Headers, then silence Sends 200 and SSE headers, then nothing Byte-level watchdog (300 s default on gateways), event-level watchdog (300 s)
③ Keep-alive pings only Keeps sending event: ping, no content Pings reset the byte-level watchdog; the event-level watchdog ignores them
④ Stalls mid-response Sends part of the answer, then stops Byte-level watchdog

Relevant doc quotes:

  • API_TIMEOUT_MS: Timeout for API requests in milliseconds (default: 600000, or 10 minutes) (env vars)
  • The first-byte deadline runs on the direct Anthropic API but not when ANTHROPIC_BASE_URL … routes them through a gateway (network-config)
  • Claude Code retries transient failures up to 10 times with exponential backoff (errors)

Method

Approach Verdict Why
Local stub reproducing each stall ✅ used Exact control over no-headers / silence / pings / mid-stream stall; ms-timestamped request log; Claude Code's --debug-file shows which timer fired
Pull the network cable, use a real gateway ❌ Not controllable, can't tell the cases apart, burns quota
Infer from docs alone ❌ not alone One measurement disagrees with the docs (the 6 minutes below)

Isolation is the same as in the 429 article: env -i, a fresh empty CLAUDE_CONFIG_DIR per run, a fake key, ANTHROPIC_BASE_URL pointing to a stub on 127.0.0.1, and HTTPS_PROXY pointing to the stub to block all outbound connections. The debug logs show telemetry (1P event logging), bootstrap and MCP registry connections all refused — nothing reached a real service.

-p runs use claude -p 'reply with exactly OK' --model haiku < /dev/null; interactive runs start the real TUI in tmux and capture a frame every 10 seconds for 25 minutes.

Results

1. No response: 6 minutes per attempt, not 10

The stub accepts the request and never sends headers. To see a single attempt, CLAUDE_CODE_MAX_RETRIES=0:

Setting Result Wall time
API_TIMEOUT_MS unset Request timed out, exit 1 360.49 s
API_TIMEOUT_MS=600000 same 360.50 s
API_TIMEOUT_MS=900000 same 360.49 s
Headers then silence (②), API_TIMEOUT_MS=900000 same 360.49 s
API_TIMEOUT_MS=20000 same 20.6 s per attempt

Unset, it is 6 minutes, not the documented 10. Raising it to 600,000 or 900,000 ms still gives 6 minutes; only lowering takes effect.

To make sure the 360 s wasn't an artifact of the setup, more controls — all still 360 s:

Control Purpose Wall time
CLAUDE_STREAM_FIRST_BYTE_TIMEOUT_MS=60000 Is the first-byte deadline active? 360.50 s (no effect, consistent with the docs for a custom base URL)
CLAUDE_ENABLE_BYTE_WATCHDOG=0 Byte-level watchdog off 360.50 s
--tools "", request body 183 KB → 91 KB Does it scale with body size? 360.48 s
All server timeouts in the Node stub disabled Rule out the stub closing the connection 360.44 s

The last one matters: Node's HTTP server defaults to a 300 s requestTimeout with a 30 s check interval, suspiciously close to 360. With those disabled it is still 360 s, and Claude Code's own log says API api_timeout after retries: Request timed out, so this is Claude Code's behavior. In the 2.1.285 binary the SDK client timeout does read API_TIMEOUT_MS with a 600000 default; I could not locate where the 360 s ceiling comes from, so this article reports the measured value only.

2. Default 10 retries: over an hour

No variables set, backend never answers — the typical "it's stuck" case. Request timeline (seconds):

0, 361, 722, 1085, 1450, 1819, …

A new attempt every 6 minutes, with gaps growing slightly later on, up to 399 s (likely exponential backoff on top of the 6 minutes; not confirmed in code). Full timeline:

0, 361, 722, 1085, 1450, 1819, 2198, 2592, 2991, 3389, 3787

After the 11th request waited its 360 s, the run failed with Request timed out at 4147.78 s (about 69 minutes), exit 1, 11 requests at the stub. The last debug line is API api_timeout after retries: Request timed out. With a backend that never answers, default Claude Code takes 69 minutes to give up.

For comparison, with API_TIMEOUT_MS=20000 both runs sent 11 requests (1 + 10 retries) and failed with Request timed out at 406.74 s and 407.87 s. Same 10 retries as in the 429 article, except each one waits out the full timeout first.

3. What interactive mode shows

A real session in tmux, backend never answering, one frame every 10 s for 25 minutes:

Time Status line
0–6 min ✢ Clauding… (4m 34s), only the timer moves (another session showed Deliberating…; the verb is random)
6–12 min ✻ API error · Retrying in 0s · attempt 1/10, countdown frozen at 0s
12–18 min ✻ API error · Retrying in 0s · attempt 2/10
18–24 min ✻ Request timed out. · Retrying in 0s · attempt 3/10
24 min on ✻ Request timed out. · Retrying in 0s · attempt 4/10

"Retrying in 0s" sitting there for 6 minutes is what makes it look frozen. The countdown only covers the backoff between attempts (under a second here); then the new request goes out and waits another 6 minutes with no further hint. Case ② (headers then silence) looks exactly the same.

4. Pings only, and mid-response stalls

Stall Aborted at Reason in the debug log Afterwards
③ Pings every 15 s, default 750 s Streaming idle warning: no chunks received for 150s at 600 s, then Streaming idle timeout: no chunks received for 300s, aborting stream at 750 s Re-sent; 7 requests by 50 min
③ Pings every 5 s, byte timeout 10 s 600 s same Re-sent; 3 requests by 25 min
④ Mid-response, default 300 s [byte-watchdog] firing: idle=300000ms → Streaming idle timeout (byte-level) Re-sent about every 5 min, logged as Request timed out; at 50 min on retry 8
④ Mid-response, byte timeout 10 s 10 s same (10000ms) Only one extra re-send, then API Error: stream idle: no bytes for 10000ms; 20.3–20.7 s total

Two takeaways:

  • Keep-alive pings don't save you: they fool the byte-level watchdog, but the event-level watchdog only counts real response events, and the abort comes even later than with silence (750 s vs 360 s).
  • On a mid-response stall, the partial answer is discarded: -p stdout contains only the final error.

(The four default cases were cut off by the experiment's 50-minute watchdog; "Afterwards" is the progress at that point, not the final result.)

5. Buffering gateways: replies over 6 minutes never arrive

Some gateways don't stream; they wait for the upstream to finish generating and then return everything at once. To Claude Code that is case ①, lasting until generation completes. The stub returns a normal reply after N seconds, with CLAUDE_CODE_MAX_RETRIES=1:

Buffer time Result Requests Wall time
300 s ✅ received, OK, exit 0 1 300.33 s
400 s ❌ cut at 361 s, re-sent, cut again at 361 s, Request timed out 2 721.29 s

If a single reply takes more than 6 minutes to generate, every attempt is cut at the 6-minute mark; no number of retries helps. Short questions fine, long tasks always stuck — that is the signature.

What to do

1. Check whether the backend answers at all. Send one streaming request with the same base URL and key Claude Code uses (use a model your backend supports):

curl -sN -w '\nfirst byte %{time_starttransfer}s, total %{time_total}s\n' \
  "$ANTHROPIC_BASE_URL/v1/messages" \
  -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-haiku-4-5","max_tokens":16,"stream":true,"messages":[{"role":"user","content":"hi"}]}'

This uses a tiny amount of quota. Normally event: lines start within seconds. No output for a long time means the backend or network is the problem, not Claude Code. Everything arriving at once at the end means your gateway buffers (section 5).

2. Press Esc when stuck. Your message stays in the conversation; just send again. Otherwise the worst case is about 69 minutes (4147.78 s).

3. Fail fast instead of hanging for an hour:

export API_TIMEOUT_MS=60000          # at most 60 s per request (only lowering works; above 6 min has no effect)
export CLAUDE_CODE_MAX_RETRIES=2     # at most 2 retries

Too small an API_TIMEOUT_MS will also cut off legitimate long replies — it limits the whole request, not just the first byte. 60 s is for troubleshooting; for daily use size it to your longest task.

4. Buffering gateway: switch to one that streams, or split the task so each reply finishes within 6 minutes. Raising API_TIMEOUT_MS doesn't help — section 1 shows raising has no effect.

5. Stuck on a command, not a spinner: that's the Bash tool waiting (2-minute default timeout), unrelated to network waits — see Command timed out after 2m 0s.

Every request in this test went to the local stub; real API cost was zero.

Pitfalls along the way

  • Nearly reported the test rig's timeout as the result. 360 s is very close to Node's default requestTimeout (300 s) plus its check interval. The number couldn't be published until the stub's own timeouts were disabled and it was still 360 s.
  • Mental arithmetic turned 6 minutes into 10. Reading 15:04:43 → 15:10:43 in the debug log, I first took it as 10 minutes, matching the docs. All time differences are now computed by script.
  • Wrong "connection closed" timestamps for the hang case. The first version logged req.on('close'), which Node fires once the request body is read, so every value was 0 s. Final numbers use request arrival times and Claude Code's own debug log.

Not verified: the direct Anthropic API without ANTHROPIC_BASE_URL (per the docs the 180 s first-byte deadline applies there; not tested); OAuth subscription logins; real HTTPS gateways (the stub is local HTTP); the source of the 360 s ceiling in code; why the ping case's re-send interval drops from 750 s to about 300 s after 1500 s; why mid-response stalls retry differently with a 10 s byte timeout vs the default (the logged error types differ — stream_idle_timeout vs Request timed out — not confirmed in code).

Get field notes like this every Saturday

Subscribe to Dev Breakfast: daily AI coding picks at 8:00, plus a Saturday roundup of this week's hands-on tests with Claude Code / Codex / local models. Written in Chinese.

Related Articles

Claude Code "Prompt is too long" vs "maximum context length": Why Auto-Compact Works for One and Not the Other (Tested)

Claude Code decides a request was too long by matching the error text. If the backend says prompt is too long or input is too long for requested model, it auto-compacts the conversation and retries — invisible to you. DeepSeek and OpenAI-style gateways say This model's maximum context length is …, which it doesn't recognize: you get API Error: 400 on every turn. Manual /compact works; the better fix is to declare the real window with CLAUDE_CODE_MAX_CONTEXT_TOKENS (non-claude- model IDs) or CLAUDE_CODE_AUTO_COMPACT_WINDOW (claude- IDs) so it compacts before hitting the limit. Tested on Claude Code 2.1.285 against a local stub.

claude-codedeepseek+6
pitfallsOct 5, 20267 min
9

Claude Code 429 "Request rejected (429)": How Long It Retries, and Why retry-after Over 60 Seconds Fails Instantly (Tested)

On a 429, Claude Code retries up to 10 times over about 3 minutes. If retry-after is 60 seconds or less it waits the full time (10 s and 60 s both recovered in testing); at 61 seconds or more it does not retry at all and fails immediately with API Error: Request rejected (429). Gateway messages such as new-api's are shown verbatim; an empty body shows status code (no body). In interactive mode every prompt sends 2 requests, so rate-limit usage doubles. All tested against a local stub on Claude Code 2.1.285.

claude-codetroubleshooting+4
pitfallsOct 4, 20268 min
12

Clash Verge TUN Mode Breaks All Internet Access: Hysteria2 Traffic Loops Back Into the TUN, and Tailscale Hijacks DNS

Clash Verge Rev 2.5.6 on macOS works fine in system-proxy mode, but as soon as TUN (virtual network adapter) mode is on, not even Baidu loads. Debugging through the mihomo core's API turned up two independent root causes stacked together. First, Hysteria2's outbound UDP isn't bound to the physical interface, so the TUN routes pull it back in and it loops. Second, Tailscale MagicDNS (100.100.100.100) has taken over system DNS, so queries go out through Tailscale's utun where Clash's dns-hijack can't see them, and come back poisoned. The fix is two Merge overrides: route-exclude-address to keep node IPs out of the TUN, and sniffer to recover domains from SNI. Every step comes with the commands and real output.

troubleshootingtailscale+7
pitfallsOct 4, 20266 min
32
Claude Code MCP Shows Connected but 0 Tools: "Invalid result for tools/list" (ttlMs / cacheScope), Reproduced and Fixed

Claude Code MCP Shows Connected but 0 Tools: "Invalid result for tools/list" (ttlMs / cacheScope), Reproduced and Fixed

An MCP server shows as connected but exposes 0 tools, and the log says Invalid result for tools/list with ttlMs and cacheScope failing validation. Reproduced with a stub server on Claude Code 2.1.280, 2.1.285 and 2.1.288: the cause is not extra fields being rejected. The server negotiated MCP 2026-07-28 and then left out fields that revision requires (resultType, ttlMs, cacheScope); adding unknown fields works fine. Whether stdio uses the new protocol is decided by a remote flag that is off by default, so the same version breaks for some users and not others, and 2.1.280 fails the same way once negotiation is on. MCP_PROTOCOL_NEGOTIATION=legacy (also works in settings.json env) restores the tools; servers fix it by adding the three fields.

mcpclaude-code+4
pitfallsOct 3, 20268 min
21

Published by Magic Tools