Magic Tools
Pitfall NotesBy CooconOctober 11, 202610 views12 min read

Claude Code Auto Mode with DeepSeek, GLM or Kimi: What Each "temporarily unavailable" Reason Means and When Retrying Is Pointless (Fault-Injection Test)

Short answers first

  • Why does the error name glm-5.3, deepseek-v4-pro or k3? Before every action with side effects (Bash, file writes, and so on), auto mode sends an extra safety-check request using your session model to the backend you configured. The name in the error is that classifier, which is simply your own session model.
  • Look at the parentheses first. (server error), (rate-limited), (overloaded) and (connection failed) mean the backend had a temporary problem; Claude Code already tried 5 times. (timed out) means the check got no response within 60 seconds.
  • No parentheses is the one to watch. The backend rejected the classifier request outright (a 4xx). Claude Code tries only once, and retrying will give the same result every time. Check your backend configuration.
  • Auto mode could not evaluate this action: the backend answered, but the model did not return a verdict in the expected format. After three of these in a row, auto mode falls back to manual approval.
  • Want a separate model for the classifier? We could not find a way. CLAUDE_CODE_AUTO_MODE_MODEL has no effect in 2.1.285. The only way to change the classifier is to change the session model.

Background

Our guide Claude Code "temporarily unavailable, so auto mode cannot determine the safety of bash": how to fix it is the most-searched page on this site over the last 28 days. Many of the search queries reaching it name a model that is not Claude: glm-5.3, glm5.2, glm-5.3-flash, claude-glm-5.3, some with a cc switch prefix. In other words, a lot of people hit this error after pointing Claude Code at Zhipu, Kimi, DeepSeek or another third-party backend through tools like cc-switch or claude-code-router.

That guide covered this case in one line: "switch to a healthy backend." What these users actually want to know is:

  1. Can a third-party model run auto mode at all? What does the classifier request look like, and what does it cost?
  2. Which failure does each reason in the parentheses map to? Which ones go away if you wait, and which ones never will?
  3. Can you give the classifier its own faster or cheaper model?

How the check works

Auto mode replaces manual approval with one model call: before an action with side effects runs, a model judges whether it is safe (details in our teardown of the auto mode classifier).

This test showed that 2.1.285 actually has two paths for that check. The debug log says so directly. On a third-party backend, these three lines appear at the first check:

[server-classifier] the platform gave no classification for this request (server_no_result); treating that as this deployment's answer: auto mode classifies locally for the rest of this session and stops sending the classifier context
[server-classifier] Bash: no server verdict for this call (server_no_result); the local classifier decides it
[server-classifier] auto mode fell back to billed classifier requests; no dialog surface to warn on, continuing in auto mode

Claude Code first looks for a verdict from the platform itself. A third-party backend never provides one (server_no_result), so for the rest of the session it uses the local classifier, which the log calls "billed classifier requests": every action that needs a check costs one separate request to your backend. Every instance of this error that third-party users see comes from that local path.

In the 2.1.285 source, the reason in parentheses is built from the HTTP status and error type:

if(e===429)return" (rate-limited)";
if(e===529)return" (overloaded)";
if(e!==void 0&&e>=500&&e<600)return" (server error)";
if(n==="wall_clock_timeout"||n==="connection_timeout"||...)return" (timed out)";
if(n==="connection_error")return" (connection failed)";

The source only tells you what should happen. How many retries, how long it waits, and what the model does after the error all have to be measured.

Test setup

A fault-injection proxy between Claude Code and a real backend. ANTHROPIC_BASE_URL points to a local proxy. Main-loop requests are forwarded unchanged to the real backend. Classifier requests (system prompt starting with You are a security monitor for autonomous AI coding agents) are answered according to the test case: 500, 429, hang, connection reset, and so on. The main loop runs on a real model, and only the classifier sees the faults.

Real backends: DeepSeek and MiMo. Both expose an Anthropic-compatible endpoint (/anthropic), we have keys for both, and running them is cheap. GLM was not tested because we do not have a Zhipu key. The code above does not distinguish between backends, so the error text is built the same way for all of them; what differs between backends is which failures happen most often. Anything this article says about GLM is inferred from that, not measured.

Isolation: env -i with a fresh, empty CLAUDE_CONFIG_DIR for every run. Claude Code only holds a fake key and the proxy injects the real one. HTTPS_PROXY also points to the proxy, which blocks every outbound connection (in practice, api.anthropic.com). The working directory lives under /tmp.

What we ruled out:

  • A stub with no real backend: it cannot give real classifier latency, token usage or cache behavior, which is exactly what third-party users care about.
  • Running inside the repo: we fell into this on the first run; see Pitfalls below.

Every run asks for the same Bash command, one that needs a check: mkdir -p /tmp/automode-lab-out && python3 -c "print(6*7)" > .... Our first choice, echo hello > out.txt, writes to the working directory, and the log showed Skipping auto mode classifier for Bash: would be allowed in acceptEdits mode, so the classifier never ran.

Results

1. Third-party models can run auto mode; here is the classifier request

Third-party models are allowed into auto mode. The gate log shows modelSupported=true … canEnterAuto=true for both DeepSeek and MiMo. With no faults injected, the command is judged safe and runs. This is the classifier request the proxy captured:

Field Measured in 2.1.285
model the session model (deepseek-v4-pro / mimo-v2.5)
stream no
max_tokens 2112
stop_sequences ["</block>"]
thinking not sent
system 3 blocks, 132,470 characters total; the large middle block is 132,144 characters and carries cache_control: ephemeral
expected output <block>no</block> or <block>yes</block>

DeepSeek returned <block>no with stop_reason: stop_sequence. MiMo ignored the stop sequence and returned the full <block>no</block> with end_turn. Both added a thinking block even though the request did not ask for one, and both verdicts parsed correctly.

Compared with our August capture on 2.1.241 (max_tokens: 64, stop at </severity>, 116,879-character prompt), the format has changed to a binary <block> verdict.

2. What one check costs in tokens and time

We ran three Bash commands in one session (python3 -c "print(1+1)" and so on) and counted main-loop and classifier input tokens separately:

DeepSeek v4-pro, 3 actions Main loop Classifier
Total input tokens (incl. cache hits) 71,407 89,043
Of which not cached 18,031 38,099
Time per check — 4.45 s / 1.96 s / 1.75 s

Check by check: the first had 29,664 tokens, none cached; the second 8,305 uncached and 21,376 cached; the third 130 uncached and 29,568 cached. The cache warms up within a session. But in our DeepSeek runs, the first check of every new session paid in full: in all 6 new sessions where we could read usage, the first check showed cache_read_input_tokens of 0. The system prompt was byte-for-byte identical across sessions (we diffed it), so the cause is on the backend side and cannot be seen from the client.

MiMo did hit its cache across sessions (from the second session on, the first check hit 28,672 tokens), but it was much slower: 11.5 s, 9.3 s, 4.7 s, 11.9 s, 4.5 s and 11.5 s over six checks. DeepSeek's 11 successful checks took between 1.75 and 4.45 s.

Bottom line: with auto mode on, classifier input is about as large as the main loop's, or larger. That is because the classifier's system prompt (about 30K tokens) is longer than the main loop's system prompt plus tool definitions (about 17.6K tokens). If you pay per token on a third-party API, budget for that.

3. Fault injection: which failure produces which message

The main loop is forwarded normally and only the classifier request fails. The table shows the text the model received as the tool result (the shared ending "Wait a moment and then try this action again. If it keeps failing, continue with other tasks…" is trimmed), and how many classifier requests Claude Code sent per action:

Injected fault Error text (key part) Classifier requests per action
HTTP 500 deepseek-v4-pro is temporarily unavailable (server error), so auto mode cannot determine the safety of Bash right now. 5 (gaps of about 0.5 / 0.9 / 1.7 / 3.6 s)
HTTP 429 … temporarily unavailable (rate-limited), so auto mode … 5
HTTP 529 … temporarily unavailable (overloaded), so auto mode … 5
Connection reset … temporarily unavailable (connection failed), so auto mode … 5
No response … temporarily unavailable (timed out), so auto mode … 1, gives up after 60 s
HTTP 400 deepseek-v4-pro is temporarily unavailable, so auto mode … (no parentheses) 1, no retry
HTTP 404 same, no parentheses 1, no retry
200 with a reply in prose ("Sure, this command is safe to run.", sent in Chinese) Auto mode could not evaluate this action and is blocking it for safety — run with --debug for details. This is not a judgment that the action is unsafe. 10 (parse failures are retried too)
200 with stop_reason: max_tokens and no verdict same 10

Three things you can use right away:

  • No parentheses means the backend rejected the request (4xx). Claude Code treats it as unrecoverable and tries once. Our older guide grouped "no parentheses" with the other reasons and said they are all handled the same way; this test shows that was wrong, and that guide has been corrected. On a third-party backend, the usual cause is something in the classifier request the backend does not accept: a model name missing from its mapping, no support for non-streaming requests or stop_sequences, or about 30K tokens of input exceeding the model's context. Note that the main loop can work fine while the classifier fails: both use the same model name, but the requests look completely different (see the table in section 1).
  • (overloaded) was missing from our old variant list. It comes from a 529 and is now added.
  • could not evaluate is not a network problem. The backend answered; the model just did not output <block>yes/no</block>. A thinking model that spends its whole 2112 max_tokens budget on reasoning, or a model that ignores the format instruction, ends up here.

4. The 60-second timeout

When the classifier request hung, the proxy saw the client disconnect at 60,002 ms and 60,000 ms (two runs), after a single attempt, followed by (timed out). In another case the classifier response was delayed by 45 seconds; the check passed and the command ran.

So if your backend takes more than a minute to process 30K tokens of input (say, queueing at peak hours), every action that needs a check will hang for a full 60 seconds before the error appears.

5. A retry-after header makes every action wait silently

With a 429 carrying retry-after: 30, the five classifier attempts were spaced exactly 30 seconds apart (30,009 / 60,019 / 90,029 / 120,039 ms). Each action waited 2 minutes before reporting (rate-limited), and in -p mode nothing was printed in the meantime. If your plan limits concurrency or request rate, remember that auto mode adds one extra request per action.

6. Three "could not evaluate" results in a row bring back manual approval

There is another important difference between "temporarily unavailable" and "could not evaluate": the latter counts as a block. From the debug log:

[WARN] Classifier denial limit exceeded, falling back to prompting: 3 consecutive actions were blocked. Please review the transcript before continuing.

After that, the same command came back as This Bash command contains multiple operations. The following parts require approval: mkdir -p /tmp/automode-lab-out, python3 -c "print(6*7)". In an interactive session you would get an approval prompt. In headless -p mode nobody can approve it, so it is denied.

7. "temporarily unavailable" has no circuit breaker, so the model keeps retrying

"temporarily unavailable" does not count toward that block limit, and its message tells the model to "Wait a moment and then try this action again". DeepSeek v4-pro retried immediately:

  • Persistent 500: 38 main-loop turns and 185 classifier requests in 400 seconds, until our watchdog killed the run.
  • Persistent 400 (no retry, no backoff): 44 main-loop turns in 100 seconds.
  • Persistent 404: the model gave up on its own after 6 attempts, replying "I've retried the command 6 times…".

When to give up is the model's own call, and a different model may behave differently. What is certain is that Claude Code itself sets no limit in this situation. While the classifier stays broken, the main loop keeps spending tokens. If you see the same error scrolling past again and again in an interactive session, press Esc.

8. A separate classifier model: CLAUDE_CODE_AUTO_MODE_MODEL does nothing

The binary contains an environment variable called CLAUDE_CODE_AUTO_MODE_MODEL, which sounds like a way to pick the classifier model. We set it to deepseek-flash while keeping the session model at deepseek-v4-pro. The classifier request still used deepseek-v4-pro, and the log still read classifier_request_started … model=deepseek-v4-pro. Reading the minified source, the classifier model comes from, in order: server-delivered config (which a third-party backend never receives) → a built-in default mapping → the session model. For third-party users, the classifier is the session model; the only way to change it is /model.

What to do, by message

What you see What it means What to do
(server error) / (overloaded) / (connection failed) Temporary backend failure; already retried 5 times Wait and try again; if it persists, switch backend or session model
(rate-limited) The backend is throttling you Auto mode doubles your request count; check your plan's concurrency or rate limit, and consider leaving auto mode at peak times
(timed out) No classifier response within 60 s The backend is too slow (30K-token input plus queueing); switch to a faster model or backend
No parentheses The backend rejected the classifier request (4xx); retrying won't help Check the model-name mapping, whether the backend supports non-streaming and stop_sequences, and whether the context fits 30K+ tokens; look at the classifier_request_finished lines in claude --debug-file /tmp/cc.log
Auto mode could not evaluate this action The backend answered, but the model gave no verdict Switch to a session model that reliably follows the output format; after 3 in a row you are back to manual approval
The same error scrolling past again and again The model is retrying and Claude Code sets no limit Press Esc, use the rows above, or switch to manual / acceptEdits mode

The four steps from our older guide (wait and retry, do read-only work first, switch session model, leave auto mode) still apply to third-party users. What this test adds is the first column: read the parentheses before deciding whether to retry.

Pitfalls

  • A working directory inside the repo pulled in project settings. On the first run the working directory was tmp/…/runs/r1/work, and the debug log showed Applying permission update: Adding 22 allow rule(s). Claude Code had walked up the tree to the repo's .claude/settings.local.json, and it also dropped Bash(npm run *) as a dangerous rule (Ignoring dangerous permission … (bypasses classifier)). The file itself was not modified (its mtime did not change), but the test environment was no longer clean. All later runs used /tmp.
  • The test command was auto-allowed by the acceptEdits rule. echo hello > out.txt in the working directory never reaches the classifier; a command that writes to /tmp and runs python does.
  • Passing accept-encoding through the proxy hid the responses. On the first forward, the upstream replied with gzip and the proxy could not read the usage. Dropping that request header gave us plain text.
  • source .env executes cron expressions. The project .env has unquoted values like X_CRON=*/5 * * * *; when sourced, the shell expands * into file names and tries to run them as a command. It did no harm here; we switched to reading only the keys we need with grep.

Get field notes like this every Saturday

Subscribe to Dev Breakfast: daily AI coding picks at 8:00, plus a Saturday roundup of this week's hands-on tests with Claude Code / Codex / local models. Written in Chinese.

Related Articles

Claude Code "Prompt is too long" vs "maximum context length": Why Auto-Compact Works for One and Not the Other (Tested)

Claude Code decides a request was too long by matching the error text. If the backend says prompt is too long or input is too long for requested model, it auto-compacts the conversation and retries — invisible to you. DeepSeek and OpenAI-style gateways say This model's maximum context length is …, which it doesn't recognize: you get API Error: 400 on every turn. Manual /compact works; the better fix is to declare the real window with CLAUDE_CODE_MAX_CONTEXT_TOKENS (non-claude- model IDs) or CLAUDE_CODE_AUTO_COMPACT_WINDOW (claude- IDs) so it compacts before hitting the limit. Tested on Claude Code 2.1.285 against a local stub.

claude-codedeepseek+6
pitfallsOct 5, 20267 min
154

Claude Code Stuck on a Spinner With No Response: It Waits 6 Minutes per Attempt, and 10 Retries Can Hang It for Over an Hour (Tested)

When Claude Code spins without output or sits at Retrying in 0s, the backend is usually not sending anything. Tested on 2.1.285 (API key + ANTHROPIC_BASE_URL): with no response headers it waits 6 minutes (360 s) per attempt before timing out — setting API_TIMEOUT_MS to 600000 or 900000 doesn't change that; only lower values work. With the default 10 retries it can hang for over an hour. A gateway that buffers the reply until generation finishes will never deliver a reply that takes over 6 minutes. Press Esc to interrupt. All tested against a local stub.

claude-codetroubleshooting+5
pitfallsOct 5, 20269 min
142

Claude Code 429 "Request rejected (429)": How Long It Retries, and Why retry-after Over 60 Seconds Fails Instantly (Tested)

On a 429, Claude Code retries up to 10 times over about 3 minutes. If retry-after is 60 seconds or less it waits the full time (10 s and 60 s both recovered in testing); at 61 seconds or more it does not retry at all and fails immediately with API Error: Request rejected (429). Gateway messages such as new-api's are shown verbatim; an empty body shows status code (no body). In interactive mode every prompt sends 2 requests, so rate-limit usage doubles. All tested against a local stub on Claude Code 2.1.285.

claude-codetroubleshooting+4
pitfallsOct 4, 20268 min
132

Clash Verge TUN Mode Breaks All Internet Access: Hysteria2 Traffic Loops Back Into the TUN, and Tailscale Hijacks DNS

Clash Verge Rev 2.5.6 on macOS works fine in system-proxy mode, but as soon as TUN (virtual network adapter) mode is on, not even Baidu loads. Debugging through the mihomo core's API turned up two independent root causes stacked together. First, Hysteria2's outbound UDP isn't bound to the physical interface, so the TUN routes pull it back in and it loops. Second, Tailscale MagicDNS (100.100.100.100) has taken over system DNS, so queries go out through Tailscale's utun where Clash's dns-hijack can't see them, and come back poisoned. The fix is two Merge overrides: route-exclude-address to keep node IPs out of the TUN, and sniffer to recover domains from SNI. Every step comes with the commands and real output.

troubleshootingtailscale+7
pitfallsOct 4, 20266 min
128

Published by Magic Tools