Claude Code Auto Mode with DeepSeek, GLM or Kimi: What Each "temporarily unavailable" Reason Means and When Retrying Is Pointless (Fault-Injection Test)
Short answers first
- Why does the error name
glm-5.3,deepseek-v4-proork3? Before every action with side effects (Bash, file writes, and so on), auto mode sends an extra safety-check request using your session model to the backend you configured. The name in the error is that classifier, which is simply your own session model.- Look at the parentheses first.
(server error),(rate-limited),(overloaded)and(connection failed)mean the backend had a temporary problem; Claude Code already tried 5 times.(timed out)means the check got no response within 60 seconds.- No parentheses is the one to watch. The backend rejected the classifier request outright (a 4xx). Claude Code tries only once, and retrying will give the same result every time. Check your backend configuration.
Auto mode could not evaluate this action: the backend answered, but the model did not return a verdict in the expected format. After three of these in a row, auto mode falls back to manual approval.- Want a separate model for the classifier? We could not find a way.
CLAUDE_CODE_AUTO_MODE_MODELhas no effect in 2.1.285. The only way to change the classifier is to change the session model.
Background
Our guide Claude Code "temporarily unavailable, so auto mode cannot determine the safety of bash": how to fix it is the most-searched page on this site over the last 28 days. Many of the search queries reaching it name a model that is not Claude: glm-5.3, glm5.2, glm-5.3-flash, claude-glm-5.3, some with a cc switch prefix. In other words, a lot of people hit this error after pointing Claude Code at Zhipu, Kimi, DeepSeek or another third-party backend through tools like cc-switch or claude-code-router.
That guide covered this case in one line: "switch to a healthy backend." What these users actually want to know is:
- Can a third-party model run auto mode at all? What does the classifier request look like, and what does it cost?
- Which failure does each reason in the parentheses map to? Which ones go away if you wait, and which ones never will?
- Can you give the classifier its own faster or cheaper model?
How the check works
Auto mode replaces manual approval with one model call: before an action with side effects runs, a model judges whether it is safe (details in our teardown of the auto mode classifier).
This test showed that 2.1.285 actually has two paths for that check. The debug log says so directly. On a third-party backend, these three lines appear at the first check:
[server-classifier] the platform gave no classification for this request (server_no_result); treating that as this deployment's answer: auto mode classifies locally for the rest of this session and stops sending the classifier context
[server-classifier] Bash: no server verdict for this call (server_no_result); the local classifier decides it
[server-classifier] auto mode fell back to billed classifier requests; no dialog surface to warn on, continuing in auto mode
Claude Code first looks for a verdict from the platform itself. A third-party backend never provides one (server_no_result), so for the rest of the session it uses the local classifier, which the log calls "billed classifier requests": every action that needs a check costs one separate request to your backend. Every instance of this error that third-party users see comes from that local path.
In the 2.1.285 source, the reason in parentheses is built from the HTTP status and error type:
if(e===429)return" (rate-limited)";
if(e===529)return" (overloaded)";
if(e!==void 0&&e>=500&&e<600)return" (server error)";
if(n==="wall_clock_timeout"||n==="connection_timeout"||...)return" (timed out)";
if(n==="connection_error")return" (connection failed)";
The source only tells you what should happen. How many retries, how long it waits, and what the model does after the error all have to be measured.
Test setup
A fault-injection proxy between Claude Code and a real backend. ANTHROPIC_BASE_URL points to a local proxy. Main-loop requests are forwarded unchanged to the real backend. Classifier requests (system prompt starting with You are a security monitor for autonomous AI coding agents) are answered according to the test case: 500, 429, hang, connection reset, and so on. The main loop runs on a real model, and only the classifier sees the faults.
Real backends: DeepSeek and MiMo. Both expose an Anthropic-compatible endpoint (/anthropic), we have keys for both, and running them is cheap. GLM was not tested because we do not have a Zhipu key. The code above does not distinguish between backends, so the error text is built the same way for all of them; what differs between backends is which failures happen most often. Anything this article says about GLM is inferred from that, not measured.
Isolation: env -i with a fresh, empty CLAUDE_CONFIG_DIR for every run. Claude Code only holds a fake key and the proxy injects the real one. HTTPS_PROXY also points to the proxy, which blocks every outbound connection (in practice, api.anthropic.com). The working directory lives under /tmp.
What we ruled out:
- A stub with no real backend: it cannot give real classifier latency, token usage or cache behavior, which is exactly what third-party users care about.
- Running inside the repo: we fell into this on the first run; see Pitfalls below.
Every run asks for the same Bash command, one that needs a check: mkdir -p /tmp/automode-lab-out && python3 -c "print(6*7)" > .... Our first choice, echo hello > out.txt, writes to the working directory, and the log showed Skipping auto mode classifier for Bash: would be allowed in acceptEdits mode, so the classifier never ran.
Results
1. Third-party models can run auto mode; here is the classifier request
Third-party models are allowed into auto mode. The gate log shows modelSupported=true … canEnterAuto=true for both DeepSeek and MiMo. With no faults injected, the command is judged safe and runs. This is the classifier request the proxy captured:
| Field | Measured in 2.1.285 |
|---|---|
model |
the session model (deepseek-v4-pro / mimo-v2.5) |
stream |
no |
max_tokens |
2112 |
stop_sequences |
["</block>"] |
thinking |
not sent |
| system | 3 blocks, 132,470 characters total; the large middle block is 132,144 characters and carries cache_control: ephemeral |
| expected output | <block>no</block> or <block>yes</block> |
DeepSeek returned <block>no with stop_reason: stop_sequence. MiMo ignored the stop sequence and returned the full <block>no</block> with end_turn. Both added a thinking block even though the request did not ask for one, and both verdicts parsed correctly.
Compared with our August capture on 2.1.241 (max_tokens: 64, stop at </severity>, 116,879-character prompt), the format has changed to a binary <block> verdict.
2. What one check costs in tokens and time
We ran three Bash commands in one session (python3 -c "print(1+1)" and so on) and counted main-loop and classifier input tokens separately:
| DeepSeek v4-pro, 3 actions | Main loop | Classifier |
|---|---|---|
| Total input tokens (incl. cache hits) | 71,407 | 89,043 |
| Of which not cached | 18,031 | 38,099 |
| Time per check | — | 4.45 s / 1.96 s / 1.75 s |
Check by check: the first had 29,664 tokens, none cached; the second 8,305 uncached and 21,376 cached; the third 130 uncached and 29,568 cached. The cache warms up within a session. But in our DeepSeek runs, the first check of every new session paid in full: in all 6 new sessions where we could read usage, the first check showed cache_read_input_tokens of 0. The system prompt was byte-for-byte identical across sessions (we diffed it), so the cause is on the backend side and cannot be seen from the client.
MiMo did hit its cache across sessions (from the second session on, the first check hit 28,672 tokens), but it was much slower: 11.5 s, 9.3 s, 4.7 s, 11.9 s, 4.5 s and 11.5 s over six checks. DeepSeek's 11 successful checks took between 1.75 and 4.45 s.
Bottom line: with auto mode on, classifier input is about as large as the main loop's, or larger. That is because the classifier's system prompt (about 30K tokens) is longer than the main loop's system prompt plus tool definitions (about 17.6K tokens). If you pay per token on a third-party API, budget for that.
3. Fault injection: which failure produces which message
The main loop is forwarded normally and only the classifier request fails. The table shows the text the model received as the tool result (the shared ending "Wait a moment and then try this action again. If it keeps failing, continue with other tasks…" is trimmed), and how many classifier requests Claude Code sent per action:
| Injected fault | Error text (key part) | Classifier requests per action |
|---|---|---|
| HTTP 500 | deepseek-v4-pro is temporarily unavailable (server error), so auto mode cannot determine the safety of Bash right now. |
5 (gaps of about 0.5 / 0.9 / 1.7 / 3.6 s) |
| HTTP 429 | … temporarily unavailable (rate-limited), so auto mode … |
5 |
| HTTP 529 | … temporarily unavailable (overloaded), so auto mode … |
5 |
| Connection reset | … temporarily unavailable (connection failed), so auto mode … |
5 |
| No response | … temporarily unavailable (timed out), so auto mode … |
1, gives up after 60 s |
| HTTP 400 | deepseek-v4-pro is temporarily unavailable, so auto mode … (no parentheses) |
1, no retry |
| HTTP 404 | same, no parentheses | 1, no retry |
| 200 with a reply in prose ("Sure, this command is safe to run.", sent in Chinese) | Auto mode could not evaluate this action and is blocking it for safety — run with --debug for details. This is not a judgment that the action is unsafe. |
10 (parse failures are retried too) |
200 with stop_reason: max_tokens and no verdict |
same | 10 |
Three things you can use right away:
- No parentheses means the backend rejected the request (4xx). Claude Code treats it as unrecoverable and tries once. Our older guide grouped "no parentheses" with the other reasons and said they are all handled the same way; this test shows that was wrong, and that guide has been corrected. On a third-party backend, the usual cause is something in the classifier request the backend does not accept: a model name missing from its mapping, no support for non-streaming requests or
stop_sequences, or about 30K tokens of input exceeding the model's context. Note that the main loop can work fine while the classifier fails: both use the same model name, but the requests look completely different (see the table in section 1). (overloaded)was missing from our old variant list. It comes from a 529 and is now added.could not evaluateis not a network problem. The backend answered; the model just did not output<block>yes/no</block>. A thinking model that spends its whole 2112max_tokensbudget on reasoning, or a model that ignores the format instruction, ends up here.
4. The 60-second timeout
When the classifier request hung, the proxy saw the client disconnect at 60,002 ms and 60,000 ms (two runs), after a single attempt, followed by (timed out). In another case the classifier response was delayed by 45 seconds; the check passed and the command ran.
So if your backend takes more than a minute to process 30K tokens of input (say, queueing at peak hours), every action that needs a check will hang for a full 60 seconds before the error appears.
5. A retry-after header makes every action wait silently
With a 429 carrying retry-after: 30, the five classifier attempts were spaced exactly 30 seconds apart (30,009 / 60,019 / 90,029 / 120,039 ms). Each action waited 2 minutes before reporting (rate-limited), and in -p mode nothing was printed in the meantime. If your plan limits concurrency or request rate, remember that auto mode adds one extra request per action.
6. Three "could not evaluate" results in a row bring back manual approval
There is another important difference between "temporarily unavailable" and "could not evaluate": the latter counts as a block. From the debug log:
[WARN] Classifier denial limit exceeded, falling back to prompting: 3 consecutive actions were blocked. Please review the transcript before continuing.
After that, the same command came back as This Bash command contains multiple operations. The following parts require approval: mkdir -p /tmp/automode-lab-out, python3 -c "print(6*7)". In an interactive session you would get an approval prompt. In headless -p mode nobody can approve it, so it is denied.
7. "temporarily unavailable" has no circuit breaker, so the model keeps retrying
"temporarily unavailable" does not count toward that block limit, and its message tells the model to "Wait a moment and then try this action again". DeepSeek v4-pro retried immediately:
- Persistent 500: 38 main-loop turns and 185 classifier requests in 400 seconds, until our watchdog killed the run.
- Persistent 400 (no retry, no backoff): 44 main-loop turns in 100 seconds.
- Persistent 404: the model gave up on its own after 6 attempts, replying "I've retried the command 6 times…".
When to give up is the model's own call, and a different model may behave differently. What is certain is that Claude Code itself sets no limit in this situation. While the classifier stays broken, the main loop keeps spending tokens. If you see the same error scrolling past again and again in an interactive session, press Esc.
8. A separate classifier model: CLAUDE_CODE_AUTO_MODE_MODEL does nothing
The binary contains an environment variable called CLAUDE_CODE_AUTO_MODE_MODEL, which sounds like a way to pick the classifier model. We set it to deepseek-flash while keeping the session model at deepseek-v4-pro. The classifier request still used deepseek-v4-pro, and the log still read classifier_request_started … model=deepseek-v4-pro. Reading the minified source, the classifier model comes from, in order: server-delivered config (which a third-party backend never receives) → a built-in default mapping → the session model. For third-party users, the classifier is the session model; the only way to change it is /model.
What to do, by message
| What you see | What it means | What to do |
|---|---|---|
(server error) / (overloaded) / (connection failed) |
Temporary backend failure; already retried 5 times | Wait and try again; if it persists, switch backend or session model |
(rate-limited) |
The backend is throttling you | Auto mode doubles your request count; check your plan's concurrency or rate limit, and consider leaving auto mode at peak times |
(timed out) |
No classifier response within 60 s | The backend is too slow (30K-token input plus queueing); switch to a faster model or backend |
| No parentheses | The backend rejected the classifier request (4xx); retrying won't help | Check the model-name mapping, whether the backend supports non-streaming and stop_sequences, and whether the context fits 30K+ tokens; look at the classifier_request_finished lines in claude --debug-file /tmp/cc.log |
Auto mode could not evaluate this action |
The backend answered, but the model gave no verdict | Switch to a session model that reliably follows the output format; after 3 in a row you are back to manual approval |
| The same error scrolling past again and again | The model is retrying and Claude Code sets no limit | Press Esc, use the rows above, or switch to manual / acceptEdits mode |
The four steps from our older guide (wait and retry, do read-only work first, switch session model, leave auto mode) still apply to third-party users. What this test adds is the first column: read the parentheses before deciding whether to retry.
Pitfalls
- A working directory inside the repo pulled in project settings. On the first run the working directory was
tmp/…/runs/r1/work, and the debug log showedApplying permission update: Adding 22 allow rule(s). Claude Code had walked up the tree to the repo's.claude/settings.local.json, and it also droppedBash(npm run *)as a dangerous rule (Ignoring dangerous permission … (bypasses classifier)). The file itself was not modified (its mtime did not change), but the test environment was no longer clean. All later runs used/tmp. - The test command was auto-allowed by the acceptEdits rule.
echo hello > out.txtin the working directory never reaches the classifier; a command that writes to/tmpand runs python does. - Passing accept-encoding through the proxy hid the responses. On the first forward, the upstream replied with gzip and the proxy could not read the usage. Dropping that request header gave us plain text.
source .envexecutes cron expressions. The project.envhas unquoted values likeX_CRON=*/5 * * * *; when sourced, the shell expands*into file names and tries to run them as a command. It did no harm here; we switched to reading only the keys we need withgrep.
Related
- Claude Code "temporarily unavailable, so auto mode cannot determine the safety of bash": how to fix it
- Inside the Claude Code auto mode classifier: its 116K-character system prompt, section by section
- Claude Code auto mode uses a model to review a model: when that model is down, even cat stops working
- Claude Code with DeepSeek as the backend: tested
- Claude Code 429: what happens and how long it retries (tested)