Magic Tools
Pitfall NotesBy CooconOctober 5, 20268 views7 min read

Claude Code "Prompt is too long" vs "maximum context length": Why Auto-Compact Works for One and Not the Other (Tested)

Short answers first

  • Prompt is too long · …: Claude Code recognized a context overflow. In a multi-turn conversation it auto-compacts older turns and retries; usually you do nothing. If it says single exchange cannot be compacted, the system prompt, tool definitions or attachments themselves are too big — cut MCP tools or attachments, or /clear.
  • API Error: 400 This model's maximum context length is … (common with DeepSeek and gateways): Claude Code doesn't recognize this text, won't auto-compact, and every following turn fails the same way. Type /compact — tested, it recovers.
  • Permanent fix: tell Claude Code the model's real window so it compacts before hitting it:
    • Model ID not starting with claude- (e.g. glm-4.6, deepseek-…): export CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072 (use your model's real limit)
    • Gateway model ID starting with claude-: export CLAUDE_CODE_AUTO_COMPACT_WINDOW=131072 (minimum 100000)
  • An overflow about max_tokens: nothing to do; Claude Code lowers max_tokens and retries automatically.

Background

As a conversation grows, or when you attach many files, you eventually hit the context limit. In Claude Code that shows up in two very different ways:

Prompt is too long · the request is ~210000 tokens (limit 200000) but this conversation is only ~12807 tokens — …
API Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1053707 tokens (1045707 in the messages, 8000 in the completion).

Many people never see the first one because Claude Code compacts in the background. People on DeepSeek or a gateway see the second one over and over until they /clear. The difference isn't the model — it's the error text the backend returns.

Why: matching on error text

The errors docs mention that Bedrock's overflow text is Input is too long for requested model., and Before v2.1.217, Claude Code didn't recognize the Bedrock wording, so auto-compact never triggered on it. Auto-compaction is triggered by recognizing the text.

The recognizers in the 2.1.285 binary, verbatim:

function BWn(e){let n=e.toLowerCase();return n.includes("prompt is too long")||n.includes("input is too long for requested model")}
function jWn(e){return e.toLowerCase().includes("context window")}          // used for 413 only
function M0r(e){return e.toLowerCase().includes("input length and `max_tokens` exceed context limit")}
function kA(e){ … return BWn(e.message)||bI(e.message,"prompt_too_long")}

So Claude Code recognizes only these:

Backend text (case-insensitive) What Claude Code does
contains prompt is too long context overflow → auto-compact
contains input is too long for requested model (Bedrock) same
HTTP 413 containing context window same
input length and \max_tokens` exceed context limit: A + B > C` parses the numbers, lowers max_tokens, retries
anything else an ordinary 400, shown as-is

DeepSeek's overflow error is This model's maximum context length is 1048576 tokens. However, you requested … tokens (… in the messages, … in the completion) (captured from DeepSeek's Anthropic-compatible endpoint in our DeepSeek 1M test), and OpenAI-style gateways pass through similar maximum context length text. None of it matches.

Method

Approach Verdict Why
Local stub returning each overflow text ✅ used Exact control of the text; we can see whether Claude Code sends a compaction request (debug log source=compact), how many messages the retry carries, and what max_tokens becomes
Actually fill 1M tokens against DeepSeek ❌ Costs real money each time and tests only one wording
Read the binary only ❌ not alone Shows what is recognized, not the full flow afterwards

Isolation as in the 429 article: env -i, a fresh empty CLAUDE_CONFIG_DIR, a fake key, ANTHROPIC_BASE_URL and HTTPS_PROXY both pointing to the local stub.

Multi-turn setup: in one isolated config dir, run 3 turns against an all-OK stub (claude -p → claude -p -c → claude -p -c) to build history, then switch to a stub that returns the overflow error once and OK afterwards for turn 4. A single-turn conversation has no history to compact, so it can't show auto-compaction.

About the texts: the Anthropic one follows the binary's token-parsing regex prompt is too long[^0-9]*(\d+)\s*tokens?\s*>\s*(\d+); the DeepSeek one splices the two fragments we captured earlier (whether anything sat between them isn't verified, but neither fragment contains a recognized phrase); the OpenAI-style one uses the common code: context_length_exceeded shape.

Results

1. Single turn: recognized vs not

Each text ran twice with identical results (claude -p, exit and stdout verbatim):

Backend returns Exit Claude Code prints
Anthropic prompt is too long: 210000 tokens > 200000 maximum 1 Prompt is too long · the request is ~210000 tokens (limit 200000) but this conversation is only ~12539 tokens — the rest is system prompt, tool definitions, and attachment content. A single-exchange …
Bedrock Input is too long for requested model. 1 Prompt is too long · this conversation is a single exchange and cannot be compacted — the request size comes mostly from system prompt, tool definitions, or attachments.
HTTP 413 Request exceeds the model's context window 1 Prompt is too long
input length and \max_tokens` exceed context limit: 190000 + 32000 > 200000` 0 OK (auto-retry, section 3)
DeepSeek This model's maximum context length is … 1 API Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1053707 tokens (1045707 in the messages, 8000 in the completion).
OpenAI-style gateway context_length_exceeded 1 API Error: 400 This model's maximum context length is 131072 tokens. …

Recognized errors are rewritten as Prompt is too long · … with a token breakdown ("the request is 210K, but the conversation is only 12.5K; the rest is system prompt, tools and attachments"). Unrecognized ones are a bare API Error: 400.

In interactive mode:

❯ reply with exactly OK
  ⎿  Prompt is too long · the request is ~210000 tokens (limit 200000) but this conversation is only ~12807 tokens — the
     rest is system prompt, tool definitions, and attachment content. A single-exchange conversation cannot be
     compacted; reduce attached files/tools or start with less context. · /clear to start fresh
❯ reply with exactly OK
⏺ API Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1053707 tokens
  (1045707 in the messages, 8000 in the completion).

2. Multi-turn: recognized texts auto-compact, unrecognized fail

Three normal turns, then turn 4's first request gets the overflow error, everything after is OK:

Backend text Turn 4 exit Requests at the stub Debug log
Anthropic 0, prints OK 400 (7 messages) → compaction request → retry (5 messages) → OK source=compact, Forked agent [reactive-compact]
Bedrock 0 400 → compaction → retry → OK source=compact
413 + context window 0 400 → compaction → retry → OK source=compact
DeepSeek 1, API Error: 400 … 1 request only no source=compact
OpenAI-style gateway 1 1 request only no source=compact

With a recognized text, auto-compaction is invisible: -p prints just OK, exit 0. With an unrecognized text, Claude Code doesn't even try to compact; the next turn sends the same long history and fails again.

(Footnote: with a non-claude- model ID such as glm-4.6, the unrecognized 400 is re-sent once unchanged. The debug log says why: an unrecognised HTTP 400 arrived while the dangerous-tool-use beta was on the request; retrying once without it — it drops an auto-mode request header and retries. Nothing to do with context; it still fails and still doesn't compact.)

3. max_tokens overflow: lowered automatically

When the backend returns input length and \max_tokens` exceed context limit: 190000 + 32000 > 200000, the retry's max_tokens` goes from 32000 to 9000: 200000 limit − 190000 input = 10000, minus a 1000 margin. One retry succeeds, single- or multi-turn, exit 0.

4. Manual /compact after an unrecognized error: works

After turn 4 failed on the DeepSeek text (exit 1), turn 5 was claude -p -c '/compact': one compaction request (source=compact), exit 0. Turn 6 carried a single message (the compacted summary), exit 0.

5. Compact early: declare the real window

Better not to wait for the backend error. The stub reported 128,000 input tokens per turn for the first 3 turns (a nearly full conversation); turn 4 was all OK. Does Claude Code compact before sending the main request?

Model ID Setting Turn 4 Verdict
glm-4.6 none main request sent directly (11 messages) control: it thinks there's room
glm-4.6 CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072 compaction first; main request carried 2 messages, exit 0 ✅ works
claude-sonnet-4-5 CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072 main request sent directly, no compaction ❌ no effect on claude- IDs (matches the docs)
claude-sonnet-4-5 CLAUDE_CODE_AUTO_COMPACT_WINDOW=131072 compaction first, exit 0 ✅ works
claude-sonnet-4-5 same, but 60,000 tokens per turn no compaction control: below threshold, no compaction

Per the docs, CLAUDE_CODE_MAX_CONTEXT_TOKENS applies directly to IDs that don't start with claude-; for IDs that resolve to a known Claude model it only applies together with DISABLE_COMPACT (which turns off all compaction). For a gateway that hands you a claude- ID, CLAUDE_CODE_AUTO_COMPACT_WINDOW is the better fit (docs: 100000–1000000, plain integers only — 500k reads as 500 and is raised to the minimum).

What to do

1. Identify which error it is. Prompt is too long · … is recognized and handled in multi-turn conversations; API Error: 400 This model's maximum context length … is not, and needs you.

2. Unrecognized: type /compact. Tested to work; the next turn is back to normal. Or /clear if you don't need the history.

3. On DeepSeek, other non-Anthropic models, or a gateway, set this permanently (use your model's real context limit):

# Model ID not starting with claude- (glm-4.6, deepseek-…, kimi-…)
export CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072

# Gateway model ID starting with claude-, but the real limit behind it is smaller
export CLAUDE_CODE_AUTO_COMPACT_WINDOW=131072

Claude Code then compacts as it nears the limit and never hits the error text it can't recognize. Take the limit from the model's official docs; many backends count input + output together, so leave some margin.

4. Overflow on the very first turn (single exchange cannot be compacted): compaction can't help; the bulk is system prompt, tool definitions and attachments. Enable fewer MCP servers, @ fewer large files, or read big files in parts.

5. Gateway developers: if your gateway translates protocols, include prompt is too long in the overflow error text; Claude Code users then get auto-compaction for free.

All requests went to the local stub. Separately, to check how DeepSeek handles an oversized max_tokens, we called DeepSeek directly twice (it silently clamps max_tokens and answers normally; about 190 tokens total), not through Claude Code.

Pitfalls along the way

  • Single-turn runs can't show auto-compaction. The first batch was all single-turn; recognized texts just errored out, which nearly led to "auto-compaction doesn't happen". The hint single exchange cannot be compacted was the clue: there was no history to compact.
  • In interactive mode the error went to the side request. The "error once, then OK" stub was consumed by the second request each prompt sends, so the main request got OK. Recording the error screen needed an "always error" stub.
  • An accidental success almost misled the conclusion. Testing manual /compact with glm-4.6, turn 4 unexpectedly succeeded — the debug log showed the beta-header retry had simply received the stub's next OK response. Re-testing with an "always error" stub confirmed unrecognized texts never trigger compaction, under any model ID.

Not verified: DeepSeek's full overflow text (we used a splice of two captured fragments); real gateways' exact texts (they vary); compaction summary quality (the stub's summary is just OK); DISABLE_COMPACT + CLAUDE_CODE_MAX_CONTEXT_TOKENS (turns off all compaction, not recommended, untested); OAuth subscription logins; the local pre-check where Claude Code rejects an oversized prompt before sending it (see the DeepSeek 1M test).

Get field notes like this every Saturday

Subscribe to Dev Breakfast: daily AI coding picks at 8:00, plus a Saturday roundup of this week's hands-on tests with Claude Code / Codex / local models. Written in Chinese.

Related Articles

Claude Code Stuck on a Spinner With No Response: It Waits 6 Minutes per Attempt, and 10 Retries Can Hang It for Over an Hour (Tested)

When Claude Code spins without output or sits at Retrying in 0s, the backend is usually not sending anything. Tested on 2.1.285 (API key + ANTHROPIC_BASE_URL): with no response headers it waits 6 minutes (360 s) per attempt before timing out — setting API_TIMEOUT_MS to 600000 or 900000 doesn't change that; only lower values work. With the default 10 retries it can hang for over an hour. A gateway that buffers the reply until generation finishes will never deliver a reply that takes over 6 minutes. Press Esc to interrupt. All tested against a local stub.

claude-codetroubleshooting+5
pitfallsOct 5, 20269 min
6

Claude Code 429 "Request rejected (429)": How Long It Retries, and Why retry-after Over 60 Seconds Fails Instantly (Tested)

On a 429, Claude Code retries up to 10 times over about 3 minutes. If retry-after is 60 seconds or less it waits the full time (10 s and 60 s both recovered in testing); at 61 seconds or more it does not retry at all and fails immediately with API Error: Request rejected (429). Gateway messages such as new-api's are shown verbatim; an empty body shows status code (no body). In interactive mode every prompt sends 2 requests, so rate-limit usage doubles. All tested against a local stub on Claude Code 2.1.285.

claude-codetroubleshooting+4
pitfallsOct 4, 20268 min
11

Clash Verge TUN Mode Breaks All Internet Access: Hysteria2 Traffic Loops Back Into the TUN, and Tailscale Hijacks DNS

Clash Verge Rev 2.5.6 on macOS works fine in system-proxy mode, but as soon as TUN (virtual network adapter) mode is on, not even Baidu loads. Debugging through the mihomo core's API turned up two independent root causes stacked together. First, Hysteria2's outbound UDP isn't bound to the physical interface, so the TUN routes pull it back in and it loops. Second, Tailscale MagicDNS (100.100.100.100) has taken over system DNS, so queries go out through Tailscale's utun where Clash's dns-hijack can't see them, and come back poisoned. The fix is two Merge overrides: route-exclude-address to keep node IPs out of the TUN, and sniffer to recover domains from SNI. Every step comes with the commands and real output.

troubleshootingtailscale+7
pitfallsOct 4, 20266 min
31
Claude Code MCP Shows Connected but 0 Tools: "Invalid result for tools/list" (ttlMs / cacheScope), Reproduced and Fixed

Claude Code MCP Shows Connected but 0 Tools: "Invalid result for tools/list" (ttlMs / cacheScope), Reproduced and Fixed

An MCP server shows as connected but exposes 0 tools, and the log says Invalid result for tools/list with ttlMs and cacheScope failing validation. Reproduced with a stub server on Claude Code 2.1.280, 2.1.285 and 2.1.288: the cause is not extra fields being rejected. The server negotiated MCP 2026-07-28 and then left out fields that revision requires (resultType, ttlMs, cacheScope); adding unknown fields works fine. Whether stdio uses the new protocol is decided by a remote flag that is off by default, so the same version breaks for some users and not others, and 2.1.280 fails the same way once negotiation is on. MCP_PROTOCOL_NEGOTIATION=legacy (also works in settings.json env) restores the tools; servers fix it by adding the three fields.

mcpclaude-code+4
pitfallsOct 3, 20268 min
20

Published by Magic Tools