Claude Code "Prompt is too long" vs "maximum context length": Why Auto-Compact Works for One and Not the Other (Tested)
Short answers first
Prompt is too long · …: Claude Code recognized a context overflow. In a multi-turn conversation it auto-compacts older turns and retries; usually you do nothing. If it says single exchange cannot be compacted, the system prompt, tool definitions or attachments themselves are too big — cut MCP tools or attachments, or/clear.API Error: 400 This model's maximum context length is …(common with DeepSeek and gateways): Claude Code doesn't recognize this text, won't auto-compact, and every following turn fails the same way. Type/compact— tested, it recovers.- Permanent fix: tell Claude Code the model's real window so it compacts before hitting it:
- Model ID not starting with
claude-(e.g.glm-4.6,deepseek-…):export CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072(use your model's real limit)- Gateway model ID starting with
claude-:export CLAUDE_CODE_AUTO_COMPACT_WINDOW=131072(minimum 100000)- An overflow about
max_tokens: nothing to do; Claude Code lowersmax_tokensand retries automatically.
Background
As a conversation grows, or when you attach many files, you eventually hit the context limit. In Claude Code that shows up in two very different ways:
Prompt is too long · the request is ~210000 tokens (limit 200000) but this conversation is only ~12807 tokens — …
API Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1053707 tokens (1045707 in the messages, 8000 in the completion).
Many people never see the first one because Claude Code compacts in the background. People on DeepSeek or a gateway see the second one over and over until they /clear. The difference isn't the model — it's the error text the backend returns.
Why: matching on error text
The errors docs mention that Bedrock's overflow text is Input is too long for requested model., and Before v2.1.217, Claude Code didn't recognize the Bedrock wording, so auto-compact never triggered on it. Auto-compaction is triggered by recognizing the text.
The recognizers in the 2.1.285 binary, verbatim:
function BWn(e){let n=e.toLowerCase();return n.includes("prompt is too long")||n.includes("input is too long for requested model")}
function jWn(e){return e.toLowerCase().includes("context window")} // used for 413 only
function M0r(e){return e.toLowerCase().includes("input length and `max_tokens` exceed context limit")}
function kA(e){ … return BWn(e.message)||bI(e.message,"prompt_too_long")}
So Claude Code recognizes only these:
| Backend text (case-insensitive) | What Claude Code does |
|---|---|
contains prompt is too long |
context overflow → auto-compact |
contains input is too long for requested model (Bedrock) |
same |
HTTP 413 containing context window |
same |
input length and \max_tokens` exceed context limit: A + B > C` |
parses the numbers, lowers max_tokens, retries |
| anything else | an ordinary 400, shown as-is |
DeepSeek's overflow error is This model's maximum context length is 1048576 tokens. However, you requested … tokens (… in the messages, … in the completion) (captured from DeepSeek's Anthropic-compatible endpoint in our DeepSeek 1M test), and OpenAI-style gateways pass through similar maximum context length text. None of it matches.
Method
| Approach | Verdict | Why |
|---|---|---|
| Local stub returning each overflow text | ✅ used | Exact control of the text; we can see whether Claude Code sends a compaction request (debug log source=compact), how many messages the retry carries, and what max_tokens becomes |
| Actually fill 1M tokens against DeepSeek | ❌ | Costs real money each time and tests only one wording |
| Read the binary only | ❌ not alone | Shows what is recognized, not the full flow afterwards |
Isolation as in the 429 article: env -i, a fresh empty CLAUDE_CONFIG_DIR, a fake key, ANTHROPIC_BASE_URL and HTTPS_PROXY both pointing to the local stub.
Multi-turn setup: in one isolated config dir, run 3 turns against an all-OK stub (claude -p → claude -p -c → claude -p -c) to build history, then switch to a stub that returns the overflow error once and OK afterwards for turn 4. A single-turn conversation has no history to compact, so it can't show auto-compaction.
About the texts: the Anthropic one follows the binary's token-parsing regex prompt is too long[^0-9]*(\d+)\s*tokens?\s*>\s*(\d+); the DeepSeek one splices the two fragments we captured earlier (whether anything sat between them isn't verified, but neither fragment contains a recognized phrase); the OpenAI-style one uses the common code: context_length_exceeded shape.
Results
1. Single turn: recognized vs not
Each text ran twice with identical results (claude -p, exit and stdout verbatim):
| Backend returns | Exit | Claude Code prints |
|---|---|---|
Anthropic prompt is too long: 210000 tokens > 200000 maximum |
1 | Prompt is too long · the request is ~210000 tokens (limit 200000) but this conversation is only ~12539 tokens — the rest is system prompt, tool definitions, and attachment content. A single-exchange … |
Bedrock Input is too long for requested model. |
1 | Prompt is too long · this conversation is a single exchange and cannot be compacted — the request size comes mostly from system prompt, tool definitions, or attachments. |
HTTP 413 Request exceeds the model's context window |
1 | Prompt is too long |
input length and \max_tokens` exceed context limit: 190000 + 32000 > 200000` |
0 | OK (auto-retry, section 3) |
DeepSeek This model's maximum context length is … |
1 | API Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1053707 tokens (1045707 in the messages, 8000 in the completion). |
OpenAI-style gateway context_length_exceeded |
1 | API Error: 400 This model's maximum context length is 131072 tokens. … |
Recognized errors are rewritten as Prompt is too long · … with a token breakdown ("the request is 210K, but the conversation is only 12.5K; the rest is system prompt, tools and attachments"). Unrecognized ones are a bare API Error: 400.
In interactive mode:
❯ reply with exactly OK
⎿ Prompt is too long · the request is ~210000 tokens (limit 200000) but this conversation is only ~12807 tokens — the
rest is system prompt, tool definitions, and attachment content. A single-exchange conversation cannot be
compacted; reduce attached files/tools or start with less context. · /clear to start fresh
❯ reply with exactly OK
⏺ API Error: 400 This model's maximum context length is 1048576 tokens. However, you requested 1053707 tokens
(1045707 in the messages, 8000 in the completion).
2. Multi-turn: recognized texts auto-compact, unrecognized fail
Three normal turns, then turn 4's first request gets the overflow error, everything after is OK:
| Backend text | Turn 4 exit | Requests at the stub | Debug log |
|---|---|---|---|
| Anthropic | 0, prints OK |
400 (7 messages) → compaction request → retry (5 messages) → OK | source=compact, Forked agent [reactive-compact] |
| Bedrock | 0 | 400 → compaction → retry → OK | source=compact |
| 413 + context window | 0 | 400 → compaction → retry → OK | source=compact |
| DeepSeek | 1, API Error: 400 … |
1 request only | no source=compact |
| OpenAI-style gateway | 1 | 1 request only | no source=compact |
With a recognized text, auto-compaction is invisible: -p prints just OK, exit 0. With an unrecognized text, Claude Code doesn't even try to compact; the next turn sends the same long history and fails again.
(Footnote: with a non-claude- model ID such as glm-4.6, the unrecognized 400 is re-sent once unchanged. The debug log says why: an unrecognised HTTP 400 arrived while the dangerous-tool-use beta was on the request; retrying once without it — it drops an auto-mode request header and retries. Nothing to do with context; it still fails and still doesn't compact.)
3. max_tokens overflow: lowered automatically
When the backend returns input length and \max_tokens` exceed context limit: 190000 + 32000 > 200000, the retry's max_tokens` goes from 32000 to 9000: 200000 limit − 190000 input = 10000, minus a 1000 margin. One retry succeeds, single- or multi-turn, exit 0.
4. Manual /compact after an unrecognized error: works
After turn 4 failed on the DeepSeek text (exit 1), turn 5 was claude -p -c '/compact': one compaction request (source=compact), exit 0. Turn 6 carried a single message (the compacted summary), exit 0.
5. Compact early: declare the real window
Better not to wait for the backend error. The stub reported 128,000 input tokens per turn for the first 3 turns (a nearly full conversation); turn 4 was all OK. Does Claude Code compact before sending the main request?
| Model ID | Setting | Turn 4 | Verdict |
|---|---|---|---|
glm-4.6 |
none | main request sent directly (11 messages) | control: it thinks there's room |
glm-4.6 |
CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072 |
compaction first; main request carried 2 messages, exit 0 | ✅ works |
claude-sonnet-4-5 |
CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072 |
main request sent directly, no compaction | ❌ no effect on claude- IDs (matches the docs) |
claude-sonnet-4-5 |
CLAUDE_CODE_AUTO_COMPACT_WINDOW=131072 |
compaction first, exit 0 | ✅ works |
claude-sonnet-4-5 |
same, but 60,000 tokens per turn | no compaction | control: below threshold, no compaction |
Per the docs, CLAUDE_CODE_MAX_CONTEXT_TOKENS applies directly to IDs that don't start with claude-; for IDs that resolve to a known Claude model it only applies together with DISABLE_COMPACT (which turns off all compaction). For a gateway that hands you a claude- ID, CLAUDE_CODE_AUTO_COMPACT_WINDOW is the better fit (docs: 100000–1000000, plain integers only — 500k reads as 500 and is raised to the minimum).
What to do
1. Identify which error it is. Prompt is too long · … is recognized and handled in multi-turn conversations; API Error: 400 This model's maximum context length … is not, and needs you.
2. Unrecognized: type /compact. Tested to work; the next turn is back to normal. Or /clear if you don't need the history.
3. On DeepSeek, other non-Anthropic models, or a gateway, set this permanently (use your model's real context limit):
# Model ID not starting with claude- (glm-4.6, deepseek-…, kimi-…)
export CLAUDE_CODE_MAX_CONTEXT_TOKENS=131072
# Gateway model ID starting with claude-, but the real limit behind it is smaller
export CLAUDE_CODE_AUTO_COMPACT_WINDOW=131072
Claude Code then compacts as it nears the limit and never hits the error text it can't recognize. Take the limit from the model's official docs; many backends count input + output together, so leave some margin.
4. Overflow on the very first turn (single exchange cannot be compacted): compaction can't help; the bulk is system prompt, tool definitions and attachments. Enable fewer MCP servers, @ fewer large files, or read big files in parts.
5. Gateway developers: if your gateway translates protocols, include prompt is too long in the overflow error text; Claude Code users then get auto-compaction for free.
All requests went to the local stub. Separately, to check how DeepSeek handles an oversized max_tokens, we called DeepSeek directly twice (it silently clamps max_tokens and answers normally; about 190 tokens total), not through Claude Code.
Pitfalls along the way
- Single-turn runs can't show auto-compaction. The first batch was all single-turn; recognized texts just errored out, which nearly led to "auto-compaction doesn't happen". The hint single exchange cannot be compacted was the clue: there was no history to compact.
- In interactive mode the error went to the side request. The "error once, then OK" stub was consumed by the second request each prompt sends, so the main request got OK. Recording the error screen needed an "always error" stub.
- An accidental success almost misled the conclusion. Testing manual
/compactwithglm-4.6, turn 4 unexpectedly succeeded — the debug log showed the beta-header retry had simply received the stub's next OK response. Re-testing with an "always error" stub confirmed unrecognized texts never trigger compaction, under any model ID.
Not verified: DeepSeek's full overflow text (we used a splice of two captured fragments); real gateways' exact texts (they vary); compaction summary quality (the stub's summary is just OK); DISABLE_COMPACT + CLAUDE_CODE_MAX_CONTEXT_TOKENS (turns off all compaction, not recommended, untested); OAuth subscription logins; the local pre-check where Claude Code rejects an oversized prompt before sending it (see the DeepSeek 1M test).