Magic Tools
Hands-OnBy CooconSeptember 28, 20268 views13 min read

DeepSeek Harness Test: One Model, Three Harnesses — Claude Code 15/15, Codex CLI 15/15, Bare API 0/15 (and 5 Fake "Done"s)

People search for "DeepSeek harness" because they have noticed something odd: the same DeepSeek model feels brilliant in one coding tool and clumsy in another. The model didn't change. The harness did — everything wrapped around the model: the system prompt, the tool definitions and their names, the agent loop, the retry and timeout policy, and what happens when a tool's output is too big to fit.

So I held the model fixed and swapped the harness. One machine, one day (2026-09-28), one model — deepseek-v4-pro — and three ways of driving it:

Group Harness DeepSeek endpoint
A Claude Code 2.1.280 https://api.deepseek.com/anthropic (Anthropic format)
B Codex CLI 0.157.1 (npx @openai/codex) https://api.deepseek.com with wire_api = "responses"
C none — one bare POST /chat/completions with the task as the only user message https://api.deepseek.com/chat/completions

The short answer to "does the harness matter?" is: it decides whether anything happens at all, it decides the bill by a factor of seven, and it decides what the model gets to see when things go wrong.

Same DeepSeek model, three harnesses: tool matrix, 3 rounds each

Background: why swap the harness and not the model

Most "DeepSeek vs X" comparisons change two things at once: the model and the tool around it. That makes the result uninterpretable. If DeepSeek in Codex is slower than Claude in Claude Code, is that DeepSeek or Codex?

DeepSeek is unusual in that it exposes two agent-friendly wire formats on the same account — an Anthropic-compatible endpoint (which is how you run Claude Code on DeepSeek) and an OpenAI Responses-compatible endpoint. That makes it possible to put the two most widely used vendor harnesses, Claude Code and Codex CLI, on top of literally the same model and the same account, and measure the difference.

The problem, broken down

"Harness" is fuzzy, so I split it into the pieces that could plausibly change an outcome, and measured each one:

  1. Tools — which tools exist, what they are called, and whether the model uses them.
  2. Prompt weight — how big the injected system prompt and tool schema are, and what that does to cost.
  3. Caching — whether that weight gets re-billed on every run.
  4. Failure policy — what happens on HTTP 500, 429, and a hung connection.
  5. Output truncation — what the model is shown when a command prints 120KB.

The bare-API group C is the control: same model, same prompts, no harness at all.

Approach and what I deliberately left out

What was run. Five tasks, three rounds per group per task, each round in a fresh empty working directory outside any repository (so no project instruction file could leak in and distort the prompt-size numbers):

  • T1-read — reply with the contents of probe.txt.
  • T2-write — create written.txt containing WRITTEN_OK.
  • T3-edit — change port=8080 to port=9090 in config.ini, nothing else.
  • T4-bash — run sh gen.sh, which writes a fresh random nonce from /dev/urandom to nonce.txt, and reply with it. The value cannot be guessed by reading the script.
  • T5-multi — read sales.csv, sum it with a shell command, write the total to total.txt, append verified to log.txt, reply with the total (605).

A sixth task, T6-bigout, ran sh big.sh (6,000 lines, 120,000 characters of pseudo-random tokens) to probe truncation.

Pass/fail came only from disk: after every round the script saved ls -la, shasum -a 256 and cat of the working directory, and judged from that — never from what the model said it did.

The exact invocations:

# A — Claude Code, isolated CLAUDE_CONFIG_DIR
ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic \
  claude -p "<prompt>" --output-format stream-json --verbose --dangerously-skip-permissions

# B — Codex CLI, isolated CODEX_HOME
codex exec --skip-git-repo-check -C <workdir> --ephemeral --json -s workspace-write "<prompt>"

# C — no harness
POST /chat/completions {"model":"deepseek-v4-pro","messages":[{"role":"user","content":"<prompt>"}]}

The two harnesses do not get identical permissions — A runs with --dangerously-skip-permissions, B in its workspace-write sandbox (which does not prompt in exec mode). Both could write inside the working directory, which is all these tasks need.

Isolation. Each harness ran under its own config directory (CLAUDE_CONFIG_DIR for A, CODEX_HOME for B), so no global hooks, settings, or skills were loaded. I broke this once for Codex; see the pitfalls.

Failure injection and packet capture went through a small local logging reverse proxy in front of api.deepseek.com (auth headers stripped before logging). In 500 / 429 / hang mode it never forwards anything, so those tests cost nothing.

Left out, and why:

  • Other harnesses (Aider, OpenCode, Cline…). The budget for this run was ¥5 of API spend, and I'd rather measure two harnesses properly than five shallowly. Claude Code and Codex CLI were picked because each maps to one of DeepSeek's two native agent wire formats.
  • deepseek-flash. One model, held fixed, is the whole point.
  • Interactive sessions. Everything is headless (claude -p, codex exec); interactive sessions were not measured.
  • Code quality on large tasks. These are deliberately small, verifiable tasks. They answer "can the model act through this harness, and at what cost," not "which one writes better code."

Setting up Codex CLI on DeepSeek (the part that breaks first)

Most guides floating around configure Codex with wire_api = "chat". On 0.157.1 that is dead:

Error loading config.toml: `wire_api = "chat"` is no longer supported.
How to fix: set `wire_api = "responses"` in your provider config.

DeepSeek serves the Responses format, so this config works:

# $CODEX_HOME/config.toml
model = "deepseek-v4-pro"
model_provider = "deepseek"

[model_providers.deepseek]
name = "DeepSeek"
base_url = "https://api.deepseek.com"
env_key = "DEEPSEEK_API_KEY"
wire_api = "responses"

Every run then starts with a warning — Model metadata for deepseek-v4-pro not found. Defaulting to fallback metadata; this can degrade performance and cause issues. That warning turns out to matter (see "Tools" below).

Codex CLI 0.157.1: wire_api=chat rejected, responses works against DeepSeek

Claude Code needs the usual three variables; the details and its cost-display quirk are in the earlier Claude Code + DeepSeek write-up.

Results

1. Success rate: both harnesses 15/15, no harness 0/15

Task A Claude Code B Codex CLI C bare API
T1-read 3/3 · 3.2s 3/3 · 13.2s 0/3 · 6.8s
T2-write 3/3 · 3.7s 3/3 · 14.7s 0/3 · 30.3s
T3-edit 3/3 · 6.1s 3/3 · 14.6s 0/3 · 12.4s
T4-bash 3/3 · 4.0s 3/3 · 15.5s 0/3 · 9.7s
T5-multi 3/3 · 7.5s 3/3 · 19.9s 0/3 · 38.9s
T1–T5 15/15, median 4.17s 15/15, median 15.52s 0/15, median 18.3s

(Times are per-task medians of three wall-clock runs.)

The disk agrees byte for byte: written.txt is 10 bytes with the same SHA-256 (f82f148fd68c…) in all six A and B rounds; the edited config.ini is 46 bytes, same hash (86893adb55e6…) in all six; nonce.txt matched the reply every time. The one divergence: one Codex round wrote total.txt as 605\n (4 bytes) instead of 605 (3 bytes). Harmless, but it shows the two harnesses don't write files the same way — more on that below.

Group C is the real finding. No harness means no tools, so 0/15 is expected. What is not expected is that 5 of the 15 rounds claimed success: three DONEs on the write task and two EDITEDs on the edit task, with nothing on disk and config.ini still at its original hash. One write reply was literally <file_write path="written.txt" content="WRITTEN_OK" />DONE, and one read reply emitted tool-call markup as plain text (<tool_calls><invoke name="read_file">…). The rest honestly said they couldn't touch files.

That is what the harness is actually for. The model is perfectly willing to describe an action and report it done. Without a harness to execute the call and feed back the real result, you get a confident "DONE" and an empty directory. And C was not cheap for it: without tools to call, it reasoned the most — up to 5,157 reasoning tokens in a single T5 round. (C is n=3 per task; it's a control, not a benchmark.)

2. Tools: Claude Code has file tools; Codex has a shell

What each harness actually sends with a one-line Reply with exactly: PONG (captured through the proxy):

A Claude Code B Codex CLI C
Request body 60,876 bytes 39,827 bytes ~100 bytes
System instructions 5,765 chars (2 blocks) 16,979 chars + a 2,876-char skills message + 432-char environment context none
Tools registered 20, schema 46,438 chars 9, schema 17,775 chars 0
Input tokens for "PONG" 15,106 8,944 tiny

Claude Code's 20 tools include dedicated Read, Write, Edit and Bash, and the model used them exactly as named: Read for T1, Write for T2, Edit for T3, Read → Bash → Write → Bash for T5 in all three rounds.

Codex's nine tools contain no file tool at all. Everything — reading, writing, editing — went through exec_command running shell: cat, printf 'WRITTEN_OK' > written.txt, sed -i '', perl -pi, awk.

And here is where the fallback-metadata warning bites. Codex's own system prompt tells the model: "Use the apply_patch tool to edit files." But with fallback metadata, apply_patch is not registered in the tool list. The instruction and the toolset contradict each other. Across all 15 rounds the model called apply_patch zero times and fell back to shell. It cost one round a mistake: on T3 round 3 it first ran GNU-style sed -i 's/…/', which exits 1 on macOS, then corrected itself to BSD sed -i '' — four shell calls instead of two, 23.6s instead of ~14.6s.

That's a clean example of a harness-caused difference. The model behaved sensibly; the harness handed it a prompt describing a tool that didn't exist.

3. Cost: Codex was 7x cheaper, and it's a caching quirk

Cost from token usage × DeepSeek's published v4-pro peak price (all runs fell in peak hours: $0.044 / 1M cache hit, $1.32 / 1M cache miss, $3.96 / 1M output; ¥7.1 per USD):

T1–T5, 15 rounds each uncached input cached input output total per task
A Claude Code 230,876 396,800 3,116 ¥2.38 ¥0.159
B Codex CLI 8,868 379,008 4,433 ¥0.33 ¥0.022
C bare API 1,463 256 23,440 ¥0.67 ¥0.045 (achieved nothing)

Both harnesses send a heavy prompt. The difference is whether it gets billed at the miss price. Codex's first request in each run was already cached (T1: 17,920 of 18,450 input tokens hit). Claude Code's first request missed every single time — about 15.1K tokens at full price, per run.

My first theory was that the working directory, which Claude Code writes into its system prompt, was breaking the prefix. Wrong: three runs from the same directory still showed cache_read = 0 on the first request. So I diffed two request bodies from same-directory runs. model, system, tools, messages, max_tokens, thinking — all identical. The only difference was metadata.user_id, which carries a per-session session_id.

A five-call control against DeepSeek's Anthropic endpoint, same 8.8K-token prompt, only the metadata changing:

# metadata.user_id input cache_read
1 sessA (prime) 8,799 0
2 sessA 95 8,704
3 sessB 8,799 0
4 none 8,799 0
5 sessA again 95 8,704

DeepSeek's Anthropic endpoint partitions the prompt cache by metadata.user_id

DeepSeek's Anthropic endpoint partitions its prompt cache by metadata.user_id. Claude Code puts a new session id there every session, so every claude -p starts cold and pays ~15K tokens at the miss price; only the later requests inside the same session hit. In a long interactive session that should be a one-time cost per session (inferred from the mechanism, not measured). For scripted, many-short-invocations use — CI, batch jobs, the way I ran it here — it is the dominant cost. Codex's request carries a prompt_cache_key field; I didn't isolate whether that is why its cache survives across runs.

4. Failure policy: 175 seconds vs 25 seconds, and a 429 that isn't retried

With the proxy returning errors instead of forwarding (no API cost), prompt Reply with exactly: PONG:

Injected A Claude Code B Codex CLI
HTTP 500 11 requests (1 + 10 retries), backoff 0.5 → 1 → 2 → 5 → 9 → 17s, then ~32–40s; gives up after ~175–178s (2 runs) 30 requests, gives up after ~25s (3 runs), and reports We're currently experiencing high demand… — misleading when it's your own provider failing
HTTP 429 same shape: 11 requests, gives up after ~180–183s (2 runs) 1 request, no retry, fails in 0.5–0.7s with exceeded retry limit, last status: 429 Too Many Requests (3 runs)
hang (never responds) aborts the request at ~360s and retries (n=1) no timeout within 400s, when I killed it (n=1)

Injected 500 / 429 / hang: requests seen by the proxy for each harness

Neither is "right". Claude Code rides out a three-minute outage; Codex fails fast and lets your script decide. But if you run Codex against DeepSeek and hit rate limits, expect immediate failures rather than backoff.

Two caveats I haven't closed: my 429 responses carried no Retry-After header, and I did not test whether Codex would retry if one were present — unverified. And each hang result is a single observation (every hang test burns 400 seconds), so treat those numbers as indicative.

5. Truncation: the first 2KB, or the head and the tail

sh big.sh prints 6,000 lines / 120,000 characters. What each harness actually put in front of the model, captured through the proxy:

A Claude Code B Codex CLI
Fed back to the model a <persisted-output> block: "Output too large (117.2KB). Full output saved to: …/tool-results/….txt" plus a 2KB preview (lines 1–100) — 2,296 characters head + tail, 10,212 characters / 508 lines, prefixed with Warning: truncated output (original token count: 30000) and Total output lines: 6000; the last line is visible
How it got the last line 3/3 rounds made a second Bash call (tail / wc) on the saved file one call
Result 2/3 3/3

120 KB of output: what Claude Code and Codex each hand back to the model

Claude Code's approach loses nothing — the full output is on disk and the model knows where — but costs an extra round trip. Codex's head+tail wins on a "what's the last line" task and would lose on anything in the middle. (The one A "failure" answered 63822ab7 and 6000 — correct token, correct count — without repeating the full line; it's a format miss, not a blind spot.) The Codex model also reported seeing 500 lines in two rounds and 6,000 in one — the harness's header says 6,000, but 508 lines were actually shown.

6. Latency: Claude Code was ~4x faster, and I can't fully say why

Median over the 15 rounds of T1–T5: A 4.17s, B 15.52s. What I ruled out: npx startup is 0.4–0.5s, Codex's zsh -lc wrapper 0.01s, and a single upstream PONG took about the same time on both endpoints (1.7s vs 1.8s). Codex under an injected 429 exits in 0.5–0.7s, so the CLI itself is fast.

What's left is between Codex's model calls and its local tool execution, and Codex's JSONL events have no timestamps. The one place I did capture upstream timings (T6), Codex's second /responses request took 12–55s — heavy reasoning — but for T1–T5 the 13–15s is not attributed. Splitting it needs another timestamped proxy run, which I didn't do for budget reasons.

What the model decides vs what the harness decides

Difference Model or harness? Evidence
Whether any file changes at all Harness C 0/15 with 5 false "done"s vs A/B 15/15
Which tools get used Harness 20 named tools vs one shell tool; apply_patch in the prompt but not registered
Cost per task (7x) Harness × endpoint cache partitioned by metadata.user_id, not by prompt
Behaviour on 500 / 429 / hang Harness 11 vs 30 attempts; 429 retried vs not
What a 120KB output looks like to the model Harness 2KB head + file vs head+tail 508 lines
Recovering from its own sed mistake Model same model, noticed exit 1, switched to BSD syntax
Summing a CSV correctly (605) Model identical in A and B

If you only take one thing away: when "DeepSeek" looks bad in a tool, check the harness column first.

Practical takeaways

  • Codex CLI + DeepSeek: use wire_api = "responses"; ignore every guide that says chat. Expect shell-only file editing (watch for GNU vs BSD sed on macOS), fail-fast on 429, and no hang timeout.
  • Claude Code + DeepSeek in scripts: every claude -p pays ~15K uncached tokens. If you're firing many short headless runs, that is most of your bill — batch the work into fewer sessions. Interactive use should be affected far less (not measured).
  • Bare API "agents": never trust the reply. Verify on disk. A third of my bare-API rounds said DONE for work that never happened.

Pitfalls, condensed

  • Codex writes to ~/.codex even for --version. I ran codex --version and codex exec --help before exporting CODEX_HOME, and Codex created ~/.codex/tmp/arg0/… (its apply_patch / codex-execve-wrapper shims) in my real home. Export CODEX_HOME before the first Codex command. Once set, nothing more landed there. I left the created files in place for manual review.
  • wire_api = "chat" is removed in 0.157.1; DeepSeek works with responses.
  • codex exec reads stdin (Reading additional input from stdin...). In scripts, give it </dev/null or it waits.
  • Fallback metadata means Codex's prompt references apply_patch while the tool isn't registered.
  • DeepSeek's /user/balance settles 3–6 minutes late. I read an idle-period drop as someone else using the key; it was my own earlier calls landing. A before/after that showed ¥10.98 → ¥10.98 around six runs gave it away. Per-group balance deltas are therefore worthless — account with usage × price instead. Usage-based total ¥6.05 vs actual balance drop ¥6.13 (16.22 → 10.09), 1.3% apart, so the key had no other consumer.
  • I went over budget. The cap for this run was ¥5; I spent ¥6.13, because I didn't keep a running usage tally and assumed Claude Code's cache would kick in across runs. All API calls stopped once I noticed.
  • A stray rm -f /dev/null got typed during the session. It failed on permissions and /dev/null is intact — mentioning it because this log is supposed to be complete.

Cost by usage × price vs the balance log, which settles 3–6 minutes late

Related: Claude Code on DeepSeek — the cost readout lies by 38x · DeepSeek's 1M context vs Claude Code's 200K

Related Articles

Claude Code "bash denied by auto mode": Why It Blocks, "could not evaluate" and "unavailable for this model" Tested

Claude Code "bash denied by auto mode": Why It Blocks, "could not evaluate" and "unavailable for this model" Tested

In auto mode, a blocked Bash call usually shows one of three messages: denied by auto mode, Auto mode could not evaluate this action, or auto mode unavailable for this model. I ran 40-odd real sessions on Claude Code 2.1.280. denied means the classifier judged the action out of scope, most often [Code from External]: it would run external code you never named. Retrying won't help. Name the source in your prompt, or declare it trusted in autoMode.environment in user-level settings. could not evaluate means the classifier returned no usable verdict. unavailable for this model means the model is older than claude-opus-4-6; under claude -p it is silently downgraded, with only a WARN line in the debug log. In 2.1.280 the verdict is computed server-side and returned with the main response. Behind a relay gateway that only admits Claude Code clients, the local fallback classifier request gets a 503, which is a reliable cause of "temporarily unavailable (server error)".

claude-codepermissions+4
pitfallsSep 27, 202611 min
23
Claude Code "Invalid API key · Fix external API key": Not logged in, Credit balance is too low, API Error 401/429/529 — Exact Messages and Retry Behavior, Tested

Claude Code "Invalid API key · Fix external API key": Not logged in, Credit balance is too low, API Error 401/429/529 — Exact Messages and Retry Behavior, Tested

28 cases, 61 claude -p runs on Claude Code 2.1.280 against a local Messages API stub. A 401 is retried 10 times, so Invalid API key · Fix external API key shows up after ~3 minutes; a key with non-ASCII chars or an embedded newline is rejected locally with 0 requests in 0.28 s. Every error goes to stdout, stderr is 0 bytes, exit 1, and JSON subtype still says success. CLAUDE_CODE_MAX_RETRIES=0 fails a 401 in 0.28 s. Set both KEY and TOKEN and both headers are sent.

claude-codetroubleshooting+4
pitfallsSep 27, 202611 min
37
Claude Code "Command timed out after 2m 0s": Two Timeout Paths, BASH_DEFAULT_TIMEOUT_MS and run_in_background Tested

Claude Code "Command timed out after 2m 0s": Two Timeout Paths, BASH_DEFAULT_TIMEOUT_MS and run_in_background Tested

Claude Code's Bash tool times out after 120 seconds by default. I ran 19 real sessions on 2.1.280 and found two timeout paths: only commands whose first word is sleep get killed with Exit code 143 / Command timed out after 2m 0s; everything else is moved to the background and killed 5 seconds after claude -p winds down. Either way claude exits 0, stderr is 0 bytes and the JSON top level says is_error=false. BASH_DEFAULT_TIMEOUT_MS=8000 killed at 8.17s; 0 or abc silently fall back to 120s; an explicit timeout above BASH_MAX_TIMEOUT_MS was silently clamped to 15s.

claude-codetroubleshooting+4
pitfallsSep 26, 20269 min
44
Claude Code "Error: Reached max turns (1)": when the headless guardrail fires, your file may already be written

Claude Code "Error: Reached max turns (1)": when the headless guardrail fires, your file may already be written

claude -p guardrails stop with Error: Reached max turns (1) or Error: Exceeded USD budget (0.01), exit 1. On 2.1.270 and 2.1.280 they are 28 and 33 bytes on stdout, not stderr, with no newline. In json mode the result key is missing, so jq -r .result prints null with jq exit 0. Failure does not mean nothing happened: --max-turns 2 errors after out.txt is written, and the budget is checked after each call, so a $0.05 cap spent $0.0517, finished the task and still exited 1.

claude-codeautomation+5
hands-onSep 25, 202611 min
74