DeepSeek Harness Test: One Model, Three Harnesses — Claude Code 15/15, Codex CLI 15/15, Bare API 0/15 (and 5 Fake "Done"s)
People search for "DeepSeek harness" because they have noticed something odd: the same DeepSeek model feels brilliant in one coding tool and clumsy in another. The model didn't change. The harness did — everything wrapped around the model: the system prompt, the tool definitions and their names, the agent loop, the retry and timeout policy, and what happens when a tool's output is too big to fit.
So I held the model fixed and swapped the harness. One machine, one day (2026-09-28), one model — deepseek-v4-pro — and three ways of driving it:
| Group | Harness | DeepSeek endpoint |
|---|---|---|
| A | Claude Code 2.1.280 | https://api.deepseek.com/anthropic (Anthropic format) |
| B | Codex CLI 0.157.1 (npx @openai/codex) |
https://api.deepseek.com with wire_api = "responses" |
| C | none — one bare POST /chat/completions with the task as the only user message |
https://api.deepseek.com/chat/completions |
The short answer to "does the harness matter?" is: it decides whether anything happens at all, it decides the bill by a factor of seven, and it decides what the model gets to see when things go wrong.

Background: why swap the harness and not the model
Most "DeepSeek vs X" comparisons change two things at once: the model and the tool around it. That makes the result uninterpretable. If DeepSeek in Codex is slower than Claude in Claude Code, is that DeepSeek or Codex?
DeepSeek is unusual in that it exposes two agent-friendly wire formats on the same account — an Anthropic-compatible endpoint (which is how you run Claude Code on DeepSeek) and an OpenAI Responses-compatible endpoint. That makes it possible to put the two most widely used vendor harnesses, Claude Code and Codex CLI, on top of literally the same model and the same account, and measure the difference.
The problem, broken down
"Harness" is fuzzy, so I split it into the pieces that could plausibly change an outcome, and measured each one:
- Tools — which tools exist, what they are called, and whether the model uses them.
- Prompt weight — how big the injected system prompt and tool schema are, and what that does to cost.
- Caching — whether that weight gets re-billed on every run.
- Failure policy — what happens on HTTP 500, 429, and a hung connection.
- Output truncation — what the model is shown when a command prints 120KB.
The bare-API group C is the control: same model, same prompts, no harness at all.
Approach and what I deliberately left out
What was run. Five tasks, three rounds per group per task, each round in a fresh empty working directory outside any repository (so no project instruction file could leak in and distort the prompt-size numbers):
- T1-read — reply with the contents of
probe.txt. - T2-write — create
written.txtcontainingWRITTEN_OK. - T3-edit — change
port=8080toport=9090inconfig.ini, nothing else. - T4-bash — run
sh gen.sh, which writes a fresh random nonce from/dev/urandomtononce.txt, and reply with it. The value cannot be guessed by reading the script. - T5-multi — read
sales.csv, sum it with a shell command, write the total tototal.txt, appendverifiedtolog.txt, reply with the total (605).
A sixth task, T6-bigout, ran sh big.sh (6,000 lines, 120,000 characters of pseudo-random tokens) to probe truncation.
Pass/fail came only from disk: after every round the script saved ls -la, shasum -a 256 and cat of the working directory, and judged from that — never from what the model said it did.
The exact invocations:
# A — Claude Code, isolated CLAUDE_CONFIG_DIR
ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic \
claude -p "<prompt>" --output-format stream-json --verbose --dangerously-skip-permissions
# B — Codex CLI, isolated CODEX_HOME
codex exec --skip-git-repo-check -C <workdir> --ephemeral --json -s workspace-write "<prompt>"
# C — no harness
POST /chat/completions {"model":"deepseek-v4-pro","messages":[{"role":"user","content":"<prompt>"}]}
The two harnesses do not get identical permissions — A runs with --dangerously-skip-permissions, B in its workspace-write sandbox (which does not prompt in exec mode). Both could write inside the working directory, which is all these tasks need.
Isolation. Each harness ran under its own config directory (CLAUDE_CONFIG_DIR for A, CODEX_HOME for B), so no global hooks, settings, or skills were loaded. I broke this once for Codex; see the pitfalls.
Failure injection and packet capture went through a small local logging reverse proxy in front of api.deepseek.com (auth headers stripped before logging). In 500 / 429 / hang mode it never forwards anything, so those tests cost nothing.
Left out, and why:
- Other harnesses (Aider, OpenCode, Cline…). The budget for this run was ¥5 of API spend, and I'd rather measure two harnesses properly than five shallowly. Claude Code and Codex CLI were picked because each maps to one of DeepSeek's two native agent wire formats.
deepseek-flash. One model, held fixed, is the whole point.- Interactive sessions. Everything is headless (
claude -p,codex exec); interactive sessions were not measured. - Code quality on large tasks. These are deliberately small, verifiable tasks. They answer "can the model act through this harness, and at what cost," not "which one writes better code."
Setting up Codex CLI on DeepSeek (the part that breaks first)
Most guides floating around configure Codex with wire_api = "chat". On 0.157.1 that is dead:
Error loading config.toml: `wire_api = "chat"` is no longer supported.
How to fix: set `wire_api = "responses"` in your provider config.
DeepSeek serves the Responses format, so this config works:
# $CODEX_HOME/config.toml
model = "deepseek-v4-pro"
model_provider = "deepseek"
[model_providers.deepseek]
name = "DeepSeek"
base_url = "https://api.deepseek.com"
env_key = "DEEPSEEK_API_KEY"
wire_api = "responses"
Every run then starts with a warning — Model metadata for deepseek-v4-pro not found. Defaulting to fallback metadata; this can degrade performance and cause issues. That warning turns out to matter (see "Tools" below).

Claude Code needs the usual three variables; the details and its cost-display quirk are in the earlier Claude Code + DeepSeek write-up.
Results
1. Success rate: both harnesses 15/15, no harness 0/15
| Task | A Claude Code | B Codex CLI | C bare API |
|---|---|---|---|
| T1-read | 3/3 · 3.2s | 3/3 · 13.2s | 0/3 · 6.8s |
| T2-write | 3/3 · 3.7s | 3/3 · 14.7s | 0/3 · 30.3s |
| T3-edit | 3/3 · 6.1s | 3/3 · 14.6s | 0/3 · 12.4s |
| T4-bash | 3/3 · 4.0s | 3/3 · 15.5s | 0/3 · 9.7s |
| T5-multi | 3/3 · 7.5s | 3/3 · 19.9s | 0/3 · 38.9s |
| T1–T5 | 15/15, median 4.17s | 15/15, median 15.52s | 0/15, median 18.3s |
(Times are per-task medians of three wall-clock runs.)
The disk agrees byte for byte: written.txt is 10 bytes with the same SHA-256 (f82f148fd68c…) in all six A and B rounds; the edited config.ini is 46 bytes, same hash (86893adb55e6…) in all six; nonce.txt matched the reply every time. The one divergence: one Codex round wrote total.txt as 605\n (4 bytes) instead of 605 (3 bytes). Harmless, but it shows the two harnesses don't write files the same way — more on that below.
Group C is the real finding. No harness means no tools, so 0/15 is expected. What is not expected is that 5 of the 15 rounds claimed success: three DONEs on the write task and two EDITEDs on the edit task, with nothing on disk and config.ini still at its original hash. One write reply was literally <file_write path="written.txt" content="WRITTEN_OK" />DONE, and one read reply emitted tool-call markup as plain text (<tool_calls><invoke name="read_file">…). The rest honestly said they couldn't touch files.
That is what the harness is actually for. The model is perfectly willing to describe an action and report it done. Without a harness to execute the call and feed back the real result, you get a confident "DONE" and an empty directory. And C was not cheap for it: without tools to call, it reasoned the most — up to 5,157 reasoning tokens in a single T5 round. (C is n=3 per task; it's a control, not a benchmark.)
2. Tools: Claude Code has file tools; Codex has a shell
What each harness actually sends with a one-line Reply with exactly: PONG (captured through the proxy):
| A Claude Code | B Codex CLI | C | |
|---|---|---|---|
| Request body | 60,876 bytes | 39,827 bytes | ~100 bytes |
| System instructions | 5,765 chars (2 blocks) | 16,979 chars + a 2,876-char skills message + 432-char environment context | none |
| Tools registered | 20, schema 46,438 chars | 9, schema 17,775 chars | 0 |
| Input tokens for "PONG" | 15,106 | 8,944 | tiny |
Claude Code's 20 tools include dedicated Read, Write, Edit and Bash, and the model used them exactly as named: Read for T1, Write for T2, Edit for T3, Read → Bash → Write → Bash for T5 in all three rounds.
Codex's nine tools contain no file tool at all. Everything — reading, writing, editing — went through exec_command running shell: cat, printf 'WRITTEN_OK' > written.txt, sed -i '', perl -pi, awk.
And here is where the fallback-metadata warning bites. Codex's own system prompt tells the model: "Use the apply_patch tool to edit files." But with fallback metadata, apply_patch is not registered in the tool list. The instruction and the toolset contradict each other. Across all 15 rounds the model called apply_patch zero times and fell back to shell. It cost one round a mistake: on T3 round 3 it first ran GNU-style sed -i 's/…/', which exits 1 on macOS, then corrected itself to BSD sed -i '' — four shell calls instead of two, 23.6s instead of ~14.6s.
That's a clean example of a harness-caused difference. The model behaved sensibly; the harness handed it a prompt describing a tool that didn't exist.
3. Cost: Codex was 7x cheaper, and it's a caching quirk
Cost from token usage × DeepSeek's published v4-pro peak price (all runs fell in peak hours: $0.044 / 1M cache hit, $1.32 / 1M cache miss, $3.96 / 1M output; ¥7.1 per USD):
| T1–T5, 15 rounds each | uncached input | cached input | output | total | per task |
|---|---|---|---|---|---|
| A Claude Code | 230,876 | 396,800 | 3,116 | ¥2.38 | ¥0.159 |
| B Codex CLI | 8,868 | 379,008 | 4,433 | ¥0.33 | ¥0.022 |
| C bare API | 1,463 | 256 | 23,440 | ¥0.67 | ¥0.045 (achieved nothing) |
Both harnesses send a heavy prompt. The difference is whether it gets billed at the miss price. Codex's first request in each run was already cached (T1: 17,920 of 18,450 input tokens hit). Claude Code's first request missed every single time — about 15.1K tokens at full price, per run.
My first theory was that the working directory, which Claude Code writes into its system prompt, was breaking the prefix. Wrong: three runs from the same directory still showed cache_read = 0 on the first request. So I diffed two request bodies from same-directory runs. model, system, tools, messages, max_tokens, thinking — all identical. The only difference was metadata.user_id, which carries a per-session session_id.
A five-call control against DeepSeek's Anthropic endpoint, same 8.8K-token prompt, only the metadata changing:
| # | metadata.user_id |
input | cache_read |
|---|---|---|---|
| 1 | sessA (prime) | 8,799 | 0 |
| 2 | sessA | 95 | 8,704 |
| 3 | sessB | 8,799 | 0 |
| 4 | none | 8,799 | 0 |
| 5 | sessA again | 95 | 8,704 |

DeepSeek's Anthropic endpoint partitions its prompt cache by metadata.user_id. Claude Code puts a new session id there every session, so every claude -p starts cold and pays ~15K tokens at the miss price; only the later requests inside the same session hit. In a long interactive session that should be a one-time cost per session (inferred from the mechanism, not measured). For scripted, many-short-invocations use — CI, batch jobs, the way I ran it here — it is the dominant cost. Codex's request carries a prompt_cache_key field; I didn't isolate whether that is why its cache survives across runs.
4. Failure policy: 175 seconds vs 25 seconds, and a 429 that isn't retried
With the proxy returning errors instead of forwarding (no API cost), prompt Reply with exactly: PONG:
| Injected | A Claude Code | B Codex CLI |
|---|---|---|
| HTTP 500 | 11 requests (1 + 10 retries), backoff 0.5 → 1 → 2 → 5 → 9 → 17s, then ~32–40s; gives up after ~175–178s (2 runs) | 30 requests, gives up after ~25s (3 runs), and reports We're currently experiencing high demand… — misleading when it's your own provider failing |
| HTTP 429 | same shape: 11 requests, gives up after ~180–183s (2 runs) | 1 request, no retry, fails in 0.5–0.7s with exceeded retry limit, last status: 429 Too Many Requests (3 runs) |
| hang (never responds) | aborts the request at ~360s and retries (n=1) | no timeout within 400s, when I killed it (n=1) |

Neither is "right". Claude Code rides out a three-minute outage; Codex fails fast and lets your script decide. But if you run Codex against DeepSeek and hit rate limits, expect immediate failures rather than backoff.
Two caveats I haven't closed: my 429 responses carried no Retry-After header, and I did not test whether Codex would retry if one were present — unverified. And each hang result is a single observation (every hang test burns 400 seconds), so treat those numbers as indicative.
5. Truncation: the first 2KB, or the head and the tail
sh big.sh prints 6,000 lines / 120,000 characters. What each harness actually put in front of the model, captured through the proxy:
| A Claude Code | B Codex CLI | |
|---|---|---|
| Fed back to the model | a <persisted-output> block: "Output too large (117.2KB). Full output saved to: …/tool-results/….txt" plus a 2KB preview (lines 1–100) — 2,296 characters |
head + tail, 10,212 characters / 508 lines, prefixed with Warning: truncated output (original token count: 30000) and Total output lines: 6000; the last line is visible |
| How it got the last line | 3/3 rounds made a second Bash call (tail / wc) on the saved file |
one call |
| Result | 2/3 | 3/3 |

Claude Code's approach loses nothing — the full output is on disk and the model knows where — but costs an extra round trip. Codex's head+tail wins on a "what's the last line" task and would lose on anything in the middle. (The one A "failure" answered 63822ab7 and 6000 — correct token, correct count — without repeating the full line; it's a format miss, not a blind spot.) The Codex model also reported seeing 500 lines in two rounds and 6,000 in one — the harness's header says 6,000, but 508 lines were actually shown.
6. Latency: Claude Code was ~4x faster, and I can't fully say why
Median over the 15 rounds of T1–T5: A 4.17s, B 15.52s. What I ruled out: npx startup is 0.4–0.5s, Codex's zsh -lc wrapper 0.01s, and a single upstream PONG took about the same time on both endpoints (1.7s vs 1.8s). Codex under an injected 429 exits in 0.5–0.7s, so the CLI itself is fast.
What's left is between Codex's model calls and its local tool execution, and Codex's JSONL events have no timestamps. The one place I did capture upstream timings (T6), Codex's second /responses request took 12–55s — heavy reasoning — but for T1–T5 the 13–15s is not attributed. Splitting it needs another timestamped proxy run, which I didn't do for budget reasons.
What the model decides vs what the harness decides
| Difference | Model or harness? | Evidence |
|---|---|---|
| Whether any file changes at all | Harness | C 0/15 with 5 false "done"s vs A/B 15/15 |
| Which tools get used | Harness | 20 named tools vs one shell tool; apply_patch in the prompt but not registered |
| Cost per task (7x) | Harness × endpoint | cache partitioned by metadata.user_id, not by prompt |
| Behaviour on 500 / 429 / hang | Harness | 11 vs 30 attempts; 429 retried vs not |
| What a 120KB output looks like to the model | Harness | 2KB head + file vs head+tail 508 lines |
Recovering from its own sed mistake |
Model | same model, noticed exit 1, switched to BSD syntax |
| Summing a CSV correctly (605) | Model | identical in A and B |
If you only take one thing away: when "DeepSeek" looks bad in a tool, check the harness column first.
Practical takeaways
- Codex CLI + DeepSeek: use
wire_api = "responses"; ignore every guide that sayschat. Expect shell-only file editing (watch for GNU vs BSDsedon macOS), fail-fast on 429, and no hang timeout. - Claude Code + DeepSeek in scripts: every
claude -ppays ~15K uncached tokens. If you're firing many short headless runs, that is most of your bill — batch the work into fewer sessions. Interactive use should be affected far less (not measured). - Bare API "agents": never trust the reply. Verify on disk. A third of my bare-API rounds said DONE for work that never happened.
Pitfalls, condensed
- Codex writes to
~/.codexeven for--version. I rancodex --versionandcodex exec --helpbefore exportingCODEX_HOME, and Codex created~/.codex/tmp/arg0/…(itsapply_patch/codex-execve-wrappershims) in my real home. ExportCODEX_HOMEbefore the first Codex command. Once set, nothing more landed there. I left the created files in place for manual review. wire_api = "chat"is removed in 0.157.1; DeepSeek works withresponses.codex execreads stdin (Reading additional input from stdin...). In scripts, give it</dev/nullor it waits.- Fallback metadata means Codex's prompt references
apply_patchwhile the tool isn't registered. - DeepSeek's
/user/balancesettles 3–6 minutes late. I read an idle-period drop as someone else using the key; it was my own earlier calls landing. A before/after that showed ¥10.98 → ¥10.98 around six runs gave it away. Per-group balance deltas are therefore worthless — account with usage × price instead. Usage-based total ¥6.05 vs actual balance drop ¥6.13 (16.22 → 10.09), 1.3% apart, so the key had no other consumer. - I went over budget. The cap for this run was ¥5; I spent ¥6.13, because I didn't keep a running usage tally and assumed Claude Code's cache would kick in across runs. All API calls stopped once I noticed.
- A stray
rm -f /dev/nullgot typed during the session. It failed on permissions and/dev/nullis intact — mentioning it because this log is supposed to be complete.

Related: Claude Code on DeepSeek — the cost readout lies by 38x · DeepSeek's 1M context vs Claude Code's 200K