Magic Tools

Hands-On

Every article here has actually been run by the author: reproduction conditions, tested environment, methodology, results and analysis are all documented. No secondhand claims — when something doesn't reproduce, we say so.

Repro conditionsEnvironmentMethodologyResultsAnalysis
llama.cpp llama-server vs Ollama vllama.cpp b11376 / Ollama 0.35.1✅ ReproducibleTested 2026-10-03

llama.cpp's llama-server or Ollama? Same GGUF on a 24GB Mac mini, Tested: Ollama Runs llama-server Under the Hood

llama-server (b11376) and Ollama (0.35.1) loading the same GGUF file on a 24GB Mac mini. Ollama 0.35's inference process is its own bundled llama-server, and with the same settings speed is identical (4B decode 36.5 vs 36.3 tok/s, 27B 6.3 on both). Every difference comes from defaults: Ollama defaults to a 4096-token context and silently cuts an over-long prompt for library models down to 2050 tokens, leaving only a WARN line in its log; llama-server sizes context to fill memory (101,120 for the 4B, 14GB+ footprint) and returns 400 when a prompt doesn't fit. llama-server runs 4 parallel slots by default while Ollama queues, and Ollama with 32k context plus 4-way parallelism allocates 32k per slot, a 21GB footprint. Recommended flags for 24GB machines at the end.

Read the report
Claude Code / Codex CLI v2.1.280 / 0.157.1✅ ReproducibleTested 2026-09-28

DeepSeek Harness Test: One Model, Three Harnesses — Claude Code 15/15, Codex CLI 15/15, Bare API 0/15 (and 5 Fake "Done"s)

Same DeepSeek model (deepseek-v4-pro), three harnesses, five tasks (read / write / edit / run a command / multi-step), three rounds each, every side effect checked on disk. Claude Code on DeepSeek's Anthropic endpoint: 15/15, median 4.17s, ¥0.159 per task. Codex CLI 0.157.1 on the Responses endpoint: 15/15, median 15.52s, ¥0.022 per task — one seventh. Bare chat/completions: 0/15, and 5 of those rounds replied DONE or EDITED with nothing on disk. The differences are the harness: DeepSeek partitions its prompt cache by metadata.user_id, so every `claude -p` pays ~15K uncached tokens; Codex has no file tools and does everything through shell; on HTTP 500 Claude Code retries 10 times over ~175s while Codex quits in ~25s, and on 429 Codex doesn't retry; on 120KB of output Claude Code shows the first 2KB, Codex head + tail. And wire_api = "chat" is gone in Codex 0.157.1 — use responses.

Read the report
Drosophila_brain_model (official repo of Shiu et al. 2024) vcommit 91bdd1e / FlyWire v783 / Brian2 2.10.1✅ ReproducibleTested 2026-09-28

How to Download the Fruit Fly Brain and Run It: 139,000 Neurons in 33 Seconds on a Mac mini

Downloading the fruit fly brain means two files: a list of 138,639 neurons and a 15.09-million-row connection table whose synapse counts add up to the 54.5 million in the Nature paper. I ran the official repo from scratch: clone 57 s, install 21 s, the sugar-taste example in 33 s. Switching to the public v783 data as the README says crashes with a KeyError, and a missing output folder only errors after the whole run. Silencing one descending neuron raised feeding output by 30%.

Read the report
FlyWire whole-brain LIF (Eon fly-brain / Shiu 2024) + NeuroMechFly v2 (flygym 2.1) vFlyWire v783 connectome / Brian2 C++ backend / MuJoCo 3.13 (Python) + 3.9 (WASM)✅ ReproducibleTested 2026-09-26

I Ran a 139,000-Neuron Fruit Fly Brain on a Mac mini. With No Input at All, Does It Move on Its Own?

I rebuilt Eon's uploaded fly from open parts: a 139k-neuron LIF brain in a MuJoCo body. With input it acts: feeding 0.66 s after tasting sugar, a giant-fiber escape 1.89 s into a looming ball. With no input it is silent. Adding noise and adaptation makes it move on its own, but v1 ticks like a clock (burst CV 0.05). v2 raises CV to 1.2 and walks for up to 4.6 s.

Read the report
Claude Code v2.1.280✅ ReproducibleTested 2026-09-25

Claude Code "Error: Reached max turns (1)": when the headless guardrail fires, your file may already be written

claude -p guardrails stop with Error: Reached max turns (1) or Error: Exceeded USD budget (0.01), exit 1. On 2.1.270 and 2.1.280 they are 28 and 33 bytes on stdout, not stderr, with no newline. In json mode the result key is missing, so jq -r .result prints null with jq exit 0. Failure does not mean nothing happened: --max-turns 2 errors after out.txt is written, and the budget is checked after each call, so a $0.05 cap spent $0.0517, finished the task and still exited 1.

Read the report
Claude Code v2.1.270✅ ReproducibleTested 2026-09-20

DeepSeek Says 1M, Claude Code Says 200K: I Measured Both and Neither Number Is the Real Limit

DeepSeek advertises a 1M context window. Point Claude Code at it and Claude Code reports contextWindow 200000 for the same model. I measured what actually happens. DeepSeek's real ceiling is 1,048,576 tokens — literally 2^20, not one million — and it covers input plus your max_tokens budget, proven with a controlled pair. A needle planted at position zero was retrieved correctly at 1,039,744 tokens. Claude Code refuses client-side long before that, in 25ms with zero API calls, and its gate is not on tokens at all: it fires at roughly 480,000 characters. Feed it high-entropy text and 478,000 characters sails through carrying 309,567 real tokens — 55% past the 200K window it just claimed. And in ordinary use you reach none of these, because Bash output over exactly 30,000 characters never enters context at all.

Read the report
Claude Code v2.1.270✅ ReproducibleTested 2026-09-20

Running Claude Code on DeepSeek: Everything Works, But the Cost Readout Lies by 38x

DeepSeek ships an Anthropic-format endpoint, so you can point Claude Code at it with three environment variables. I ran the whole thing on a real machine: every local tool (Read / Write / Bash / Glob / Edit / subagents) works and produces real side effects, so the short answer is yes, it works. The long answer is the part nobody measured — Claude Code bills DeepSeek tokens at Claude Sonnet rates. Ten identical turns: Claude Code reported $1.71, DeepSeek's actual balance dropped ¥0.32 (≈$0.045). That is a 38x over-report, measured against the invoice, not a price list. Also inside: the official docs are wrong about unknown model names (they 400, they don't fall back), v4-pro returns thinking blocks by default so a small max_tokens looks like an empty reply, and one failure that looks like DeepSeek's fault but isn't.

Read the report
MiniMax T2A v2 vs Azure Neural TTS vspeech-02-turbo / zh-CN-YunxiNeural + en-US-GuyNeural✅ ReproducibleTested 2026-09-05

MiniMax T2A v2 vs Azure Neural TTS, Benchmarked: 6–10× Latency Gap on the Same Text

My video factory wires up both MiniMax and Azure TTS. This benchmark measures them head to head: same Chinese + English text, 5 rounds each, unified 24kHz/mono/16bit output. MiniMax median latency 1.0–2.2s, RTF 0.08–0.18 (5–12× faster than real time); Azure median 6–21s, RTF near 1.0, tail jittering to 27s. Two causes, both with evidence: Azure Neural synthesizes at roughly real-time pace, and its eastasia endpoint is a trans-Pacific hop from mainland China (TLS jitter to 0.8s).

Read the report
Claude Code v2.1.241🌗 PartialTested 2026-08-31

Reproducing an Injection Chain That Cracks Claude Code Auto Mode: the Model Refuses the Malicious Binary, Then Writes Code That Pwns Itself

In late August embracethered published an attack chain where a plain 'summarize this page' request drags auto-mode Claude Code to a 60–80% code-execution rate — while Anthropic's commissioned third-party test reported 0.00%. I took the chain apart and tested it stage by stage in an isolated environment: the endpoint that nudges the model from WebFetch to curl, and the crux — the model's own 'safe' decision to refuse the unknown binary and write its own Python decoder instead lands straight on a same-name struct.py planted in the extracted directory. The deterministic parts (branching + module-shadow poison + mitigation controls) reproduce fully on my machine with real evidence; the live end couldn't complete a full RCE here because the classifier rate-limited and failed closed — flagged honestly. Ends with mitigations that actually help.

Read the report
Claude Code v2.1.241✅ ReproducibleTested 2026-08-29

Cracking Open Claude Code's Auto-Mode Classifier: A 116K-Char System Prompt, Dissected Line by Line

My earlier retest confirmed auto mode calls the session model as a classifier before each risky Bash — but what it receives stayed a black box. This time I captured the full request: a 116,879-char system prompt opening 'You are a security monitor for autonomous AI coding agents.' I quote it verbatim to dissect the threat model, two-tier rules (1 HARD BLOCK / 68 SOFT BLOCK / 17 ALLOW), and two-stage evaluation — stage 1 grades harm only, stage 2 layers intent on top. Every number read out this session.

Read the report
Claude Code v2.1.223🔍 Not reproducedTested 2026-08-06

claude install Creates a Broken %h Symlink? Clean Debian Test Comes Out Fine (Not Reproduced)

A Fedora user reports on GitHub that the official one-line installer leaves ~/.local/bin/claude as a broken symlink pointing at the literal %h/.local/share/claude/versions/2.1.220 — the %h placeholder never expanded to the home directory. We ran the exact same install command in a clean Debian 12 container: the resulting symlink is correct (%h properly expanded), and install.sh now ships 2.1.223 (the report was against 2.1.220). This is a not-reproduced field report: full environment, commands, and raw output are included, along with three candidate explanations and a one-minute manual fix if you are currently stuck on the broken link.

Read the report
Claude Code v2.1.220✅ ReproducibleTested 2026-08-06

Claude Code Says No conversation found to continue — Your -p Sessions Are Being Filtered Out

You run a headless claude -p command, then type claude --continue to pick it up interactively — and get No conversation found to continue, even though the session file sits right there on disk. This is a regression introduced in v2.1.90: the --resume picker's 'hide -p/SDK sessions' filter was mistakenly applied to --continue, which should continue the latest session unconditionally. We reproduced it end to end on macOS + v2.1.220 (the GitHub issue reports Fedora, so both platforms confirm), and verified two workarounds: headless -p --continue has always worked, and prefixing CLAUDE_CODE_ENTRYPOINT=sdk-cli lets interactive --continue bypass the filter with full history loaded.

Read the report
Claude Code v2.1.220✅ ReproducibleTested 2026-08-06

Which Model Is Claude Code Actually Using? settings.json vs ANTHROPIC_MODEL vs --model, Tested

Multi-agent setups, CI, and batch scripts all rest on one assumption: the session actually runs on the model you configured. But the model can be set in four places — settings.json, the ANTHROPIC_MODEL environment variable, the --model flag, and /model in-session — and the docs never give you one table saying which wins. Meanwhile a GitHub issue reports settings.json silently ignored on Windows, with a full day of work run on the wrong model and discarded. Using modelUsage as hard evidence, we tested the chain layer by layer: --model > ANTHROPIC_MODEL > project settings.json > built-in default, with the [1m] suffix form working too — plus a verification method more trustworthy than any UI hint.

Read the report
Claude Code v2.1.220✅ ReproducibleTested 2026-08-05

The Claude Code Hooks Stdin Trap: Python Heredocs Eat Your Hook JSON

The docs say hooks receive JSON via stdin — true. But parse it with a python heredoc (python3 - <<'EOF') and you hit a silent failure: the heredoc redirects python's stdin to the script itself, so json.load(sys.stdin) reads nothing. A best-practice hook then swallows the exception and exits 0 — no error anywhere, just a hook that 'somehow does not work'. The real debugging session, the two-line fix, and a lesson: silently fault-tolerant code is your enemy at debugging time.

Read the report
Claude Code v2.1.220✅ ReproducibleTested 2026-07-28

Claude Code's Auto Mode Judges a Model With a Model — When the Model Is Down, You Can't Even Run cat

Auto permission mode calls your session model to judge whether each Bash command is safe — the judge and the worker are the same model. When it goes unavailable you land in a counterintuitive half-paralysis: reading files works, but you can't run a single cat. This is a log of a real debugging session in which I proposed three entirely reasonable hypotheses and knocked all three down with controlled experiments, leaving exactly one dependable way out. Retested 2026-08-29 on v2.1.241 with a local logging proxy: the trap is half-fixed — routine read/write commands no longer touch the classifier, but the classifier itself, its session-model binding, and its fail-closed behavior are all still there, with captured payloads and a deterministic breaker experiment at the end.

Read the report