Real-world engineering pitfalls documented as they happened: symptom, root cause, and verified fix. Calendar math, time zones, CI/CD, LLM apps — every post is a hands-on postmortem from building and running this site, not theory.
Loading...
In auto mode, a blocked Bash call usually shows one of three messages: denied by auto mode, Auto mode could not evaluate this action, or auto mode unavailable for this model. I ran 40-odd real sessions on Claude Code 2.1.280. denied means the classifier judged the action out of scope, most often [Code from External]: it would run external code you never named. Retrying won't help. Name the source in your prompt, or declare it trusted in autoMode.environment in user-level settings. could not evaluate means the classifier returned no usable verdict. unavailable for this model means the model is older than claude-opus-4-6; under claude -p it is silently downgraded, with only a WARN line in the debug log. In 2.1.280 the verdict is computed server-side and returned with the main response. Behind a relay gateway that only admits Claude Code clients, the local fallback classifier request gets a 503, which is a reliable cause of "temporarily unavailable (server error)".
28 cases, 61 claude -p runs on Claude Code 2.1.280 against a local Messages API stub. A 401 is retried 10 times, so Invalid API key · Fix external API key shows up after ~3 minutes; a key with non-ASCII chars or an embedded newline is rejected locally with 0 requests in 0.28 s. Every error goes to stdout, stderr is 0 bytes, exit 1, and JSON subtype still says success. CLAUDE_CODE_MAX_RETRIES=0 fails a 401 in 0.28 s. Set both KEY and TOKEN and both headers are sent.
Claude Code's Bash tool times out after 120 seconds by default. I ran 19 real sessions on 2.1.280 and found two timeout paths: only commands whose first word is sleep get killed with Exit code 143 / Command timed out after 2m 0s; everything else is moved to the background and killed 5 seconds after claude -p winds down. Either way claude exits 0, stderr is 0 bytes and the JSON top level says is_error=false. BASH_DEFAULT_TIMEOUT_MS=8000 killed at 8.17s; 0 or abc silently fall back to 120s; an explicit timeout above BASH_MAX_TIMEOUT_MS was silently clamped to 15s.
Six causes behind Claude Code's MCP "Failed to connect", reproduced on 2.1.280: mcp list exits 0 on failure, CONNECTION_CLOSED hides two causes that only --debug-file reveals, and one server that never handshakes pushes claude -p wall time 4.63s → 36.21s, duration_ms only 4323.5 → 5930.
A Los Angeles VPS running sing-box (VLESS-REALITY + Hysteria2) lost its VPN the day after Tailscale was installed. systemctl, ports and certificates were all fine. The root cause was in /etc/resolv.conf: Tailscale manages DNS by default, the tailnet had no global nameservers, and when dhclient renewed its lease tailscaled read an empty resolv.conf and dropped its upstream list. From then on every public domain got SERVFAIL, and the REALITY handshake could not even resolve www.apple.com. Full timeline, the evidence for each step, three fixes, and the rules we added to CLAUDE.md so an AI assistant (Claude Code) does not walk into this again.
Claude Code auto mode pops up 'temporarily unavailable, so auto mode cannot determine the safety of bash'? First, the conclusion: it's not your command that's dangerous; it's the safety classifier (an additional model call) that's temporarily unavailable. This article provides a four-step fix, a complete variant lookup for model name × tool name × reason, and a mechanism explanation for why read-only operations are unaffected.
A production pipeline went dark two mornings in a row: logs stopped at 07:04, and every poll after that said 'record exists for today, skipping.' The root cause was PM2's cron_restart on a wake-every-5-minutes dispatcher — it kills the running instance (and its children) before starting a new one. Includes an 8-minute local reproduction: the cron_restart group started 9 times and finished 0; the resident-loop control finished 4 out of 4.
I exported ANTHROPIC_BASE_URL in .zshrc to point at a self-hosted API gateway, and Claude Code kept talking to Google Vertex anyway. On the same machine, a launchd-managed web UI insisted it wasn't authenticated at all. Neither bug was in the gateway — both were in the gap between 'I set the env var' and 'the process actually has it.'
We built WeChat QR login on an Official Account, passed every local test, then hit 48001 in production. After ruling out four false suspects (IP whitelist, token cache, the console permission page, business domain), WeChat's official rid diagnosis revealed the truth: the parametric QR API only serves verified non-individual Service Accounts — a verified enterprise Subscription Account doesn't qualify, and the types can't be converted. Includes the rid/quota diagnostic toolkit and a workaround any subscription account can use: fixed QR + 6-digit code, with hand-minted NextAuth database sessions.
A mascot scene card driven by SwiftUI's TimelineView had two defects: the snail mirror-flipped in place the first time you started a sound, and the background hard-cut when you stopped it. The first one took two rounds to bottom out — round one blamed a paused timeline desyncing the two clocks, and the flip survived the fix. The actual cause was one line of phase normalization, `raw < 0 ? raw + 1 : raw`, which took the few-dozen-millisecond fact that timeline.date trails Date() by a frame and wrapped it into a phase of 0.9999 — and facing direction happens to be a discontinuous function of phase at zero. This postmortem covers why .transition doesn't work inside TimelineView, how to express animation state as a pure function of time, and why a crossfade should be fade-in only, with no fade-out.
Building explainer videos with TTS + auto subtitles: the first 30 seconds were perfectly synced, then the subtitles drifted further and further behind, and the tail of the .srt degraded to 00:00:00 timestamps. The first instinct — patch the alignment (normalize numbers, loosen similarity thresholds) — treats symptoms. This postmortem shows the real cause: any 'whole-clip TTS + whisper transcription + post-hoc matching' pipeline drifts by construction, because it hands timing — which the generation step could produce directly — to a lossy reconstruction chain. The fix: sentence-level TTS with sample-accurate concatenation, where each subtitle's timestamp is the running sum of real wav sample counts. Sentence boundaries reset the error, so drift is physically impossible. Includes splitting rules, caching design, and the limits of the approach.
Two Cloudflare stdio MCP processes pegged a CPU core each, with a stack of zombie instances from old sessions behind them. How to diagnose runaway MCP servers, remove them from Claude Code cleanly, and why low-frequency ops are better served by the REST API than a resident MCP.