Articles
Industry news, technical articles, and product introductions
📚 Claude Tutorials Hub
40+ step-by-step Claude guides — prompt engineering, Claude Code, API, agents. Browse by topic →
Loading...
Industry news, technical articles, and product introductions
📚 Claude Tutorials Hub
40+ step-by-step Claude guides — prompt engineering, Claude Code, API, agents. Browse by topic →
Loading...
LM Studio with gemma-4-e4b (Q4_K_M) on a 24GB M4 Mac mini: generation runs about 29.5 tok/s, and turning GPU offload from max to off costs only 11%. What really feels slow is prompt prefill — a 13K-token prompt waited 39.4 seconds for the first token (about 330 tokens/s), and 3.4–3.7x longer on CPU. A second request with the same prefix got its first token in 0.1 s thanks to prefix caching. With --parallel 4 and four simultaneous requests, each dropped to 9.7 tok/s.
On a 24GB M4 Mac mini, Ollama 0.19 sees only 17.8 GiB of VRAM and defaults to a 4096-token context. With qwen3:4b, going from 4k to 32k context grows memory from 3.73GB to 9.89GB; OLLAMA_NUM_PARALLEL=4 at 8k uses exactly as much as a single 32k slot while ollama ps still shows 8192; Flash Attention + q4_0 KV cache brings 32k down to 4.31GB and 64k to 5.77GB with no speed loss. Ask for a context far beyond RAM and Ollama doesn't refuse — it starts allocating, and free memory fell to 23% within 16 seconds. Running fully on CPU was only 25% slower.
Qwen3.8 ships a trained multi-token-prediction head, and llama.cpp can mount it with one flag — no separate 2B draft model. I benchmarked it against DFlash 2 on the same 24GB Mac mini M4: memory does drop (16.0GB vs 19.4GB peak), but speed goes backwards — prose falls from 6.0 to 4.5 tok/s (-24%) while DFlash 2 delivers 1.8–1.9x on the same machine. Draft acceptance is healthy (59–85%); the loss is in Metal's verify path — batch-8 decode amortizes at just 1.13x, measured.
A Mac mini at home runs Claude Code around the clock — but how do you take over a session from a browser when you're away? This is a real setup that has been live for a week and in daily use: claudecodeui as the web UI (chosen over the official web version, ttyd, and code-server), an SSH reverse tunnel pushing it to a VPS, and nginx adding TLS plus login rate limiting to turn it into an ordinary URL. Includes full configs, real operating numbers (five days of tunnel uptime with zero drops, 170MB RSS), a <synthetic> placeholder bug hit and fixed within the first week, and an honest for-and-against on why not Tailscale.
One week after DFlash 2 shipped, I got Qwen3.8-27B with speculative decoding fully working on a 24GB Mac mini M4: 6.5 tok/s to 11.7–12.2 tok/s at 4-bit, a stable 1.8–1.9x. This post covers the exact deployment commands, three controlled benchmark rounds, the GB-by-GB memory budget, and the three concrete reasons the official 2.7–3.4x number shrinks on consumer Apple Silicon. An Aug 29 retest adds a block-size and draft-precision sweep: block-size 8 collapses to 1.11x (the official cliff warning is real), block-size 3 beats the default, and an 8-bit draft loses to 4-bit.