Articles
Industry news, technical articles, and product introductions
📚 Claude Tutorials Hub
40+ step-by-step Claude guides — prompt engineering, Claude Code, API, agents. Browse by topic →
Loading...
Industry news, technical articles, and product introductions
📚 Claude Tutorials Hub
40+ step-by-step Claude guides — prompt engineering, Claude Code, API, agents. Browse by topic →
Loading...
On a 24GB M4 Mac mini, Ollama 0.19 sees only 17.8 GiB of VRAM and defaults to a 4096-token context. With qwen3:4b, going from 4k to 32k context grows memory from 3.73GB to 9.89GB; OLLAMA_NUM_PARALLEL=4 at 8k uses exactly as much as a single 32k slot while ollama ps still shows 8192; Flash Attention + q4_0 KV cache brings 32k down to 4.31GB and 64k to 5.77GB with no speed loss. Ask for a context far beyond RAM and Ollama doesn't refuse — it starts allocating, and free memory fell to 23% within 16 seconds. Running fully on CPU was only 25% slower.
One week after DFlash 2 shipped, I got Qwen3.8-27B with speculative decoding fully working on a 24GB Mac mini M4: 6.5 tok/s to 11.7–12.2 tok/s at 4-bit, a stable 1.8–1.9x. This post covers the exact deployment commands, three controlled benchmark rounds, the GB-by-GB memory budget, and the three concrete reasons the official 2.7–3.4x number shrinks on consumer Apple Silicon. An Aug 29 retest adds a block-size and draft-precision sweep: block-size 8 collapses to 1.11x (the official cliff warning is real), block-size 3 beats the default, and an 8-bit draft loses to 4-bit.