What LLMs can a Mac mini M4 24GB run?
At the recommended quantization with 16K context, the largest model a Mac mini M4 24GB can run is Nemotron 3.5 Lightning (30B-A3B) (30B, about 16.4 GB); 5 models on the list run comfortably. With 120GB/s of memory bandwidth, the decode ceiling is roughly bandwidth ÷ weight size; measured speeds usually land at 70-85% of that. This page includes 9 first-hand measurements we ran on this exact machine (tok/s, peak memory, all verifiable in the source articles). Data verified 2026-10-03.
Specs: 24GB · 120GB/s · Our primary test machine (base M4, 10-core GPU, macOS 26.3)
First-hand benchmarks9 runs
Measured by us on this exact machine with real model files; every row links to the article with commands, versions and full readings.
| Model | Quant / runtime | Setup | Decode tok/s | Prefill | Peak mem | Source |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | 4-bit (MLX) MLX | Plain autoregressive baseline, 3 runs | 6.4–6.5 | — | — | article 2026-08-29 |
| Qwen3.8-27B | 4-bit (MLX) MLX | DFlash 2 draft model (2B, 4-bit) speculative decoding, block-size 3 block-size 8 drops to 1.11x; an 8-bit draft is slower (1.62x) | 11.6–12.21.8–1.9x | — | 19.4GB | article 2026-08-29 |
| Qwen3.8-27B | UD-Q4_K_S (GGUF, 15.4GB) llama.cpp b96806d (Metal) | Plain autoregressive baseline (prose / code) | 5.9–6 | — | 15.4GB | article 2026-09-02 |
| Qwen3.8-27B | UD-Q4_K_S (GGUF) llama.cpp b96806d (Metal) | Native MTP head speculative decoding, n-max 2 Prose -24%, code -4%: a net slowdown on Metal | 4.5–5.70.76–0.96x | — | 16GB | article 2026-09-02 |
| Qwen3.8-27B | UD-Q4_K_S (GGUF) llama.cpp b96806d (Metal) | DFlash 2 GGUF draft (1.14GB), n-max 3 The recommended n-max 7 hits a reproducible Metal OOM at 8K context | 3–4.60.5–0.77x | — | — | article 2026-09-03 |
| Qwen3.8-27B | UD-Q4_K_S (GGUF) llama.cpp llama-server b11376 | Defaults (--fit caps context at 30K, 4 slots) | 6.3 | 56 | — | article 2026-10-03 |
| Qwen3.8-27B | UD-Q4_K_S (GGUF, 自导入) Ollama 0.35.1 | Defaults (4096 context) Ollama 0.35's inference process is its bundled llama-server; same flags, same speed | 6.3 | 57 | — | article 2026-10-03 |
| Qwen3 4B | Q4_K_M (GGUF) llama.cpp llama-server b11376 | Defaults (--fit opens 101K context, 4 slots) With 4 concurrent requests: 11.6 tok/s each, 40.9 tok/s aggregate | 36.5 | 376 | — | article 2026-10-03 |
| Qwen3 4B | Q4_K_M (GGUF, 官方库 qwen3:4b) Ollama 0.35.1 | Defaults (4096 context, single slot, queued) Over-long prompts are silently truncated to 2050 tokens; the only trace is one WARN log line | 36.3 | 367 | — | article 2026-10-03 |
Model fit list (recommended quant, 16K context)
Comfortable = ≤ 65% of memory; Tight = ≤ 82% (macOS wires only ~2/3-3/4 of unified memory to the GPU; thresholds calibrated on our 24GB measurements); Lower quant = the recommended quant doesn't fit but a smaller one does. Ceiling = 120GB/s ÷ weight size; single-request decoding cannot exceed it.
| Model | Quant | Needs | Verdict | Decode ceiling tok/s |
|---|---|---|---|---|
| Qwen3-4B | Q4_K_M (GGUF) | 6.0 GB | Comfortable | ≤ 53 |
| Qwen3.5-9B | Q4_K_M (GGUF) | 7.5 GB | Comfortable | ≤ 22 |
| Llama 3.1 8B | Q4_K_M (GGUF) | 8.0 GB | Comfortable | ≤ 26 |
| Qwen3-14B | Q4_K_M (GGUF) | 12.4 GB | Comfortable | ≤ 14 |
| gpt-oss-20B | MXFP4 | 12.6 GB | Comfortable | ≤ 12 |
| Nemotron 3.5 Lightning (30B-A3B) | NVFP4 | 16.4 GB | Tight | ≤ 8.1 |
| Gemma 4 26B-A4B | Q4_K_M (GGUF) | 16.5 GB | Tight | ≤ 8.2 |
| Qwen3.8-27B | Q4_K_M (GGUF) | 17.3 GB | Tight | ≤ 7.8 |
| Mistral Small 3.x 24B | Q4_K_M (GGUF) | 17.6 GB | Tight | ≤ 8.8 |
| GLM-4.7-Flash (30B-A3B) | Q3_K_M (GGUF) | 16.0 GB | Lower quant | ≤ 8.8 |
| Qwen3-Coder-30B-A3B | Q3_K_M (GGUF) | 16.4 GB | Lower quant | ≤ 9.0 |
| Gemma 4 31B | Q3_K_M (GGUF) | 16.5 GB | Lower quant | ≤ 8.8 |
| Qwen3.6-35B-A3B | Q3_K_M (GGUF) | 17.6 GB | Lower quant | ≤ 7.6 |
| Qwen3-32B | Q3_K_M (GGUF) | 19.5 GB | Lower quant | ≤ 8.6 |
| Gemma 3 27B | Q2_K (GGUF) | 18.1 GB | Lower quant | ≤ 14 |
| Llama 3.3 70B | Q4_K_M (GGUF) | 47.9 GB | Won't fit | — |
| gpt-oss-120B | MXFP4 | 63.5 GB | Won't fit | — |
| Mistral Leanstral 1.5 (119B-A6B MoE) | Q4_K_M (GGUF) | 73.4 GB | Won't fit | — |
| Qwen3.8-Flash-Next (180B MoE) | Q4_K_M (GGUF) | 111 GB | Won't fit | — |
| Qwen3-235B-A22B | Q4_K_M (GGUF) | 147 GB | Won't fit | — |
| DeepSeek V4 Flash (304B MoE) | Q4_K_M (GGUF) | 187 GB | Won't fit | — |
| GLM-4.5 (355B MoE) | Q4_K_M (GGUF) | 224 GB | Won't fit | — |
| DeepSeek V3 / R1 (671B MoE) | Q4_K_M (GGUF) | 413 GB | Won't fit | — |
| GLM-5.3 (753B MoE) | Q4_K_M (GGUF) | 463 GB | Won't fit | — |
| Kimi K2 (1T MoE) | Q4_K_M (GGUF) | 631 GB | Won't fit | — |
| DeepSeek V4 Pro (1.6T MoE) | Q4_K_M (GGUF) | 983 GB | Won't fit | — |
| Kimi K3 (2.8T MoE) | MXFP4 | 1.46 TB | Won't fit | — |
Benchmarked a Mac mini M4 24GB yourself?
Send us the model, quant, runtime version, tok/s and peak memory. Verified readings are added to this page with credit. First-hand numbers with a command or log only.
Submit a benchmarkFAQ
What is the largest LLM a Mac mini M4 24GB can run?
At the recommended quantization with 16K context, the ceiling is Nemotron 3.5 Lightning (30B-A3B) (about 16.4 GB, 68% of 24GB). macOS lets the GPU wire only ~2/3-3/4 of unified memory by default; near the limit, raise iogpu.wired_limit_mb and close memory-heavy apps.
Which models run best on a Mac mini M4 24GB?
Models under 16GB leave room for long context without closing other apps, e.g. gpt-oss-20B, Qwen3.5-9B, Qwen3-14B, Llama 3.1 8B.
How many tokens per second does a Mac mini M4 24GB get?
Single-request decoding is bandwidth-bound: ceiling ≈ 120GB/s ÷ weight size. A 15GB 27B 4-bit model tops out around 8.0 tok/s; measured speeds are usually 70-85% of that, and speculative decoding (a draft model) adds another 1.5-2x.
Are these numbers measured or estimated?
The fit table is a planning estimate (same formula as the calculator). The first-hand benchmark table contains readings we took on this exact machine with real model files; every row links to the source article.