Magic Tools
LLM VRAM Calculator

What LLMs can a Mac mini M4 24GB run?

At the recommended quantization with 16K context, the largest model a Mac mini M4 24GB can run is Nemotron 3.5 Lightning (30B-A3B) (30B, about 16.4 GB); 5 models on the list run comfortably. With 120GB/s of memory bandwidth, the decode ceiling is roughly bandwidth ÷ weight size; measured speeds usually land at 70-85% of that. This page includes 9 first-hand measurements we ran on this exact machine (tok/s, peak memory, all verifiable in the source articles). Data verified 2026-10-03.

Specs: 24GB · 120GB/s · Our primary test machine (base M4, 10-core GPU, macOS 26.3)

First-hand benchmarks9 runs

Measured by us on this exact machine with real model files; every row links to the article with commands, versions and full readings.

ModelQuant / runtimeSetupDecode tok/sPrefillPeak memSource
Qwen3.8-27B
4-bit (MLX)
MLX
Plain autoregressive baseline, 3 runs
6.4–6.5——article
2026-08-29
Qwen3.8-27B
4-bit (MLX)
MLX
DFlash 2 draft model (2B, 4-bit) speculative decoding, block-size 3
block-size 8 drops to 1.11x; an 8-bit draft is slower (1.62x)
11.6–12.21.8–1.9x—19.4GBarticle
2026-08-29
Qwen3.8-27B
UD-Q4_K_S (GGUF, 15.4GB)
llama.cpp b96806d (Metal)
Plain autoregressive baseline (prose / code)
5.9–6—15.4GBarticle
2026-09-02
Qwen3.8-27B
UD-Q4_K_S (GGUF)
llama.cpp b96806d (Metal)
Native MTP head speculative decoding, n-max 2
Prose -24%, code -4%: a net slowdown on Metal
4.5–5.70.76–0.96x—16GBarticle
2026-09-02
Qwen3.8-27B
UD-Q4_K_S (GGUF)
llama.cpp b96806d (Metal)
DFlash 2 GGUF draft (1.14GB), n-max 3
The recommended n-max 7 hits a reproducible Metal OOM at 8K context
3–4.60.5–0.77x——article
2026-09-03
Qwen3.8-27B
UD-Q4_K_S (GGUF)
llama.cpp llama-server b11376
Defaults (--fit caps context at 30K, 4 slots)
6.356—article
2026-10-03
Qwen3.8-27B
UD-Q4_K_S (GGUF, 自导入)
Ollama 0.35.1
Defaults (4096 context)
Ollama 0.35's inference process is its bundled llama-server; same flags, same speed
6.357—article
2026-10-03
Qwen3 4B
Q4_K_M (GGUF)
llama.cpp llama-server b11376
Defaults (--fit opens 101K context, 4 slots)
With 4 concurrent requests: 11.6 tok/s each, 40.9 tok/s aggregate
36.5376—article
2026-10-03
Qwen3 4B
Q4_K_M (GGUF, 官方库 qwen3:4b)
Ollama 0.35.1
Defaults (4096 context, single slot, queued)
Over-long prompts are silently truncated to 2050 tokens; the only trace is one WARN log line
36.3367—article
2026-10-03

Model fit list (recommended quant, 16K context)

Comfortable = ≤ 65% of memory; Tight = ≤ 82% (macOS wires only ~2/3-3/4 of unified memory to the GPU; thresholds calibrated on our 24GB measurements); Lower quant = the recommended quant doesn't fit but a smaller one does. Ceiling = 120GB/s ÷ weight size; single-request decoding cannot exceed it.

ModelQuantNeedsVerdictDecode ceiling tok/s
Qwen3-4BQ4_K_M (GGUF)6.0 GBComfortable≤ 53
Qwen3.5-9BQ4_K_M (GGUF)7.5 GBComfortable≤ 22
Llama 3.1 8BQ4_K_M (GGUF)8.0 GBComfortable≤ 26
Qwen3-14BQ4_K_M (GGUF)12.4 GBComfortable≤ 14
gpt-oss-20BMXFP412.6 GBComfortable≤ 12
Nemotron 3.5 Lightning (30B-A3B)NVFP416.4 GBTight≤ 8.1
Gemma 4 26B-A4BQ4_K_M (GGUF)16.5 GBTight≤ 8.2
Qwen3.8-27BQ4_K_M (GGUF)17.3 GBTight≤ 7.8
Mistral Small 3.x 24BQ4_K_M (GGUF)17.6 GBTight≤ 8.8
GLM-4.7-Flash (30B-A3B)Q3_K_M (GGUF)16.0 GBLower quant≤ 8.8
Qwen3-Coder-30B-A3BQ3_K_M (GGUF)16.4 GBLower quant≤ 9.0
Gemma 4 31BQ3_K_M (GGUF)16.5 GBLower quant≤ 8.8
Qwen3.6-35B-A3BQ3_K_M (GGUF)17.6 GBLower quant≤ 7.6
Qwen3-32BQ3_K_M (GGUF)19.5 GBLower quant≤ 8.6
Gemma 3 27BQ2_K (GGUF)18.1 GBLower quant≤ 14
Llama 3.3 70BQ4_K_M (GGUF)47.9 GBWon't fit—
gpt-oss-120BMXFP463.5 GBWon't fit—
Mistral Leanstral 1.5 (119B-A6B MoE)Q4_K_M (GGUF)73.4 GBWon't fit—
Qwen3.8-Flash-Next (180B MoE)Q4_K_M (GGUF)111 GBWon't fit—
Qwen3-235B-A22BQ4_K_M (GGUF)147 GBWon't fit—
DeepSeek V4 Flash (304B MoE)Q4_K_M (GGUF)187 GBWon't fit—
GLM-4.5 (355B MoE)Q4_K_M (GGUF)224 GBWon't fit—
DeepSeek V3 / R1 (671B MoE)Q4_K_M (GGUF)413 GBWon't fit—
GLM-5.3 (753B MoE)Q4_K_M (GGUF)463 GBWon't fit—
Kimi K2 (1T MoE)Q4_K_M (GGUF)631 GBWon't fit—
DeepSeek V4 Pro (1.6T MoE)Q4_K_M (GGUF)983 GBWon't fit—
Kimi K3 (2.8T MoE)MXFP41.46 TBWon't fit—

Benchmarked a Mac mini M4 24GB yourself?

Send us the model, quant, runtime version, tok/s and peak memory. Verified readings are added to this page with credit. First-hand numbers with a command or log only.

Submit a benchmark

FAQ

What is the largest LLM a Mac mini M4 24GB can run?

At the recommended quantization with 16K context, the ceiling is Nemotron 3.5 Lightning (30B-A3B) (about 16.4 GB, 68% of 24GB). macOS lets the GPU wire only ~2/3-3/4 of unified memory by default; near the limit, raise iogpu.wired_limit_mb and close memory-heavy apps.

Which models run best on a Mac mini M4 24GB?

Models under 16GB leave room for long context without closing other apps, e.g. gpt-oss-20B, Qwen3.5-9B, Qwen3-14B, Llama 3.1 8B.

How many tokens per second does a Mac mini M4 24GB get?

Single-request decoding is bandwidth-bound: ceiling ≈ 120GB/s ÷ weight size. A 15GB 27B 4-bit model tops out around 8.0 tok/s; measured speeds are usually 70-85% of that, and speculative decoding (a draft model) adds another 1.5-2x.

Are these numbers measured or estimated?

The fit table is a planning estimate (same formula as the calculator). The first-hand benchmark table contains readings we took on this exact machine with real model files; every row links to the source article.

Hands-on articles for the Mac mini M4 24GB

Other machines