Magic Tools
LLM VRAM Calculator

What LLMs can a Mac mini M4 / MacBook Air M4 16GB run?

At the recommended quantization with 16K context, the largest model a Mac mini M4 / MacBook Air M4 16GB can run is gpt-oss-20B (21B, about 12.6 GB); 3 models on the list run comfortably. With 120GB/s of memory bandwidth, the decode ceiling is roughly bandwidth ÷ weight size; measured speeds usually land at 70-85% of that. We have not benchmarked this machine ourselves yet — submissions welcome. Data verified 2026-10-03.

Specs: 16GB · 120GB/s · Base M4 unified memory at 120GB/s; macOS lets the GPU wire only ~2/3-3/4 of RAM by default

Model fit list (recommended quant, 16K context)

Comfortable = ≤ 65% of memory; Tight = ≤ 82% (macOS wires only ~2/3-3/4 of unified memory to the GPU; thresholds calibrated on our 24GB measurements); Lower quant = the recommended quant doesn't fit but a smaller one does. Ceiling = 120GB/s ÷ weight size; single-request decoding cannot exceed it.

ModelQuantNeedsVerdictDecode ceiling tok/s
Qwen3-4BQ4_K_M (GGUF)6.0 GBComfortable≤ 53
Qwen3.5-9BQ4_K_M (GGUF)7.5 GBComfortable≤ 22
Llama 3.1 8BQ4_K_M (GGUF)8.0 GBComfortable≤ 26
Qwen3-14BQ4_K_M (GGUF)12.4 GBTight≤ 14
gpt-oss-20BMXFP412.6 GBTight≤ 12
Nemotron 3.5 Lightning (30B-A3B)Q2_K (GGUF)11.4 GBLower quant≤ 12
Gemma 4 26B-A4BQ3_K_M (GGUF)13.1 GBLower quant≤ 11
Qwen3.8-27BQ2_K (GGUF)10.8 GBLower quant≤ 14
Mistral Small 3.x 24BQ2_K (GGUF)11.8 GBLower quant≤ 15
GLM-4.7-Flash (30B-A3B)Q2_K (GGUF)12.5 GBLower quant≤ 12
Qwen3-Coder-30B-A3BQ2_K (GGUF)12.9 GBLower quant≤ 12
Gemma 4 31BQ2_K (GGUF)13.0 GBLower quant≤ 12
Qwen3.6-35B-A3BQ4_K_M (GGUF)22.4 GBWon't fit—
Qwen3-32BQ4_K_M (GGUF)23.7 GBWon't fit—
Gemma 3 27BQ4_K_M (GGUF)24.6 GBWon't fit—
Llama 3.3 70BQ4_K_M (GGUF)47.9 GBWon't fit—
gpt-oss-120BMXFP463.5 GBWon't fit—
Mistral Leanstral 1.5 (119B-A6B MoE)Q4_K_M (GGUF)73.4 GBWon't fit—
Qwen3.8-Flash-Next (180B MoE)Q4_K_M (GGUF)111 GBWon't fit—
Qwen3-235B-A22BQ4_K_M (GGUF)147 GBWon't fit—
DeepSeek V4 Flash (304B MoE)Q4_K_M (GGUF)187 GBWon't fit—
GLM-4.5 (355B MoE)Q4_K_M (GGUF)224 GBWon't fit—
DeepSeek V3 / R1 (671B MoE)Q4_K_M (GGUF)413 GBWon't fit—
GLM-5.3 (753B MoE)Q4_K_M (GGUF)463 GBWon't fit—
Kimi K2 (1T MoE)Q4_K_M (GGUF)631 GBWon't fit—
DeepSeek V4 Pro (1.6T MoE)Q4_K_M (GGUF)983 GBWon't fit—
Kimi K3 (2.8T MoE)MXFP41.46 TBWon't fit—

Benchmarked a Mac mini M4 / MacBook Air M4 16GB yourself?

Send us the model, quant, runtime version, tok/s and peak memory. Verified readings are added to this page with credit. First-hand numbers with a command or log only.

Submit a benchmark

FAQ

What is the largest LLM a Mac mini M4 / MacBook Air M4 16GB can run?

At the recommended quantization with 16K context, the ceiling is gpt-oss-20B (about 12.6 GB, 79% of 16GB). macOS lets the GPU wire only ~2/3-3/4 of unified memory by default; near the limit, raise iogpu.wired_limit_mb and close memory-heavy apps.

Which models run best on a Mac mini M4 / MacBook Air M4 16GB?

Models under 10GB leave room for long context without closing other apps, e.g. Qwen3.5-9B, Llama 3.1 8B, Qwen3-4B.

How many tokens per second does a Mac mini M4 / MacBook Air M4 16GB get?

Single-request decoding is bandwidth-bound: ceiling ≈ 120GB/s ÷ weight size. A 15GB 27B 4-bit model tops out around 8.0 tok/s; measured speeds are usually 70-85% of that, and speculative decoding (a draft model) adds another 1.5-2x.

Are these numbers measured or estimated?

The fit table is a planning estimate (same formula as the calculator). The first-hand benchmark table contains readings we took on this exact machine with real model files; every row links to the source article.

Other machines