Magic Tools
Pitfall NotesBy CooconOctober 5, 20264 views7 min read

Ollama Out of Memory on a Mac: How Much num_ctx, Parallelism and KV Quantization Really Cost on 24GB (Tested) — and the Trap It Won't Stop You From

Short answers first

  • Start with ollama ps: 100% GPU under PROCESSOR means it fits; something like 48%/52% CPU/GPU means part runs on the CPU. SIZE is the real memory used.
  • The usual cause is too much context: qwen3:4b uses 3.73GB at 4k context and 9.89GB at 32k. Most of that is KV cache, not the model.
  • The most effective fix: set OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 (or q4_0) before starting Ollama. Measured: 32k context went from 9.89GB to 5.51GB (q8_0) / 4.31GB (q4_0) at the same speed.
  • Check parallelism: OLLAMA_NUM_PARALLEL=4 sizes memory for 4× the context, but the CONTEXT column in ollama ps doesn't show it. Set it to 1 if you don't need concurrency.
  • Don't count on Ollama to stop you: with a context far larger than your RAM it doesn't say "not enough memory" — it starts allocating and drags the system into swapping. Do the math first.
  • On a Mac, "offloaded to CPU" isn't a disaster: qwen3:4b fully on CPU was only 25% slower than fully on GPU.

Background

Running models locally, "out of memory" is almost unavoidable. On a Mac it has two twists:

  1. Apple Silicon has unified memory with no separate VRAM, but Ollama treats whatever Metal reports as VRAM;
  2. You often don't get a clean error — you get slowness, stalls, or the whole machine swapping.

So instead of theory, this article measures on a 24GB M4 Mac mini what each setting costs and what happens past the limit.

How Ollama sees this machine

From the startup log:

msg="inference compute" library=Metal name=Metal description="Apple M4" total="17.8 GiB" available="17.8 GiB"
msg="vram-based default context" total_vram="17.8 GiB" default_num_ctx=4096

On a 24GB Mac, Ollama sees 17.8 GiB of VRAM. Per the docs, the default context depends on VRAM: under 24 GiB gets 4k, 24–48 GiB gets 32k, 48 GiB and up gets 256k. So this machine defaults to 4096, while the same page says Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens.

So people raise the context, and that's where out-of-memory starts. Relevant doc points:

  • Memory scales with OLLAMA_NUM_PARALLEL × OLLAMA_CONTEXT_LENGTH (FAQ)
  • The KV cache can be quantized: q8_0 ≈ 1/2 of f16, q4_0 ≈ 1/4, only with Flash Attention on
  • ollama ps PROCESSOR: 100% GPU / 100% CPU / 48%/52% CPU/GPU

Method

Approach Verdict Why
A separate Ollama instance, one setting at a time ✅ used An independent instance on 127.0.0.1:11500, reusing the downloaded models read-only, unloading after each test; /api/ps for memory, /api/generate for speed
Push a 27B model into the limit ❌ dropped Another job using 7GB was running on the machine; a 27B model (~17GB) would have pushed it into swap
Run in Docker ❌ Per the FAQ, GPU acceleration is not available for Docker Desktop in macOS; results would be CPU-only

Safeguards:

  • OLLAMA_NOPRUNE=1: ollama serve prunes model files it considers unused at startup; the test instance shouldn't touch your model store;
  • A memory guard checking free memory every 2 seconds and killing the test Ollama below 35% (it fired once — see below);
  • Model qwen3:4b (Q4_K_M, 2.50GB file), fixed 128 generated tokens, temperature 0.

Results

1. Context length: it's mostly KV cache

Default settings (f16 KV, no Flash Attention):

num_ctx ollama ps size PROCESSOR Generation speed
4096 (default) 3.73 GB 100% GPU 35.3 / 35.5 tok/s
8192 4.61 GB 100% GPU 35.5 tok/s
16384 6.37 GB 100% GPU 35.4 tok/s
32768 9.89 GB 100% GPU 33.8 tok/s

A 2.50GB model file takes 9.89GB at 32k context — roughly 0.21GB per extra 1k tokens for qwen3:4b (this varies a lot by model, depending on layer and KV-head count).

At 32k free memory dipped to 37%, close to the guard threshold, so f16 at 64k (~16GB by the slope) wasn't tested.

2. Flash Attention + KV quantization: same 32k, under half the memory

Setting (32k context) Size Generation speed
Default (f16, no FA) 9.89 GB 33.8 tok/s
OLLAMA_FLASH_ATTENTION=1 7.78 GB 36.3 tok/s
FA + OLLAMA_KV_CACHE_TYPE=q8_0 5.51 GB 35.4 tok/s
FA + OLLAMA_KV_CACHE_TYPE=q4_0 4.31 GB 33.6 / 33.5 tok/s
FA + q4_0, 64k context 5.77 GB 36.8 / 36.6 tok/s

With q4_0, a 64k context uses less memory than the default settings at 32k. Speed stayed between 33.5 and 36.8 tok/s. The server log confirms the settings: FlashAttention:Enabled KvSize:32768 KvCacheType:q4_0.

On quality, the docs say q8_0 usually has no noticeable impact, while q4_0 may be more noticeable at higher context sizes, and models with high GQA counts (the docs cite Qwen2) are affected more. This article measured memory and speed only, not answer quality, so start with q8_0.

3. Parallelism: an invisible 4×

Setting ollama ps CONTEXT Size
OLLAMA_NUM_PARALLEL=1, num_ctx=8192 8192 4.61 GB
OLLAMA_NUM_PARALLEL=4, num_ctx=8192 8192 9.89 GB

Four 8k slots use exactly the same memory as one 32k slot (9.89GB), yet ollama ps still says 8192. If you share one Ollama across Open WebUI and several editor plugins and raised the parallelism, this is how memory quietly multiplies.

4. Context far beyond RAM: Ollama won't stop you

I requested qwen3:4b with num_ctx=262144 (f16 KV), expecting Ollama to estimate, see it can't fit, and refuse. The server log:

load request="{Operation:fit … KvSize:262144 … GPULayers:1[ID:0 Layers:1(35..35)] …}"
load request="{Operation:fit … KvSize:262144 … GPULayers:[] …}"
load request="{Operation:alloc … KvSize:262144 … GPULayers:[] …}"
load request="{Operation:alloc … KvSize:262144 … GPULayers:15[ID:0 Layers:15(21..35)] …}"
…
msg="total memory" size="54.6 GiB"

It estimated 54.6 GiB — more than twice the machine's 24GB — and still went into allocation (Operation:alloc): first all on CPU, then 15 layers on GPU. Within 16 seconds free memory dropped from about 75% to 23%, the guard fired and killed the test Ollama. The server then logged Load failed … error="model failed to load, this may be due to resource limitations or an internal error, check ollama server logs for details" — after the kill. What would have happened without the kill wasn't tested (it would have dragged down other services on this machine).

That is the most dangerous form of "out of memory": not a clear error, but the whole machine swapping. Estimate with the slope above before setting a large context.

5. Partial CPU offload: not that slow on a Mac

Controlling GPU layers with num_gpu (qwen3:4b has 37 layers, 4k context):

GPU layers ollama ps VRAM share Generation speed
37 (all) 100% 35.4 / 35.5 tok/s
27 60% 30.9 tok/s
18 42% 27.3 tok/s
9 25% 26.3 tok/s
0 (all CPU) 0% 26.6 / 26.3 tok/s

All on CPU was only about 25% slower than all on GPU. That differs from the discrete-GPU experience: on Apple Silicon the CPU and GPU share one memory pool, so layers on either side don't need data shuttled over PCIe between VRAM and RAM (why the gap is exactly ~25% wasn't analyzed further). So mixed CPU/GPU on a Mac isn't "game over"; what to avoid is total demand exceeding physical memory, as in section 4. (This is a 4B model; larger models and long-prompt prefill may differ more — not tested.)

What to do

1. Diagnose:

ollama ps

Look at SIZE (real memory), PROCESSOR (any CPU offload), CONTEXT (per-slot context, not multiplied by parallelism).

2. Usually two variables are enough (on macOS use launchctl setenv or export before ollama serve; restart Ollama afterwards):

export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0     # q4_0 if memory is very tight, at some possible quality cost on long contexts
export OLLAMA_NUM_PARALLEL=1         # leave concurrency off unless you need it

3. Raise context as needed, not to the max. Coding tools do want around 64k; a 4B model with q4_0 measured 5.77GB, easily fitting a 24GB Mac. For 7B or 14B, re-estimate proportionally — try our VRAM calculator or what a Mac mini M4 24GB can run.

4. After setting a large context, watch memory pressure (Activity Monitor's Memory Pressure graph, or memory_pressure). Ollama won't stop you; if it goes red, ollama stop <model> right away.

5. On a Mac, CPU offload isn't the end of the world. A small model fully on CPU was only ~25% slower; exceeding physical memory is what to avoid.

Pitfalls along the way

  • Assumed Ollama would estimate and refuse. The oversized-context test was planned as "safe" because it would fail at estimation. It went straight into allocation and was stopped only by the memory guard. Lesson: on a shared machine, start the guard before any memory experiment.
  • ollama serve prunes model files at startup. The log said total unused blobs removed: 0 this time, but the test instance was restarted with OLLAMA_NOPRUNE=1 to avoid deleting anything from the local model store.
  • Prompt-processing speed is unusable here. The test prompt was about twenty tokens; the same setting gave 120 and 248 tok/s on two runs, so it isn't cited.

Not verified: layer placement for 27B-class models on a 24GB machine (not loaded because of another memory-heavy job); KV quantization's effect on answer quality; what Ollama ultimately does with an oversized context when not killed; CPU vs GPU speed during long-prompt prefill; per-1k-context memory slopes for models other than qwen3:4b. Apart from the rows showing two values, each setting ran once.

Get field notes like this every Saturday

Subscribe to Dev Breakfast: daily AI coding picks at 8:00, plus a Saturday roundup of this week's hands-on tests with Claude Code / Codex / local models. Written in Chinese.

Related Articles

LM Studio Slow on a Mac? Tested: the 39-Second Wait Is Prompt Prefill, Not GPU Offload

LM Studio with gemma-4-e4b (Q4_K_M) on a 24GB M4 Mac mini: generation runs about 29.5 tok/s, and turning GPU offload from max to off costs only 11%. What really feels slow is prompt prefill — a 13K-token prompt waited 39.4 seconds for the first token (about 330 tokens/s), and 3.4–3.7x longer on CPU. A second request with the same prefix got its first token in 0.1 s thanks to prefix caching. With --parallel 4 and four simultaneous requests, each dropped to 9.7 tok/s.

local-llmmac-mini+5
pitfallsOct 5, 20267 min
4
llama.cpp's llama-server or Ollama? Same GGUF on a 24GB Mac mini, Tested: Ollama Runs llama-server Under the Hood

llama.cpp's llama-server or Ollama? Same GGUF on a 24GB Mac mini, Tested: Ollama Runs llama-server Under the Hood

llama-server (b11376) and Ollama (0.35.1) loading the same GGUF file on a 24GB Mac mini. Ollama 0.35's inference process is its own bundled llama-server, and with the same settings speed is identical (4B decode 36.5 vs 36.3 tok/s, 27B 6.3 on both). Every difference comes from defaults: Ollama defaults to a 4096-token context and silently cuts an over-long prompt for library models down to 2050 tokens, leaving only a WARN line in its log; llama-server sizes context to fill memory (101,120 for the 4B, 14GB+ footprint) and returns 400 when a prompt doesn't fit. llama-server runs 4 parallel slots by default while Ollama queues, and Ollama with 32k context plus 4-way parallelism allocates 32k per slot, a 21GB footprint. Recommended flags for 24GB machines at the end.

benchmarkqwen+6
hands-onOct 3, 202613 min
35
Same 24GB Mac mini, Same DFlash 2: 1.9x Faster on MLX, 50% Slower (and OOM) on llama.cpp

Same 24GB Mac mini, Same DFlash 2: 1.9x Faster on MLX, 50% Slower (and OOM) on llama.cpp

Part three of my speculative-decoding trilogy on a base Mac mini M4. llama.cpp merged DFlash 2 support with official GGUF drafts — and every configuration is a net slowdown. The README-recommended n-max 7 hits a reproducible Metal OOM on 24GB; the only stable setting cuts prose from 6.0 to 3.0 tok/s, and an 83.8% acceptance rate on code still loses 23%. Same algorithm, same machine, MLX gets 1.8–1.9x. The arithmetic shows why: 0.77s per speculative step loses even at 100% acceptance.

qwendflash+6
ai-tutorialsSep 3, 20268 min
280
Qwen3.8 Native MTP on a 24GB Mac mini: 3GB Less Memory, 24% Slower — Head-to-Head With DFlash 2

Qwen3.8 Native MTP on a 24GB Mac mini: 3GB Less Memory, 24% Slower — Head-to-Head With DFlash 2

Qwen3.8 ships a trained multi-token-prediction head, and llama.cpp can mount it with one flag — no separate 2B draft model. I benchmarked it against DFlash 2 on the same 24GB Mac mini M4: memory does drop (16.0GB vs 19.4GB peak), but speed goes backwards — prose falls from 6.0 to 4.5 tok/s (-24%) while DFlash 2 delivers 1.8–1.9x on the same machine. Draft acceptance is healthy (59–85%); the loss is in Metal's verify path — batch-8 decode amortizes at just 1.13x, measured.

qwenspeculative-decoding+6
ai-tutorialsSep 2, 20268 min
324

Published by Magic Tools