Ollama Out of Memory on a Mac: How Much num_ctx, Parallelism and KV Quantization Really Cost on 24GB (Tested) — and the Trap It Won't Stop You From
Short answers first
- Start with
ollama ps:100% GPUunderPROCESSORmeans it fits; something like48%/52% CPU/GPUmeans part runs on the CPU.SIZEis the real memory used.- The usual cause is too much context: qwen3:4b uses 3.73GB at 4k context and 9.89GB at 32k. Most of that is KV cache, not the model.
- The most effective fix: set
OLLAMA_FLASH_ATTENTION=1andOLLAMA_KV_CACHE_TYPE=q8_0(orq4_0) before starting Ollama. Measured: 32k context went from 9.89GB to 5.51GB (q8_0) / 4.31GB (q4_0) at the same speed.- Check parallelism:
OLLAMA_NUM_PARALLEL=4sizes memory for 4× the context, but theCONTEXTcolumn inollama psdoesn't show it. Set it to 1 if you don't need concurrency.- Don't count on Ollama to stop you: with a context far larger than your RAM it doesn't say "not enough memory" — it starts allocating and drags the system into swapping. Do the math first.
- On a Mac, "offloaded to CPU" isn't a disaster: qwen3:4b fully on CPU was only 25% slower than fully on GPU.
Background
Running models locally, "out of memory" is almost unavoidable. On a Mac it has two twists:
- Apple Silicon has unified memory with no separate VRAM, but Ollama treats whatever Metal reports as VRAM;
- You often don't get a clean error — you get slowness, stalls, or the whole machine swapping.
So instead of theory, this article measures on a 24GB M4 Mac mini what each setting costs and what happens past the limit.
How Ollama sees this machine
From the startup log:
msg="inference compute" library=Metal name=Metal description="Apple M4" total="17.8 GiB" available="17.8 GiB"
msg="vram-based default context" total_vram="17.8 GiB" default_num_ctx=4096
On a 24GB Mac, Ollama sees 17.8 GiB of VRAM. Per the docs, the default context depends on VRAM: under 24 GiB gets 4k, 24–48 GiB gets 32k, 48 GiB and up gets 256k. So this machine defaults to 4096, while the same page says Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens.
So people raise the context, and that's where out-of-memory starts. Relevant doc points:
- Memory scales with
OLLAMA_NUM_PARALLEL×OLLAMA_CONTEXT_LENGTH(FAQ) - The KV cache can be quantized:
q8_0≈ 1/2 off16,q4_0≈ 1/4, only with Flash Attention on ollama psPROCESSOR:100% GPU/100% CPU/48%/52% CPU/GPU
Method
| Approach | Verdict | Why |
|---|---|---|
| A separate Ollama instance, one setting at a time | ✅ used | An independent instance on 127.0.0.1:11500, reusing the downloaded models read-only, unloading after each test; /api/ps for memory, /api/generate for speed |
| Push a 27B model into the limit | ❌ dropped | Another job using 7GB was running on the machine; a 27B model (~17GB) would have pushed it into swap |
| Run in Docker | ❌ | Per the FAQ, GPU acceleration is not available for Docker Desktop in macOS; results would be CPU-only |
Safeguards:
OLLAMA_NOPRUNE=1:ollama serveprunes model files it considers unused at startup; the test instance shouldn't touch your model store;- A memory guard checking free memory every 2 seconds and killing the test Ollama below 35% (it fired once — see below);
- Model
qwen3:4b(Q4_K_M, 2.50GB file), fixed 128 generated tokens,temperature 0.
Results
1. Context length: it's mostly KV cache
Default settings (f16 KV, no Flash Attention):
num_ctx |
ollama ps size |
PROCESSOR | Generation speed |
|---|---|---|---|
| 4096 (default) | 3.73 GB | 100% GPU | 35.3 / 35.5 tok/s |
| 8192 | 4.61 GB | 100% GPU | 35.5 tok/s |
| 16384 | 6.37 GB | 100% GPU | 35.4 tok/s |
| 32768 | 9.89 GB | 100% GPU | 33.8 tok/s |
A 2.50GB model file takes 9.89GB at 32k context — roughly 0.21GB per extra 1k tokens for qwen3:4b (this varies a lot by model, depending on layer and KV-head count).
At 32k free memory dipped to 37%, close to the guard threshold, so f16 at 64k (~16GB by the slope) wasn't tested.
2. Flash Attention + KV quantization: same 32k, under half the memory
| Setting (32k context) | Size | Generation speed |
|---|---|---|
| Default (f16, no FA) | 9.89 GB | 33.8 tok/s |
OLLAMA_FLASH_ATTENTION=1 |
7.78 GB | 36.3 tok/s |
FA + OLLAMA_KV_CACHE_TYPE=q8_0 |
5.51 GB | 35.4 tok/s |
FA + OLLAMA_KV_CACHE_TYPE=q4_0 |
4.31 GB | 33.6 / 33.5 tok/s |
FA + q4_0, 64k context |
5.77 GB | 36.8 / 36.6 tok/s |
With q4_0, a 64k context uses less memory than the default settings at 32k. Speed stayed between 33.5 and 36.8 tok/s. The server log confirms the settings: FlashAttention:Enabled KvSize:32768 KvCacheType:q4_0.
On quality, the docs say q8_0 usually has no noticeable impact, while q4_0 may be more noticeable at higher context sizes, and models with high GQA counts (the docs cite Qwen2) are affected more. This article measured memory and speed only, not answer quality, so start with q8_0.
3. Parallelism: an invisible 4×
| Setting | ollama ps CONTEXT |
Size |
|---|---|---|
OLLAMA_NUM_PARALLEL=1, num_ctx=8192 |
8192 | 4.61 GB |
OLLAMA_NUM_PARALLEL=4, num_ctx=8192 |
8192 | 9.89 GB |
Four 8k slots use exactly the same memory as one 32k slot (9.89GB), yet ollama ps still says 8192. If you share one Ollama across Open WebUI and several editor plugins and raised the parallelism, this is how memory quietly multiplies.
4. Context far beyond RAM: Ollama won't stop you
I requested qwen3:4b with num_ctx=262144 (f16 KV), expecting Ollama to estimate, see it can't fit, and refuse. The server log:
load request="{Operation:fit … KvSize:262144 … GPULayers:1[ID:0 Layers:1(35..35)] …}"
load request="{Operation:fit … KvSize:262144 … GPULayers:[] …}"
load request="{Operation:alloc … KvSize:262144 … GPULayers:[] …}"
load request="{Operation:alloc … KvSize:262144 … GPULayers:15[ID:0 Layers:15(21..35)] …}"
…
msg="total memory" size="54.6 GiB"
It estimated 54.6 GiB — more than twice the machine's 24GB — and still went into allocation (Operation:alloc): first all on CPU, then 15 layers on GPU. Within 16 seconds free memory dropped from about 75% to 23%, the guard fired and killed the test Ollama. The server then logged Load failed … error="model failed to load, this may be due to resource limitations or an internal error, check ollama server logs for details" — after the kill. What would have happened without the kill wasn't tested (it would have dragged down other services on this machine).
That is the most dangerous form of "out of memory": not a clear error, but the whole machine swapping. Estimate with the slope above before setting a large context.
5. Partial CPU offload: not that slow on a Mac
Controlling GPU layers with num_gpu (qwen3:4b has 37 layers, 4k context):
| GPU layers | ollama ps VRAM share |
Generation speed |
|---|---|---|
| 37 (all) | 100% | 35.4 / 35.5 tok/s |
| 27 | 60% | 30.9 tok/s |
| 18 | 42% | 27.3 tok/s |
| 9 | 25% | 26.3 tok/s |
| 0 (all CPU) | 0% | 26.6 / 26.3 tok/s |
All on CPU was only about 25% slower than all on GPU. That differs from the discrete-GPU experience: on Apple Silicon the CPU and GPU share one memory pool, so layers on either side don't need data shuttled over PCIe between VRAM and RAM (why the gap is exactly ~25% wasn't analyzed further). So mixed CPU/GPU on a Mac isn't "game over"; what to avoid is total demand exceeding physical memory, as in section 4. (This is a 4B model; larger models and long-prompt prefill may differ more — not tested.)
What to do
1. Diagnose:
ollama ps
Look at SIZE (real memory), PROCESSOR (any CPU offload), CONTEXT (per-slot context, not multiplied by parallelism).
2. Usually two variables are enough (on macOS use launchctl setenv or export before ollama serve; restart Ollama afterwards):
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0 # q4_0 if memory is very tight, at some possible quality cost on long contexts
export OLLAMA_NUM_PARALLEL=1 # leave concurrency off unless you need it
3. Raise context as needed, not to the max. Coding tools do want around 64k; a 4B model with q4_0 measured 5.77GB, easily fitting a 24GB Mac. For 7B or 14B, re-estimate proportionally — try our VRAM calculator or what a Mac mini M4 24GB can run.
4. After setting a large context, watch memory pressure (Activity Monitor's Memory Pressure graph, or memory_pressure). Ollama won't stop you; if it goes red, ollama stop <model> right away.
5. On a Mac, CPU offload isn't the end of the world. A small model fully on CPU was only ~25% slower; exceeding physical memory is what to avoid.
Pitfalls along the way
- Assumed Ollama would estimate and refuse. The oversized-context test was planned as "safe" because it would fail at estimation. It went straight into allocation and was stopped only by the memory guard. Lesson: on a shared machine, start the guard before any memory experiment.
ollama serveprunes model files at startup. The log saidtotal unused blobs removed: 0this time, but the test instance was restarted withOLLAMA_NOPRUNE=1to avoid deleting anything from the local model store.- Prompt-processing speed is unusable here. The test prompt was about twenty tokens; the same setting gave 120 and 248 tok/s on two runs, so it isn't cited.
Not verified: layer placement for 27B-class models on a 24GB machine (not loaded because of another memory-heavy job); KV quantization's effect on answer quality; what Ollama ultimately does with an oversized context when not killed; CPU vs GPU speed during long-prompt prefill; per-1k-context memory slopes for models other than qwen3:4b. Apart from the rows showing two values, each setting ran once.