Magic Tools
Hands-OnBy CooconOctober 3, 202622 views13 min read

llama.cpp's llama-server or Ollama? Same GGUF on a 24GB Mac mini, Tested: Ollama Runs llama-server Under the Hood

The problem

Running local models on a Mac comes down to a choice early on: use llama.cpp's llama-server directly, or install Ollama?

Most comparisons stop at "Ollama is simpler, llama.cpp is more flexible". Once you actually deploy, what decides the outcome is usually not which one is faster but a handful of defaults: how long the context is, what happens when a prompt is too long, how many requests run at once, how much memory gets used. When those defaults differ, the same model behaves completely differently on the two.

So on a Mac mini M4 with 24GB, I had both load the same GGUF file, with the same script and the same API (OpenAI-compatible /v1/chat/completions, streaming), and compared them item by item: deployment, API, speed, default context, long prompts, concurrency, memory, and how a 27B model behaves on a 24GB machine.

Analysis

One look at the process list before benchmarking already gave the main answer: Ollama 0.35.1's inference process is the llama-server shipped inside its own package.

The child process of ollama serve is bin/ollama/llama-server, launched with -c 4096 -np 1 Figure: process list captured by the benchmark script while Ollama was running (paths shortened). Top: qwen3:4b from Ollama's library. Bottom: a Qwen3.8 GGUF I imported with ollama create. (Chinese comments: "Ollama 0.35.1 running qwen3:4b (library model)" / "the same Ollama running the imported Qwen3.8 GGUF".)

The unpacked Ollama directory contains a llama-server (--version reports 0.5.0-dev (build 1, commit 6f767fe96)) plus MLX files such as mlx_metal_v3/v4. When a GGUF model is loaded, ollama serve launches this llama-server as its runner with these key flags:

Flag Ollama passes Meaning
-c 4096 -np 1 4096-token context, a single slot
--no-jinja --chat-template chatml Ollama applies the template itself for library models (absent for imported GGUFs)
--context-shift --keep 4 when generation fills the context, drop old tokens and keep going
-b 512 -ub 512 logical batch 512; llama-server's default is -b 2048 -ub 512
--no-webui turn off llama-server's built-in web UI

So the real question isn't "which engine is faster". It's which flags Ollama fills in for the same llama-server, and what management layer it wraps around it. Almost everything below confirms this.

Setup

  • Versions: llama.cpp official release b11376 (2026-10-03, macos-arm64 prebuilt) and Ollama official release v0.35.1 (2026-09-29, ollama-darwin.tgz). Both ran from a temp directory; the brew-installed Ollama 0.19 on the machine was left alone.
  • Models, the same file on both sides:
    • 4B: the GGUF blob behind Ollama's library qwen3:4b (2.5GB, Q4_K_M, metadata name Qwen3 4B Thinking 2507). llama-server loads the blob directly with -m (hard-linked, no extra disk).
    • 27B: the Qwen3.8-27B-UD-Q4_K_S.gguf already on this machine (15.4GB). Ollama imports it with ollama create; llama-server loads it directly.
  • Method (Python script, standard library only): each case starts one server, then measures cold start, sequential decoding (a writing prompt, max_tokens 300 / 200 for 27B, 5 / 3 runs), prefill (a ~2,500-token on-call log, 5 / 3 runs), needle-in-a-haystack (one line at the top says "the code word is 青鸟-7421", followed by filler, with the question at the end; ~15.7k tokens for 4B, ~7.6k for 27B), concurrency (4 requests at once for 4B, 2 for 27B), and memory (macOS footprint). A random UUID prefixes every prompt to avoid prefix-cache hits. temperature 0.
  • How numbers are computed: decode speed = (completion_tokens − 1) ÷ (time of last streamed chunk − time of first chunk); prefill speed = prompt_tokens ÷ time to first token, measured client-side including HTTP overhead. Memory is footprint, which excludes the mmapped weight file and reflects extra memory such as KV cache and compute buffers; for Ollama it is serve + runner combined.

Experiments

1. Deployment: one command each, the difference is in defaults

# llama.cpp: unpack the release and run; it starts listening once the model is loaded
./llama-server -m Qwen3.8-27B-UD-Q4_K_S.gguf --host 127.0.0.1 --port 8080

# Ollama: start the server; the model loads on the first request
ollama serve                         # listens on 127.0.0.1:11434 by default
ollama create qwen38-ud -f Modelfile # the Modelfile is one line: FROM /path/to/Qwen3.8-27B-UD-Q4_K_S.gguf

Two observations:

  • Importing a GGUF into Ollama copies it. ollama create took 17.5 s and copied the 15GB file into Ollama's own model directory, so the model now takes up disk twice. After import, ollama show correctly listed the tools and thinking capabilities, so it read the template from the GGUF.

  • The other direction doesn't always work. Loading the blob of Ollama's library qwen3.5:27b into upstream llama-server fails immediately:

    error loading model hyperparameters: key qwen35.rope.dimension_sections has wrong array length; expected 4, got 3
    

    The qwen3:4b blob loads fine. Ollama's bundled llama-server is not the same commit as upstream and their GGUF metadata conventions differ, so "reuse the model Ollama already downloaded in llama.cpp" won't always work.

Startup: with the model file already in the page cache, llama-server with the 4B and -c 4096 was ready in 0.6 s, and about 3 s with defaults (it first allocates 14GB of KV cache). Ollama's server answers in 0.2–0.6 s but loads the model on the first request, so the 4B's first token took 1.2–1.4 s. When the 27B's 15GB file wasn't cached, the first load took 14–16 s on either side. Also, a freshly downloaded binary is much slower on its first run (llama-server 17 s, Ollama 9.5 s; the second runs took 3.8 s and 0.3 s). That is presumably macOS checking new binaries on first launch, nothing to do with the engines, so drop the first run when measuring startup.

2. API: both OpenAI-compatible, with different details

Both serve /v1/chat/completions, so a client switches between them by changing the base URL. My script sent identical request bodies to both (stream: true + stream_options.include_usage), and both returned usage. Differences:

  • llama-server names the model with --alias, and requests use that name (m in these runs). The final streamed chunk carries timings (server-side prompt and generation speed), which is handy for benchmarking.
  • Ollama's model is a name from its model list (such as qwen3:4b), and the server loads the matching model by name. The OpenAI-compatible streaming response has no timings; for server-side timing use the native /api/chat endpoint.
  • Ollama unloads a model after 5 idle minutes by default (official FAQ: By default models are kept in memory for 5 minutes), so the next request has to reload it. llama-server keeps the model in memory for as long as the process runs.

3. Speed: identical with the same flags

Summary of speed, context, needle test, concurrency and memory for the same GGUF on both Figure: summary of all cases. Decode is the median of 5 / 3 runs. For the two 4B-default rows, the needle and memory columns come from the dedicated long-prompt cases (B1/B2). Column headers: config, engine, context, decode, prefill (client-timed), needle (找到 found / 截断 truncated / 400), concurrency (并行 parallel / 排队 queued), footprint.

Model Decode tok/s (llama-server / Ollama) Prefill tok/s, client-timed
4B, each with defaults 36.5 / 36.3 376 / 367
27B, each with defaults 6.3 / 6.3 56 / 57
27B, both at 16k, one slot 6.3 / 6.3 57 / 57

The 27B's 6.3 tok/s matches the 6.0–6.5 tok/s plain autoregressive baseline this site measured earlier for Qwen3.8-27B. Ollama's 4B prefill is about 2% slower. When I ran llama-server with Ollama's -c 4096 -np 1, its own timings showed 384–389 tok/s, still slightly above Ollama. The remaining gap may come from the -b 512 Ollama passes (llama-server defaults to -b 2048); I didn't test that separately, so it's a guess. Either way the gap is at the noise floor: you won't regret either choice on speed.

4. Default context: 4096 versus "as large as memory allows"

This is by far the biggest difference.

  • Ollama defaults to 4096. The official FAQ says By default, Ollama uses a context window size of 4096 tokens, and the server log has vram-based default context total_vram="17.8 GiB" default_num_ctx=4096, so the default is picked by VRAM tier and this machine lands in the 4096 tier.
  • llama-server defaults to -c 0, the model's own context length, which the default-on --fit then shrinks to fit memory. This 4B model declares 262,144; --fit brought it down to 101,120, shared by 4 slots. The 27B came out at 30,208.

So the same ~15.7k-token prompt gets completely different treatment:

Ollama silently truncates to 2050 tokens; llama-server with the same flags returns 400; Ollama passes the 400 through for an imported model Figure: three ways of handling the same 4096 context. The yellow line is the only hint in Ollama's server log. (Chinese headings: "same 15,742-token prompt, context 4096 everywhere"; "Ollama + qwen3:4b (library model): HTTP 200, usage shows only 2050 tokens went in"; "llama-server -c 4096 -np 1 (same flags as Ollama): rejected outright"; "Ollama + imported Qwen3.8 GGUF (7,596 tokens): no truncation, the 400 is passed through".)

  • Ollama + library model: HTTP 200 with usage.prompt_tokens of just 2050. The code word at the top was cut off; the model produced 1,333 tokens without giving it. The server log has a single line, level=WARN msg="truncating input prompt" limit=2050 prompt=15742 keep=4 new=2050, and the API caller gets no signal at all.
  • llama-server with the same -c 4096 -np 1: an immediate 400, request (15742 tokens) exceeds the available context size (4096 tokens), try increasing it.
  • Ollama + an imported GGUF (27B, 7,596 tokens): no truncation, and the 400 is passed through to the caller. The reason is in the flags table above: for library models Ollama applies the template itself (--no-jinja --chat-template chatml) and also truncates at that layer; imported models are handed to llama-server as-is, so llama-server's error path applies.

Once the context is large enough, both find the needle: llama-server defaults, llama-server -c 32768, Ollama OLLAMA_CONTEXT_LENGTH=32768, and both sides at 16k for the 27B all answered 青鸟-7421 correctly.

On the same Ollama, changing where the model came from flips over-long prompts from "silently truncated" to "error". For RAG, codebase Q&A or long-document summaries on Ollama, the silent case is the dangerous one: the answer looks normal, but the model only saw half the input.

5. Concurrency: 4 parallel slots by default vs a queue by default

Four writing requests at once (4B, 200 tokens each):

First token per request (s) Decode per request, tok/s Total throughput, tok/s
llama-server defaults (4 slots) 2.36 / 2.37 / 2.37 / 2.37 11.6 40.9
Ollama defaults (1 slot) 0.27 / 6.0 / 11.7 / 17.4 36.3 34.9
llama-server -c 32768 (4 slots sharing it) 1.02 × 4 11.6 44.0
Ollama OLLAMA_CONTEXT_LENGTH=32768 + OLLAMA_NUM_PARALLEL=4 7.53 × 4 11.8 32.7

On the 4B, llama-server's parallel slots buy about 17% more total throughput, at the cost of every request running at a third of the speed. Ollama queues by default, so the fourth request waits 17 s for its first token. On the 27B, parallelism pays off even less: 6.86 vs 5.9 tok/s total for 2 requests, with each dropping to 3.8 tok/s.

If you're one person sending one request at a time, queueing actually gives each request the best speed; parallel slots only matter when several agents or people share the server.

6. Memory: the same "32k + 4 parallel" setting differs 3x

Config (4B) Flags actually passed to llama-server footprint
Ollama defaults -c 4096 -np 1 0.7 GB
Ollama OLLAMA_CONTEXT_LENGTH=32768 -c 32768 -np 1 4.7 GB
Ollama 32k + OLLAMA_NUM_PARALLEL=4 -c 131072 -np 4 21.0 GB
llama-server -c 32768 (4 slots, auto) -c 32768, kv_unified = 'true' 6.9 GB
llama-server defaults -c 0 → fit to 101120, 4 slots 14.1 GB (after the long-prompt case only) / 18.3 GB (after the full suite)

Two traps:

  • Ollama's OLLAMA_CONTEXT_LENGTH is per request slot, multiplied by the number of parallel slots (official FAQ: Required RAM will scale by OLLAMA_NUM_PARALLEL * OLLAMA_CONTEXT_LENGTH). llama-server's -c is the total, shared across slots with a unified KV cache (kv_unified). The same "32k, 4 parallel" means 32k for each request on Ollama and 32k for all four together on llama-server. Different semantics, 3x the memory.
  • llama-server's defaults take most of a 24GB machine's memory, even for a 2.5GB model. Two reasons: --fit grows the KV cache until it nearly reaches the memory limit (default margin just 1024 MiB), and the default --cache-ram 8192 keeps up to 8GB of old prompt state in RAM, growing as more requests come in. The same default config measured 14.1GB after the long-prompt case and 18.3GB after the full suite.

For the 27B: llama-server defaults (30,208 context) 7.9GB, Ollama defaults 5.0GB, and 6.8GB on both at 16k. Qwen3.8 uses hybrid attention with only 1 full-attention layer in every 4, so its KV cache is much smaller than a full-attention model like the 4B, and the 27B's extra memory ends up lower than the 4B's defaults.

One counterexample: Ollama's library qwen3.5:27b (/api/ps reports 17.2GB, and the runner also loads a vision module via --mmproj) shows size_vram of 15.6GB on this machine, meaning about 1.6GB sits on the CPU (decimal GB from /api/ps byte counts). It decoded at 5.8 tok/s with a 21.1GB footprint. A 17GB-class model is at the edge of what a 24GB machine can hold; the 27B UD-Q4_K_S (15.4GB) fits entirely on the GPU.

7. Built-in web UI: llama-server has one, Ollama turns it off

llama-server ships with a web UI enabled by default. Open http://127.0.0.1:8080 in a browser to chat, with token counts and speed shown:

llama-server's built-in web UI showing 856 tokens, 24s, 35.03 t/s under the reply Figure: llama-server b11376's built-in web UI with the 4B (-c 8192 -np 1). The screenshot stops mid-reasoning; this 4B model gets the llama-server/Ollama relationship wrong while thinking, so treat it as a UI demo only. The prompt asks, in Chinese, for a one-sentence explanation of how llama-server and Ollama relate.

Ollama explicitly passes --no-webui when it starts the runner, so the command-line version has no web UI (the macOS desktop app has its own chat window, which I didn't test). For a web front end on Ollama you'd add something like Open WebUI.

Results

Which one to pick

Your situation Pick
Want it to just work, pull models from the library, switch models often Ollama, but raise OLLAMA_CONTEXT_LENGTH first
Manage your own GGUF files, want exact control over context / concurrency / memory, or want a built-in web UI llama-server
Backend for Claude Code-style tools, RAG, long-document summaries Either works if you set the context explicitly; watch out for silent truncation with Ollama library models
Several agents or people calling it at once llama-server has parallel slots by default; on Ollama set OLLAMA_NUM_PARALLEL and budget memory as per-slot length × slots
# llama-server: set context and slots explicitly so --fit and the prompt cache don't eat your memory
./llama-server -m model.gguf -c 16384 -np 1 --cache-ram 2048 \
  --host 127.0.0.1 --port 8080

# Ollama: set the server-wide default context (when started from the command line)
OLLAMA_CONTEXT_LENGTH=16384 ollama serve

With the 27B at -c 16384 -np 1, both sides showed a 6.8GB footprint. Add roughly 14.3GB of mapped weights (the 15.4GB file) and the total is close to 21GB. That doesn't leave much room on a 24GB machine, so close memory-hungry apps first.

I didn't test --cache-ram 2048 separately; it's a conservative value based on "the 8GB default is too much for 24GB". 0 disables the prompt cache entirely (official help: 0 - disable). The Ollama desktop app doesn't go through your shell, so per the official FAQ set environment variables with launchctl setenv and restart the app. A single request can also override the context with options.num_ctx on the native API (as shown in the official FAQ).

Security notes

  • llama-server prints security: no API key is set and CORS allows all origins at startup. That's fine while it only listens on 127.0.0.1; if you switch to --host 0.0.0.0 to expose it on your LAN, add --api-key.
  • Ollama binds 127.0.0.1:11434 by default (official FAQ). After setting OLLAMA_HOST=0.0.0.0 as many tutorials suggest, anyone on the LAN can call it, since it has no authentication of its own.

Gotchas

  • Ollama's silent truncation only happens with library models. On the same Ollama, an imported GGUF returns 400 when the prompt is too long, while a library model returns 200 with half the prompt. When a model "didn't see" earlier content, check usage.prompt_tokens first, then look for truncating input prompt in Ollama's server log.
  • OLLAMA_NUM_PARALLEL multiplies memory by the slot count. 32k with 4 slots on the 4B is 21GB; with a 27B, a 24GB machine can't possibly hold it.
  • llama-server's defaults fill your memory. 14.1–18.3GB footprint for a 4B model. On a Mac you also use for other things, set -c and --cache-ram explicitly.
  • Ollama blobs aren't guaranteed to work with upstream llama.cpp. qwen3.5:27b failed with a rope.dimension_sections length mismatch. To share a file between the two, import an upstream GGUF into Ollama.
  • This qwen3:4b is the Thinking variant, and /no_think does nothing. My first long-prompt run allowed only 32 output tokens, all spent on reasoning, so neither side answered and it almost looked like both had truncated. Raising it to 2048 showed what really happened.
  • Mistakes in the experiment itself: in one case llama-server returned 400 on the long prompt, the benchmark script didn't catch the exception and exited, and the llama-server it left behind became an orphan holding port 8080 and over ten GB of memory. The next Ollama case ran with that memory taken, and the readiness check of the case after that connected straight to the orphan. All three cases were discarded (raw records kept); the script now cleans up in try/finally and checks the port before starting, and the cases were rerun. Also, footprint excludes mmapped weights and llama-server's prompt cache grows as the tests proceed, so memory has to be compared at the same stage; the tables above say which stage.

Get field notes like this every Saturday

Subscribe to Dev Breakfast: daily AI coding picks at 8:00, plus a Saturday roundup of this week's hands-on tests with Claude Code / Codex / local models. Written in Chinese.

Related Articles

MiniMax T2A v2 vs Azure Neural TTS, Benchmarked: 6–10× Latency Gap on the Same Text

MiniMax T2A v2 vs Azure Neural TTS, Benchmarked: 6–10× Latency Gap on the Same Text

My video factory wires up both MiniMax and Azure TTS. This benchmark measures them head to head: same Chinese + English text, 5 rounds each, unified 24kHz/mono/16bit output. MiniMax median latency 1.0–2.2s, RTF 0.08–0.18 (5–12× faster than real time); Azure median 6–21s, RTF near 1.0, tail jittering to 27s. Two causes, both with evidence: Azure Neural synthesizes at roughly real-time pace, and its eastasia endpoint is a trans-Pacific hop from mainland China (TLS jitter to 0.8s).

ttsrtf+6
hands-onSep 5, 20269 min
161
Same 24GB Mac mini, Same DFlash 2: 1.9x Faster on MLX, 50% Slower (and OOM) on llama.cpp

Same 24GB Mac mini, Same DFlash 2: 1.9x Faster on MLX, 50% Slower (and OOM) on llama.cpp

Part three of my speculative-decoding trilogy on a base Mac mini M4. llama.cpp merged DFlash 2 support with official GGUF drafts — and every configuration is a net slowdown. The README-recommended n-max 7 hits a reproducible Metal OOM on 24GB; the only stable setting cuts prose from 6.0 to 3.0 tok/s, and an 83.8% acceptance rate on code still loses 23%. Same algorithm, same machine, MLX gets 1.8–1.9x. The arithmetic shows why: 0.77s per speculative step loses even at 100% acceptance.

qwendflash+6
ai-tutorialsSep 3, 20268 min
274
Qwen3.8 Native MTP on a 24GB Mac mini: 3GB Less Memory, 24% Slower — Head-to-Head With DFlash 2

Qwen3.8 Native MTP on a 24GB Mac mini: 3GB Less Memory, 24% Slower — Head-to-Head With DFlash 2

Qwen3.8 ships a trained multi-token-prediction head, and llama.cpp can mount it with one flag — no separate 2B draft model. I benchmarked it against DFlash 2 on the same 24GB Mac mini M4: memory does drop (16.0GB vs 19.4GB peak), but speed goes backwards — prose falls from 6.0 to 4.5 tok/s (-24%) while DFlash 2 delivers 1.8–1.9x on the same machine. Draft acceptance is healthy (59–85%); the loss is in Metal's verify path — batch-8 decode amortizes at just 1.13x, measured.

qwenspeculative-decoding+6
ai-tutorialsSep 2, 20268 min
321
Turn a Home Mac mini Into an Always-On Claude Code Workstation: claudecodeui + SSH Reverse Tunnel, Take Over Sessions From Any Browser

Turn a Home Mac mini Into an Always-On Claude Code Workstation: claudecodeui + SSH Reverse Tunnel, Take Over Sessions From Any Browser

A Mac mini at home runs Claude Code around the clock — but how do you take over a session from a browser when you're away? This is a real setup that has been live for a week and in daily use: claudecodeui as the web UI (chosen over the official web version, ttyd, and code-server), an SSH reverse tunnel pushing it to a VPS, and nginx adding TLS plus login rate limiting to turn it into an ordinary URL. Includes full configs, real operating numbers (five days of tunnel uptime with zero drops, 170MB RSS), a <synthetic> placeholder bug hit and fixed within the first week, and an honest for-and-against on why not Tailscale.

claude-codeclaude-code-lab+7
claudeAug 29, 202612 min
320