Magic Tools
Pitfall NotesBy CooconOctober 5, 20264 views7 min read

LM Studio Slow on a Mac? Tested: the 39-Second Wait Is Prompt Prefill, Not GPU Offload

Short answers first

  • First decide: a long wait for the first token, or slow output? LM Studio's API returns time_to_first_token and tokens_per_second, and the UI shows them too.
  • Long wait for the first token: usually a long prompt (long chat, big attached files, or LM Studio behind Claude Code / Cline, whose system prompts run to tens of thousands of tokens). A 13K-token prompt waited 39 seconds before the first token. The next request with the same prefix took 0.1 s.
  • Slow output: check whether GPU offload was turned down — though on a Mac it matters less than you'd think: 29.5 tok/s at max, 26.4 tok/s fully off.
  • Several tools sharing one LM Studio: with --parallel 4, four simultaneous requests ran at 9.7 tok/s each; with 1 they queue and each runs at full speed.
  • Still slow: use a smaller model or lower quantization; on a Mac you can also try the MLX build of the same model (see Related).

Background

"LM Studio is slow" covers at least two different problems:

  1. Slow start: nothing for a long time, then reasonable output speed;
  2. Slow output: tokens trickle out one by one.

The usual advice is "max out GPU offload", but on Apple Silicon that is often not the main cause. This article measures the common factors one at a time on a 24GB M4 Mac mini.

Two phases

A request has two parts:

  • Prefill: the model reads the whole prompt before the first token; time grows with prompt length — this is time_to_first_token.
  • Decode: tokens are then generated one at a time — this is tokens_per_second.

Settings that affect them: GPU offload (--gpu), context length (-c), parallel predictions (--parallel), and whether LM Studio reuses the previous request's prefix. The lms CLI exposes all of these, and the REST API (/api/v0/chat/completions) returns stats.tokens_per_second and stats.time_to_first_token, which is what we measured.

Method

Approach Verdict Why
Load with lms, read stats from the REST API ✅ used Load options controlled one by one; speed figures are LM Studio's own
Try things in the chat UI ❌ Hard to record and reproduce
Download an MLX build to compare ❌ not this time Another ~6GB download; an MLX vs llama.cpp comparison already exists on this site (see Related)

Constraints:

  • LM Studio's state was recorded first (no model loaded, server off) and restored at the end; the server ran on a separate port, 1239.
  • A memory guard ran lms unload --all if free memory fell below 35% (it fired once).
  • Model google/gemma-4-e4b (7.5B params, GGUF Q4_K_M, 6.33GB, llama.cpp engine 2.13.0); 200 generated tokens per request, temperature 0; each case run twice.

Results

1. Defaults: 4k context, 29.5 tok/s

Loaded with no options: context 4096, GPU offload automatic (same result as --gpu max), 29.62 / 29.46 tok/s, 0.3 s to first token on a short prompt.

--estimate-only estimates memory without loading — friendlier than Ollama, which doesn't stop you (see Ollama out of memory on a Mac):

Context LM Studio estimate
4096 6.29 GiB
32768 7.85 GiB
131072 13.18 GiB

Each estimate says This model may be loaded based on your resource guardrails settings, i.e. it checks guardrails before loading.

2. GPU offload: small effect on a Mac

4096 context, short prompt:

--gpu Output speed Time to first token
max 29.56 / 29.52 tok/s 0.32 / 0.31 s
0.75 27.27 / 26.99 0.41 / 0.39
0.5 26.98 / 26.84 0.47 / 0.46
0.25 26.38 / 26.12 0.58 / 0.69
off (all CPU) 26.16 / 26.66 0.65 / 0.65

All on CPU, output is only about 11% slower — consistent with the Ollama article: on Apple Silicon the CPU and GPU share one memory pool, so moving layers to the CPU costs far less than on a discrete GPU. Time to first token doubled, but on a short prompt that's still 0.65 s — long prompts are what matter.

3. Prompt length: the real cause of a slow start

Context 32768, prompts of about 600, 3,300 and 13,000 tokens:

Prompt tokens GPU first token GPU output CPU first token CPU output
576 1.65 s 29.42 tok/s 5.53 s 24.37 tok/s
3268 8.99 s 28.66 tok/s 33.51 s 20.62 tok/s
12961 39.44 s 25.95 tok/s not completed (see below) —
  • Prefill on GPU runs at about 330–360 tokens/s (576/1.65, 3268/8.99, 12961/39.44). A 13K-token prompt waits 39 seconds for the first token.
  • On CPU, prefill is 3.4–3.7x slower (5.53 vs 1.65, 33.51 vs 8.99). GPU offload barely affects output speed but strongly affects the first token.
  • Longer context also slows output a little: on GPU, 29.4 → 26.0 tok/s.

In the CPU + 13K-token case, free memory fell to 31% during prefill and the memory guard unloaded the model, so there's no data for it.

This is why LM Studio "hangs" behind coding tools like Claude Code or Cline: every request carries tens of thousands of tokens of system prompt and tool definitions.

4. Prefix caching: the second time is fast

Each case above sent the identical request twice. Time to first token:

Prompt tokens First request Second request
576 1.65 s 0.08 s
3268 8.99 s 0.08 s
12961 39.44 s 0.12 s

LM Studio reused the previous request's prefix. Within a conversation, as long as the earlier content doesn't change, later turns only process what's new. Conversely, if your tool changes the start of the prompt every time (say, a timestamp in the system prompt), the cache misses and every turn prefills from scratch. The CPU cases' second requests also took only 0.09–0.10 s.

5. Parallel predictions: more throughput, slower individual requests

--parallel sets how many requests the model processes at once. Four simultaneous requests (200 tokens each):

Setting Per-request speed All 4 done Total throughput
--parallel 1, 1 request 29.71 tok/s — —
--parallel 4, 1 request 29.71 tok/s — —
--parallel 1, 4 at once 29.6–29.8 tok/s each, but queued: finished at 6.99 / 13.94 / 20.87 / 27.79 s 27.79 s 28.8 tok/s
--parallel 4, 4 at once 9.66–9.69 tok/s each 21.35 s 37.5 tok/s

With parallelism, the four requests advance together, total throughput is 30% higher and the last one finishes sooner — but each user sees a third of the speed. If a chat window, editor completions and an agent all share one LM Studio with parallelism on, everything feels slow. With a single request, the setting makes no difference.

What to do

1. Identify which kind of slow. Check time_to_first_token and tokens_per_second; over the API they're in stats:

curl -s http://127.0.0.1:1234/api/v0/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"<your model ID>","messages":[{"role":"user","content":"hi"}],"max_tokens":50}' \
  | python3 -c "import json,sys; print(json.load(sys.stdin)['stats'])"

(Use your LM Studio server port; the default is 1234.)

2. Slow first token:

  • Keep prompts short: start new chats for long conversations; don't paste whole large files.
  • Behind coding tools, a first-turn wait of tens of seconds is prefill and unavoidable; later turns are much faster thanks to prefix caching. Don't let the tool put changing content at the start of the prompt.
  • Make sure GPU offload is at max — it changes prefill by 3–4x.

3. Slow output:

  • Set GPU offload back to max (about 10% on a Mac);
  • Don't set context far beyond what you need; longer context slows output somewhat;
  • Use a smaller model or lower quantization; on a Mac you can try the MLX build (our single comparison: Qwen3.8-27B ran 6.4–6.5 tok/s on MLX vs 5.9–6.0 tok/s as GGUF, about 8% faster).

4. Several tools sharing it: load with --parallel 1 (or set concurrent predictions to 1 in the UI) so requests queue at full speed each; enable parallelism only when total throughput matters more.

5. Estimate before loading: lms load <model> -c 32768 --estimate-only shows memory needs without loading.

Pitfalls along the way

  • zsh doesn't split unquoted variables. The parallel test passed arguments as for pc in "1 1" "4 1"; do python3 par.py $pc; done, and the script got a single argument; in zsh use ${=pc} or pass them separately.
  • The memory guard fired on CPU + long prompt. CPU prefill of 13K tokens used noticeably more memory; the guard unloaded the model. That data point is missing, but other services on the machine weren't affected.

Not verified: the same comparisons on the MLX engine (needs another download); larger models (only one 7.5B GGUF here); CPU prefill time for the 13K-token prompt; prefix-cache hits in real multi-turn chats or partially changed prompts; other UI settings (e.g. Flash Attention, KV cache quantization) and their effect on speed. Each case ran only twice.

Get field notes like this every Saturday

Subscribe to Dev Breakfast: daily AI coding picks at 8:00, plus a Saturday roundup of this week's hands-on tests with Claude Code / Codex / local models. Written in Chinese.

Related Articles

Ollama Out of Memory on a Mac: How Much num_ctx, Parallelism and KV Quantization Really Cost on 24GB (Tested) — and the Trap It Won't Stop You From

On a 24GB M4 Mac mini, Ollama 0.19 sees only 17.8 GiB of VRAM and defaults to a 4096-token context. With qwen3:4b, going from 4k to 32k context grows memory from 3.73GB to 9.89GB; OLLAMA_NUM_PARALLEL=4 at 8k uses exactly as much as a single 32k slot while ollama ps still shows 8192; Flash Attention + q4_0 KV cache brings 32k down to 4.31GB and 64k to 5.77GB with no speed loss. Ask for a context far beyond RAM and Ollama doesn't refuse — it starts allocating, and free memory fell to 23% within 16 seconds. Running fully on CPU was only 25% slower.

local-llmmac-mini+5
pitfallsOct 5, 20267 min
3
llama.cpp's llama-server or Ollama? Same GGUF on a 24GB Mac mini, Tested: Ollama Runs llama-server Under the Hood

llama.cpp's llama-server or Ollama? Same GGUF on a 24GB Mac mini, Tested: Ollama Runs llama-server Under the Hood

llama-server (b11376) and Ollama (0.35.1) loading the same GGUF file on a 24GB Mac mini. Ollama 0.35's inference process is its own bundled llama-server, and with the same settings speed is identical (4B decode 36.5 vs 36.3 tok/s, 27B 6.3 on both). Every difference comes from defaults: Ollama defaults to a 4096-token context and silently cuts an over-long prompt for library models down to 2050 tokens, leaving only a WARN line in its log; llama-server sizes context to fill memory (101,120 for the 4B, 14GB+ footprint) and returns 400 when a prompt doesn't fit. llama-server runs 4 parallel slots by default while Ollama queues, and Ollama with 32k context plus 4-way parallelism allocates 32k per slot, a 21GB footprint. Recommended flags for 24GB machines at the end.

benchmarkqwen+6
hands-onOct 3, 202613 min
35
Same 24GB Mac mini, Same DFlash 2: 1.9x Faster on MLX, 50% Slower (and OOM) on llama.cpp

Same 24GB Mac mini, Same DFlash 2: 1.9x Faster on MLX, 50% Slower (and OOM) on llama.cpp

Part three of my speculative-decoding trilogy on a base Mac mini M4. llama.cpp merged DFlash 2 support with official GGUF drafts — and every configuration is a net slowdown. The README-recommended n-max 7 hits a reproducible Metal OOM on 24GB; the only stable setting cuts prose from 6.0 to 3.0 tok/s, and an 83.8% acceptance rate on code still loses 23%. Same algorithm, same machine, MLX gets 1.8–1.9x. The arithmetic shows why: 0.77s per speculative step loses even at 100% acceptance.

qwendflash+6
ai-tutorialsSep 3, 20268 min
280
Qwen3.8 Native MTP on a 24GB Mac mini: 3GB Less Memory, 24% Slower — Head-to-Head With DFlash 2

Qwen3.8 Native MTP on a 24GB Mac mini: 3GB Less Memory, 24% Slower — Head-to-Head With DFlash 2

Qwen3.8 ships a trained multi-token-prediction head, and llama.cpp can mount it with one flag — no separate 2B draft model. I benchmarked it against DFlash 2 on the same 24GB Mac mini M4: memory does drop (16.0GB vs 19.4GB peak), but speed goes backwards — prose falls from 6.0 to 4.5 tok/s (-24%) while DFlash 2 delivers 1.8–1.9x on the same machine. Draft acceptance is healthy (59–85%); the loss is in Metal's verify path — batch-8 decode amortizes at just 1.13x, measured.

qwenspeculative-decoding+6
ai-tutorialsSep 2, 20268 min
324

Published by Magic Tools