LM Studio Slow on a Mac? Tested: the 39-Second Wait Is Prompt Prefill, Not GPU Offload
Short answers first
- First decide: a long wait for the first token, or slow output? LM Studio's API returns
time_to_first_tokenandtokens_per_second, and the UI shows them too.- Long wait for the first token: usually a long prompt (long chat, big attached files, or LM Studio behind Claude Code / Cline, whose system prompts run to tens of thousands of tokens). A 13K-token prompt waited 39 seconds before the first token. The next request with the same prefix took 0.1 s.
- Slow output: check whether GPU offload was turned down — though on a Mac it matters less than you'd think: 29.5 tok/s at max, 26.4 tok/s fully off.
- Several tools sharing one LM Studio: with
--parallel 4, four simultaneous requests ran at 9.7 tok/s each; with 1 they queue and each runs at full speed.- Still slow: use a smaller model or lower quantization; on a Mac you can also try the MLX build of the same model (see Related).
Background
"LM Studio is slow" covers at least two different problems:
- Slow start: nothing for a long time, then reasonable output speed;
- Slow output: tokens trickle out one by one.
The usual advice is "max out GPU offload", but on Apple Silicon that is often not the main cause. This article measures the common factors one at a time on a 24GB M4 Mac mini.
Two phases
A request has two parts:
- Prefill: the model reads the whole prompt before the first token; time grows with prompt length — this is
time_to_first_token. - Decode: tokens are then generated one at a time — this is
tokens_per_second.
Settings that affect them: GPU offload (--gpu), context length (-c), parallel predictions (--parallel), and whether LM Studio reuses the previous request's prefix. The lms CLI exposes all of these, and the REST API (/api/v0/chat/completions) returns stats.tokens_per_second and stats.time_to_first_token, which is what we measured.
Method
| Approach | Verdict | Why |
|---|---|---|
Load with lms, read stats from the REST API |
✅ used | Load options controlled one by one; speed figures are LM Studio's own |
| Try things in the chat UI | ❌ | Hard to record and reproduce |
| Download an MLX build to compare | ❌ not this time | Another ~6GB download; an MLX vs llama.cpp comparison already exists on this site (see Related) |
Constraints:
- LM Studio's state was recorded first (no model loaded, server off) and restored at the end; the server ran on a separate port, 1239.
- A memory guard ran
lms unload --allif free memory fell below 35% (it fired once). - Model
google/gemma-4-e4b(7.5B params, GGUF Q4_K_M, 6.33GB, llama.cpp engine 2.13.0); 200 generated tokens per request,temperature 0; each case run twice.
Results
1. Defaults: 4k context, 29.5 tok/s
Loaded with no options: context 4096, GPU offload automatic (same result as --gpu max), 29.62 / 29.46 tok/s, 0.3 s to first token on a short prompt.
--estimate-only estimates memory without loading — friendlier than Ollama, which doesn't stop you (see Ollama out of memory on a Mac):
| Context | LM Studio estimate |
|---|---|
| 4096 | 6.29 GiB |
| 32768 | 7.85 GiB |
| 131072 | 13.18 GiB |
Each estimate says This model may be loaded based on your resource guardrails settings, i.e. it checks guardrails before loading.
2. GPU offload: small effect on a Mac
4096 context, short prompt:
--gpu |
Output speed | Time to first token |
|---|---|---|
| max | 29.56 / 29.52 tok/s | 0.32 / 0.31 s |
| 0.75 | 27.27 / 26.99 | 0.41 / 0.39 |
| 0.5 | 26.98 / 26.84 | 0.47 / 0.46 |
| 0.25 | 26.38 / 26.12 | 0.58 / 0.69 |
| off (all CPU) | 26.16 / 26.66 | 0.65 / 0.65 |
All on CPU, output is only about 11% slower — consistent with the Ollama article: on Apple Silicon the CPU and GPU share one memory pool, so moving layers to the CPU costs far less than on a discrete GPU. Time to first token doubled, but on a short prompt that's still 0.65 s — long prompts are what matter.
3. Prompt length: the real cause of a slow start
Context 32768, prompts of about 600, 3,300 and 13,000 tokens:
| Prompt tokens | GPU first token | GPU output | CPU first token | CPU output |
|---|---|---|---|---|
| 576 | 1.65 s | 29.42 tok/s | 5.53 s | 24.37 tok/s |
| 3268 | 8.99 s | 28.66 tok/s | 33.51 s | 20.62 tok/s |
| 12961 | 39.44 s | 25.95 tok/s | not completed (see below) | — |
- Prefill on GPU runs at about 330–360 tokens/s (576/1.65, 3268/8.99, 12961/39.44). A 13K-token prompt waits 39 seconds for the first token.
- On CPU, prefill is 3.4–3.7x slower (5.53 vs 1.65, 33.51 vs 8.99). GPU offload barely affects output speed but strongly affects the first token.
- Longer context also slows output a little: on GPU, 29.4 → 26.0 tok/s.
In the CPU + 13K-token case, free memory fell to 31% during prefill and the memory guard unloaded the model, so there's no data for it.
This is why LM Studio "hangs" behind coding tools like Claude Code or Cline: every request carries tens of thousands of tokens of system prompt and tool definitions.
4. Prefix caching: the second time is fast
Each case above sent the identical request twice. Time to first token:
| Prompt tokens | First request | Second request |
|---|---|---|
| 576 | 1.65 s | 0.08 s |
| 3268 | 8.99 s | 0.08 s |
| 12961 | 39.44 s | 0.12 s |
LM Studio reused the previous request's prefix. Within a conversation, as long as the earlier content doesn't change, later turns only process what's new. Conversely, if your tool changes the start of the prompt every time (say, a timestamp in the system prompt), the cache misses and every turn prefills from scratch. The CPU cases' second requests also took only 0.09–0.10 s.
5. Parallel predictions: more throughput, slower individual requests
--parallel sets how many requests the model processes at once. Four simultaneous requests (200 tokens each):
| Setting | Per-request speed | All 4 done | Total throughput |
|---|---|---|---|
--parallel 1, 1 request |
29.71 tok/s | — | — |
--parallel 4, 1 request |
29.71 tok/s | — | — |
--parallel 1, 4 at once |
29.6–29.8 tok/s each, but queued: finished at 6.99 / 13.94 / 20.87 / 27.79 s | 27.79 s | 28.8 tok/s |
--parallel 4, 4 at once |
9.66–9.69 tok/s each | 21.35 s | 37.5 tok/s |
With parallelism, the four requests advance together, total throughput is 30% higher and the last one finishes sooner — but each user sees a third of the speed. If a chat window, editor completions and an agent all share one LM Studio with parallelism on, everything feels slow. With a single request, the setting makes no difference.
What to do
1. Identify which kind of slow. Check time_to_first_token and tokens_per_second; over the API they're in stats:
curl -s http://127.0.0.1:1234/api/v0/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"<your model ID>","messages":[{"role":"user","content":"hi"}],"max_tokens":50}' \
| python3 -c "import json,sys; print(json.load(sys.stdin)['stats'])"
(Use your LM Studio server port; the default is 1234.)
2. Slow first token:
- Keep prompts short: start new chats for long conversations; don't paste whole large files.
- Behind coding tools, a first-turn wait of tens of seconds is prefill and unavoidable; later turns are much faster thanks to prefix caching. Don't let the tool put changing content at the start of the prompt.
- Make sure GPU offload is at max — it changes prefill by 3–4x.
3. Slow output:
- Set GPU offload back to max (about 10% on a Mac);
- Don't set context far beyond what you need; longer context slows output somewhat;
- Use a smaller model or lower quantization; on a Mac you can try the MLX build (our single comparison: Qwen3.8-27B ran 6.4–6.5 tok/s on MLX vs 5.9–6.0 tok/s as GGUF, about 8% faster).
4. Several tools sharing it: load with --parallel 1 (or set concurrent predictions to 1 in the UI) so requests queue at full speed each; enable parallelism only when total throughput matters more.
5. Estimate before loading: lms load <model> -c 32768 --estimate-only shows memory needs without loading.
Pitfalls along the way
- zsh doesn't split unquoted variables. The parallel test passed arguments as
for pc in "1 1" "4 1"; do python3 par.py $pc; done, and the script got a single argument; in zsh use${=pc}or pass them separately. - The memory guard fired on CPU + long prompt. CPU prefill of 13K tokens used noticeably more memory; the guard unloaded the model. That data point is missing, but other services on the machine weren't affected.
Not verified: the same comparisons on the MLX engine (needs another download); larger models (only one 7.5B GGUF here); CPU prefill time for the 13K-token prompt; prefix-cache hits in real multi-turn chats or partially changed prompts; other UI settings (e.g. Flash Attention, KV cache quantization) and their effect on speed. Each case ran only twice.