Skip to content
llm-speed

Why is my local LLM slow? Check GPU offload, context and latency

Published 2026-09-12 · updated 2026-09-16

First identify where the time goes: loading the model, processing the prompt, waiting for a free slot, or generating tokens. Then check the actual model placement and context allocation. A GPU upgrade is a decision to test after that diagnosis; a low headline tok/s number alone cannot tell you what to buy.

Only the first request is slow

In Ollama's native API response, compare load_duration, prompt_eval_duration and eval_duration. They separate model loading, uncached prompt processing and token generation; durations are in nanoseconds. If loading dominates, test again while the model remains resident. Source: Ollama response timings.

Keep the model artifact, input length and output limit fixed. Also record whether a repeated prefix was cached: keeping weights loaded and reusing processed prompt tokens are different effects. Ollama's keep-alive documentation describes model residency. A faster second request is not by itself evidence that all future requests will be equally fast.

Long wait, then fast text

Measure time to first token separately from decode speed. In our Gemma 4 RTX 5090 study, a 95,212-token input had a median 19.901-second first-token wait, followed by 115.98 decode tok/s. Those are three repetitions of Gemma 4 12B IT QAT Q4_0 on a warm, single-request llama.cpp CUDA server, revision 9725a31, with prompt caching and thinking off. That result illustrates the distinction; it does not predict another model's delay or answer quality.

First compare a short input with your real workload. Preserve the context needed to answer correctly. If an app hides reasoning or buffers output, the first visible answer can arrive later than the first generated token. Our prefill, decode and TTFT explainer covers the measurements and full study limitations.

For a native Mac example, our Qwen3.8-27B M5 Max study records all 12 samples, including the cost of preparing a reusable prefix. Shared-workstation timings varied substantially; the article keeps the ranges and memory-pressure history visible.

Every token is slow: verify GPU placement

While your local Ollama model is loaded, run:

ollama ps

Read PROCESSOR: GPU-only, CPU-only and mixed CPU/GPU placement mean different things. This is where the model was loaded, not a live GPU utilization percentage. An empty list means no model is currently loaded; check during your request. See Ollama's placement explanation.

With llama.cpp, inspect startup logs for the selected backend and the actual GPU-offloaded layers. Requesting GPU layers is not proof that all requested work landed there. Follow the upstream performance checks for your build. Avoid copying someone else's layer count as a universal setting.

If you expected GPU use but see CPU-only placement, check GPU discovery and backend errors before changing model settings. Ollama documents log locations and GPU discovery troubleshooting for each platform. If placement is mixed, try a smaller supported artifact or a lower necessary context, then compare placement and speed again. Treat this as a diagnostic change and recheck answer usefulness.

Longer chats slow down: check allocated context

The model file is only part of the memory requirement. Context cache and runtime allocations need room too. Check the context actually allocated by your app or server, alongside the actual input tokens. Ollama shows allocated context in ollama ps; its context guide explains how increasing it changes memory needs. Inspect your installed version's setting instead of assuming a fixed default.

Change one variable at a time: first use one active request, then a shorter sufficient context, then a smaller quantized artifact if needed. Save the original settings. The memory-fit calculator can narrow candidates; choose your installed Mac memory where applicable. Its coarse estimates do not prove fit on your machine. Validate at your intended input length, not just at model load.

Why does SGLang allow less context than the model advertises?

A large model context limit does not prove that your server can hold a request that large. SGLang distinguishes --context-length from the allocated token pool controlled by --max-total-tokens. Without an explicit pool cap, allocation depends on the memory budget. Raising the context flag alone does not add memory. See the official SGLang server arguments.

  1. Record what the server actually allocated. Save its version, checkpoint, startup context limit and token-pool capacity (max_total_num_tokens in scheduler diagnostics). Keep these separate from the maximum printed on the model card.
  2. Check the agent app separately. Record its context setting, output budget and compaction trigger for the served model. The app can shorten a conversation before it reaches the server. An app's context meter is not a GPU-memory measurement.
  3. Locate the limiting layer. Compare one controlled request through the app and directly to the same server, recording actual input tokens, requested output and the exact response or error. Keep templates and other request settings consistent. App-only shortening points to client configuration; a server rejection needs server diagnostics.
  4. Leave room for the whole task. Count system text, tools, images where applicable, and generated output. Start with one active request; concurrent requests share resources. A successful short prompt does not establish long-document capacity or answer quality.

For Qwen3.8-27B, use the official SGLang cookbook for your exact checkpoint, GPU and speculative-decoding setup. Hybrid models also allocate recurrent state; the cookbook's memory-ratio calculator accounts for request length and concurrency. Increasing one pool can squeeze the other. Treat changes to vision, state precision or speculation as capability and quality changes to validate, not free capacity.

This is a diagnostic checklist, not a measured SGLang capacity result. Our RTX 5090 Qwen/Gemma study used llama.cpp with different artifacts and settings; it cannot establish the context available in an SGLang deployment.

Can I run multiple Ollama models at once?

Loading two models and sending two requests to one model are different capacity questions. Ollama can keep multiple models loaded when memory permits. GPU inference requires concurrent models to fit in VRAM; otherwise requests can wait while idle models are unloaded. Parallel requests to one model also increase its context allocation. The controls differ: OLLAMA_MAX_LOADED_MODELS limits resident models, OLLAMA_NUM_PARALLEL limits parallel requests per model, and OLLAMA_MAX_QUEUE limits waiting requests. A larger queue creates no extra memory or compute capacity. See the official concurrency explanation.

Before buying a mini PC or a larger Mac, test the actual combination: your main agent model, any smaller helper model, and the context each needs. Our single-model memory estimate does not establish that the whole combination fits or remains responsive. A small helper can still compete with the main model for resources.

Check the helper's decisions before optimizing its speed. In our small Gemma 4 E2B JSON-output test, a schema fixed formatting but one routing decision still failed. Valid output structure alone does not establish useful agent behavior.

A three-case check before changing hardware

This is our suggested diagnostic, not a measured multi-model benchmark. Use the same main-model prompt and output limit in each case. Record three repetitions, keeping cold loads and cached prompts separate.

Compare one request, parallel requests and alternating models
CaseWhat to recordWhat it helps distinguish
Main model alone, one requestLoaded models, context, first-token wait and completion time.Your baseline at the intended context.
Two overlapping requests to the main modelEach request's latency, overlap and any queue delay.Request contention or waiting, without introducing a second model.
Main model, helper, then main model againLoaded-model list after each step and the main model's load duration on return.Whether switching models triggers a reload. This case tests switching, not simultaneous generation.

Use ollama ps during each case. For a machine-readable snapshot, Ollama's running-model API exposes the loaded models, their size, size_vram and allocated context_length. It is a snapshot, not a record of peak memory or proof that two requests are generating simultaneously.

curl http://localhost:11434/api/ps

If switching produces repeated loads, investigate residency before treating it as slow decoding. If the models remain loaded but overlapping requests take longer, compare a sequential agent workflow before increasing parallelism. Finally test the actual overlapping main/helper workload: measure both answers and keep the configuration only if each still meets your latency and quality needs. Neither a larger queue nor a high single-user tok/s result demonstrates that it will.

Why might a rented GPU be slower than a hosted API?

A GPU name does not specify the serving system. Compare the exact model and quantization, actual input/output lengths, reasoning settings, cache state, active requests and measurement point. Single-user streaming speed and total throughput across a batch answer different questions. Queuing can increase the wait before generation. TNG's serving-performance analysis explains the latency/throughput trade-off.

Start with one warm request and check GPU placement on the rented server. Measure client-observed first-token and total completion time as well as server timings. If the provider does not expose its model or serving settings, report an end-to-end comparison of your task, not a GPU speed ranking. Do not extend a rental merely to chase an unmatched API number.

Save a comparison you can act on

Repeat a fixed workload three times before and after one change. Keep cold loads separate from warm and prefix-cached runs. Record errors and whether the answer still solves the task. This is a suggested diagnostic protocol, not a claim that these changes produce a particular speedup.

Model file / quantization:
Runtime / version:
Hardware / actual memory:
GPU placement / allocated context:
Input / output tokens:
Warm or cold / prompt cache / reasoning:
Active requests:
First token / decode tok/s / total time:
Change tested / answer still useful:

For Ollama generation timing, eval_count / (eval_duration / 1e9) gives output tok/s when the duration is positive. It is not first-token latency or total request speed. Compare the documented timing fields with what your client actually displays.

Once you know the constraint, look up recorded configurations or use the RTX 5090 and Mac Studio comparison to evaluate a capacity change. If your setup fills a missing measurement, record a benchmark.

Sources checked September 12, 2026. This is a diagnostic guide; the cited Gemma measurements are from September 6. No new speedup, rental-provider benchmark or hardware purchase recommendation is claimed here.