aihomelabprivacy

Tune Ollama for Real Performance: Memory, Quantization, and Hardware Hacks

🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.

Your local LLM feels sluggish because you’re probably running it with default settings—high precision, no GPU acceleration, and a context window that thrashes your disk. Ollama ships with sensible defaults, but they’re not optimized for your specific hardware. This guide walks through the levers you actually control to get faster, cheaper inference without sacrificing quality.

Install Ollama and Understand Your Baseline

  1. Install Ollama for your OS:
curl -fsSL https://ollama.com/install.sh | sh
  1. Start the Ollama service:
ollama serve

This runs Ollama on localhost:11434 by default. Leave it running in a terminal or background process.

  1. Pull a model to test with—Gemma 4 is a good middle ground:
ollama pull gemma4
  1. Run a quick chat to see baseline speed:
ollama run gemma4

Note the tokens-per-second (t/s) output. You’ll compare against this.

Quantization: Trade Precision for Speed

Quantization reduces model weights from 16-bit or 32-bit floats to 8-bit or 4-bit integers. Ollama handles this automatically when you pull a model—it grabs the smallest quantized version by default. But you can be explicit.

  1. Check what quantized variants are available on ollama.com/library. Most models ship in Q4 (4-bit), Q5, and Q8 (8-bit) versions.

  2. Pull a specific quantization level. For example, the Q4 variant of Gemma 4:

ollama pull gemma4:q4_0
  1. Run it and measure t/s again:
ollama run gemma4:q4_0 "Explain quantum computing in one paragraph"

Q4 is typically 2–4× faster than Q8 with acceptable quality loss for most tasks. Q5 splits the difference. Smaller quantizations (Q3, Q2) run on constrained hardware but degrade output noticeably.

GPU Offloading: Move Compute Off the CPU

If you have an NVIDIA GPU, Ollama can offload layers to VRAM for massive speedup. AMD and Intel GPUs are also supported; check the Ollama docs for your card.

  1. Verify your GPU is detected. On Linux, check:
ls /usr/local/cuda/bin/nvidia-smi

Ollama automatically uses NVIDIA GPUs if CUDA is installed.

  1. Control how many layers to offload with the num_gpu parameter in a chat request. Use the REST API to set this explicitly:
curl http://localhost:11434/api/chat -d '{
  "model": "gemma4:q4_0",
  "messages": [{"role": "user", "content": "test"}],
  "num_gpu": 50,
  "stream": false
}'

The num_gpu value is the number of layers to offload. Higher = more GPU use, faster inference (if VRAM allows). Start with 50 and increase until you hit out-of-memory errors, then back off.

  1. For AMD GPUs, set the OLLAMA_ROCM_PATHS environment variable before running Ollama. Refer to the official Ollama docs for your specific GPU model.

Context Window Tuning

Larger context windows let the model see more input, but they consume memory linearly. If your inference is slow, a bloated context window is often the culprit.

  1. Check the default context length for your model (usually 2048 or 4096 tokens). This is set in the model’s Modelfile.

  2. Reduce it if you don’t need the full window. Create a custom Modelfile:

cat > Modelfile << 'EOF'
FROM gemma4:q4_0
PARAMETER num_ctx 2048
EOF
  1. Build and run it:
ollama create gemma4-fast -f Modelfile
ollama run gemma4-fast

Smaller context = faster per-token inference. For chat or short tasks, 2048 is usually enough.

CPU Affinity and Thread Count

Ollama uses all available CPU threads by default. On shared hardware or systems with many cores, this can cause contention.

  1. Set the number of threads explicitly. Use the OLLAMA_NUM_THREAD environment variable:
export OLLAMA_NUM_THREAD=8
ollama serve

Start with half your core count and adjust. Fewer threads = lower latency on single queries; more threads = better throughput on concurrent requests.

  1. On Linux, pin Ollama to specific CPU cores for consistent performance:
taskset -c 0-7 ollama serve

This binds Ollama to cores 0–7. Adjust the range to match your thread count.

Memory and Swap Tuning

If inference slows down mid-conversation, you’re likely hitting swap. Increase available RAM or reduce context size.

  1. Monitor memory during inference:
watch -n 1 free -h
  1. If swap usage climbs, disable it temporarily to see the real memory floor:
sudo swapoff -a

(Re-enable with swapon -a when done.)

  1. If you’re out of RAM, reduce num_ctx or switch to an even smaller quantization (Q3 instead of Q4).

Batch Size and Request Pipelining

Ollama can process multiple requests in parallel if you’re running a service (not just chat). The REST API handles this automatically, but you can tune batch behavior.

  1. For concurrent requests, increase the batch size. Set OLLAMA_BATCH_SIZE before starting:
export OLLAMA_BATCH_SIZE=256
ollama serve

Higher batch sizes = better GPU utilization but higher latency per request. Default is 512; reduce to 128 if you want snappier single requests.

Profile Your Setup

Run a real benchmark to see where you stand:

time curl http://localhost:11434/api/chat -d '{
  "model": "gemma4:q4_0",
  "messages": [{"role": "user", "content": "Write a 500-word essay on distributed systems."}],
  "stream": false
}'

Note the total time and divide by output token count (visible in the response) to get t/s. Compare before and after each tuning step. A good target is 10–20 t/s for a 7B model on consumer hardware; 30+ t/s if you have a decent GPU.

Integrate Ollama into Your Workflow

Once tuned, expose Ollama to your internal tools via the REST API. Keep it behind a firewall—never expose port 11434 to the public internet.

  1. Use the Python SDK for scripting:
pip install ollama
from ollama import chat

response = chat(model='gemma4:q4_0', messages=[
  {'role': 'user', 'content': 'Summarize this: [text]'},
])
print(response.message.content)
  1. Use the JavaScript SDK for web apps:
npm install ollama
import ollama from "ollama";

const response = await ollama.chat({
  model: "gemma4:q4_0",
  messages: [{ role: "user", content: "test" }],
});
console.log(response.message.content);
  1. Pair Ollama with Open WebUI for a ChatGPT-like interface on your LAN:
docker run -d -p 3000:8080 ghcr.io/open-webui/open-webui:latest

Then point it to http://localhost:11434 as your Ollama backend.

Keep Ollama on your LAN only. If you need remote access, use a VPN. The REST API has no built-in auth.

Is It Worth It?

Yes, if you have the hardware. A tuned Ollama setup on a mid-range CPU (Ryzen 5700G or better) or with a GPU hits 10–30 t/s for 7B models—fast enough for real work. You trade off the polish of ChatGPT for full privacy and zero API costs. The setup is straightforward: quantize aggressively, offload to GPU if you have one, and dial in context and threads for your workload. Start with Q4 quantization and a 2048-token context; adjust from there. Most performance gains come from quantization and GPU offloading; the rest is fine-tuning.

Related guides

← All guides