Run Local LLMs on Bare Metal with Ollama
🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.
Ollama runs open-source large language models on your own hardware—no API keys, no cloud bills, no data leaving your network. You get a REST API and CLI to chat with models like Gemma 4, Llama, or Mistral. This guide gets you running on Linux bare metal in about 10 minutes.
Install Ollama on Linux
- Download and run the official installer:
curl -fsSL https://ollama.com/install.sh | sh
This installs the ollama binary and sets up a systemd service.
- Start the Ollama service:
sudo systemctl start ollama
sudo systemctl enable ollama
Verify it’s running:
sudo systemctl status ollama
Ollama listens on http://localhost:11434 by default.
Pull and Run Your First Model
- Pull a model from the library. Gemma 4 is a good starting point—smaller than Llama but capable:
ollama pull gemma4
This downloads the model weights to ~/.ollama/models (several GB, depending on the model).
- Run the model interactively:
ollama run gemma4
You’ll get a chat prompt. Type a question and press Enter. Type /bye to exit.
Query via REST API
Ollama exposes a REST API on port 11434. Test it without needing the CLI:
curl http://localhost:11434/api/chat -d '{
"model": "gemma4",
"messages": [{"role": "user", "content": "Why is the sky blue?"}],
"stream": false
}'
The response is JSON with the model’s answer in the message.content field. Set "stream": true for streaming output.
Use Ollama from Python
- Install the Python client:
pip install ollama
- Write a simple script:
from ollama import chat
response = chat(model='gemma4', messages=[
{
'role': 'user',
'content': 'Explain quantum entanglement in one sentence.',
},
])
print(response.message.content)
Run it:
python3 script.py
Use Ollama from JavaScript/Node.js
- Install the Node client:
npm i ollama
- Create a script:
import ollama from "ollama";
const response = await ollama.chat({
model: "gemma4",
messages: [{ role: "user", content: "What is machine learning?" }],
});
console.log(response.message.content);
Run it:
node script.js
Connect a Web UI
Ollama runs headless by default. Pair it with a web interface like Open WebUI for a ChatGPT-like experience:
- Install Open WebUI via Docker (if you have Docker):
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway ghcr.io/open-webui/open-webui:latest
-
Open
http://localhost:3000in your browser. -
When prompted, set the Ollama API URL to
http://host.docker.internal:11434(orhttp://localhost:11434if running Open WebUI natively).
Alternatively, install Open WebUI directly on the host if you prefer no Docker.
Network Security
By default, Ollama listens only on localhost (127.0.0.1). Keep it that way—never expose port 11434 to the public internet. If you need remote access:
- Use SSH port forwarding:
ssh -L 11434:localhost:11434 user@server - Run Ollama behind a reverse proxy on your LAN only
- Keep your homelab on a private VLAN or behind a VPN
Monitor and Manage Models
List all pulled models:
ollama list
Delete a model to free disk space:
ollama rm gemma4
Check Ollama logs:
sudo journalctl -u ollama -f
Performance Tips
- GPU acceleration: If your hardware has NVIDIA or AMD GPU support, Ollama uses it automatically. Check the logs for
GPUorCUDAmessages. - Model size: Smaller models (7B parameters) run on modest hardware. Larger models (13B+) need more VRAM and CPU.
- Context length: Longer conversations consume more memory. Adjust via the API if needed.
- Concurrent requests: Ollama serializes requests by default. Use a queue or load balancer for parallel inference.
Is It Worth It?
Yes, if you want privacy and control. You own your data, no API costs, and models run offline. The trade-off: slower inference than cloud APIs (GPT-4, Claude) and you manage the hardware. For homelab automation, local document analysis, or private AI experiments, Ollama is hard to beat. For production at scale, you’ll outgrow it—but that’s not the point of bare metal.
Gear used in this build
* Affiliate links — I earn a small commission at no cost to you. It's gear I use and would genuinely recommend. See the full disclosure.
Related video
New self-hosted AI & homelab shorts, daily.
Subscribe on YouTubeRelated guides
Run Claude Code and Codex Locally with OtoDock
Deploy OtoDock on your server to get local code generation without cloud API costs. Self-hosted alternative to Claude Code and GitHub Copilot.
Run Claude Code Locally with OtoDock
Deploy OtoDock on your own hardware to run Claude Code and Codex as local AI agents without paying per API call. Full self-hosted setup.
Self-host Claude Code agents with OtoDock
Run Claude's code execution engine on your own hardware. Replace the SaaS with OtoDock—a self-hosted agent framework that keeps your inference and execution local.