Run a Local LLM on a Mini PC: Self-Host Your Own ChatGPT
You don’t need a $2,000 GPU to run a usable local AI. A ~$400 mini PC with an integrated GPU and enough RAM will run 7B–8B models fast enough for daily use — and every token stays on your hardware. Here’s the exact setup I run.
What you need
- A mini PC with a modern APU (I use a Ryzen 7 5700G) and at least 32GB RAM (I run 64GB so I can hold a model and still run other services).
- ~20GB of disk for a couple of models.
- Linux (this guide assumes Ubuntu/Debian; Proxmox VM or bare metal both work).
No discrete GPU. On an APU the model runs on CPU + shared memory, which is why RAM matters more than anything else here.
Step 1 — install Ollama
Ollama is the laziest way to pull and serve models. One line:
curl -fsSL https://ollama.com/install.sh | sh
Pull a model and talk to it:
ollama run llama3.1:8b
The first run downloads ~4.7GB. After that it’s instant to start.
Step 2 — pick a model that actually fits
RAM is the constraint. A rough rule for a CPU-only box:
| Model size | RAM needed | Feels like |
|---|---|---|
3B (e.g. llama3.2:3b) |
~4GB | Fast, fine for summaries/classification |
7–8B (llama3.1:8b, qwen2.5:7b) |
~8GB | The sweet spot — general chat, coding help |
| 14B | ~16GB | Noticeably smarter, slower on CPU |
On the 5700G I get ~10–14 tokens/sec on an 8B model — faster than you read. Above 14B on CPU it drops below reading speed and stops being fun.
Step 3 — add a real chat UI
The terminal is fine for testing, but you want a proper interface. Open WebUI gives you a ChatGPT-style UI, pointed at your local Ollama. Run it in Docker:
docker run -d -p 3000:8080 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
--name open-webui --restart always \
ghcr.io/open-webui/open-webui:main
Open http://<your-mini-pc-ip>:3000, create the first account (it becomes admin),
and it auto-detects your Ollama models. Done — a private ChatGPT on hardware you own.
Keep it on your LAN. There’s no reason to expose this to the public internet — if you want it on your phone away from home, reach it over a VPN like Tailscale, not a port-forward.
Is it actually usable?
For chat, drafting, summarizing, classification, and coding help: yes, an 8B model on this box is genuinely useful every day. For frontier-level reasoning it won’t match GPT-class cloud models — but it’s free, private, and always available, and that trade is worth it for most of what you actually do.
That’s the whole build — and it runs alongside n8n on the same box, no extra hardware.
Gear used in this build
* Affiliate links — I earn a small commission at no cost to you. It's gear I use and would genuinely recommend. See the full disclosure.
New self-hosted AI & homelab shorts, daily.
Subscribe on YouTubeRelated guides
Run Claude Code and Codex Locally with OtoDock
Deploy OtoDock on your server to get local code generation without cloud API costs. Self-hosted alternative to Claude Code and GitHub Copilot.
Run Claude Code Locally with OtoDock
Deploy OtoDock on your own hardware to run Claude Code and Codex as local AI agents without paying per API call. Full self-hosted setup.
Self-host Claude Code agents with OtoDock
Run Claude's code execution engine on your own hardware. Replace the SaaS with OtoDock—a self-hosted agent framework that keeps your inference and execution local.