Run Ollama Locally: Skip Google's Cloud, Own Your AI
🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.
Google’s latest AI updates push cloud inference and subscription pricing. Ollama flips that script: run open-source models directly on your hardware, no API key, no monthly bill, no data leaving your network.
This guide gets you chatting with Gemma 4 or any model from the Ollama library in under 10 minutes. Everything runs locally behind your firewall.
Install Ollama
-
On Linux, run the official installer:
curl -fsSL https://ollama.com/install.sh | sh -
On macOS, use the same command or download the DMG manually from ollama.com/download.
-
On Windows, run the PowerShell installer:
irm https://ollama.com/install.ps1 | iexOr download OllamaSetup.exe from ollama.com/download.
-
Verify installation:
ollama --version
Run Your First Model
-
Start Ollama and download Gemma 4 (or any model from ollama.com/library):
ollama run gemma4The first run downloads the model—size varies by model, usually 4–40 GB. Be patient.
-
Once loaded, you’ll see a prompt. Type your question:
>>> Why is the sky blue?The model responds locally. No internet required after download.
-
Exit the chat with
Ctrl+Dor type/bye.
Access via REST API
Ollama exposes a local REST API on port 11434 by default. This lets you integrate models into scripts, web apps, or other tools.
-
Start the Ollama service (runs in background on Linux/macOS):
ollama serveOn macOS/Windows, the service starts automatically.
-
Query the API from another terminal or script:
curl http://localhost:11434/api/chat -d '{ "model": "gemma4", "messages": [{ "role": "user", "content": "Explain machine learning in one sentence." }], "stream": false }'
Use Ollama with Python
-
Install the Python client:
pip install ollama -
Write a simple script (
chat.py):from ollama import chat response = chat(model='gemma4', messages=[ { 'role': 'user', 'content': 'What is a neural network?', }, ]) print(response.message.content) -
Run it:
python chat.py
Use Ollama with JavaScript
-
Install the npm package:
npm i ollama -
Create a script (
chat.js):import ollama from "ollama"; const response = await ollama.chat({ model: "gemma4", messages: [{ role: "user", content: "What is a neural network?" }], }); console.log(response.message.content); -
Run it:
node chat.js
Connect a Web UI
Ollama is headless by default. For a ChatGPT-like interface, pair it with Open WebUI or another community frontend.
-
Using Docker (if you have Docker installed):
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway ghcr.io/open-webui/open-webui:latestOpen http://localhost:3000 in your browser.
-
Or run Ollama and Open WebUI on the same machine without Docker—see the Open WebUI repo for native setup.
-
Point the UI to your local Ollama instance at
http://localhost:11434.
Network Security
Important: By default, Ollama listens only on localhost (127.0.0.1) and is not accessible from other machines. Keep it that way. If you need to expose it to your LAN:
- Set
OLLAMA_HOST=0.0.0.0:11434as an environment variable only if your machine is behind a router firewall. - Never expose port 11434 to the public internet. Always keep Ollama on your LAN or behind a VPN.
- If you need remote access, use a VPN (WireGuard, Tailscale, OpenVPN) instead of opening the port.
Manage Models
-
List installed models:
ollama list -
Pull a specific model without running it:
ollama pull llama2See ollama.com/library for all available models.
-
Remove a model to free disk space:
ollama rm gemma4
Is It Worth It?
Yes, if you:
- Want zero API costs after initial hardware investment.
- Run inference frequently and need predictable billing.
- Care about privacy and keeping data local.
- Have a decent GPU or CPU (even a 4-core CPU works for smaller models).
No, if you:
- Need the absolute best performance or latest proprietary models (GPT-4, Claude).
- Don’t have spare hardware or don’t want to manage it.
- Only use AI occasionally (cloud APIs are cheaper per-query).
For most homelabs and self-hosted setups, Ollama is the obvious play. You own the hardware, you own the model, you own the data. Google’s cloud strategy doesn’t apply here.
Sources
Gear used in this build
* Affiliate links — I earn a small commission at no cost to you. It's gear I use and would genuinely recommend. See the full disclosure.
Related video
New self-hosted AI & homelab shorts, daily.
Subscribe on YouTubeRelated guides
Run Claude Code and Codex Locally with OtoDock
Deploy OtoDock on your server to get local code generation without cloud API costs. Self-hosted alternative to Claude Code and GitHub Copilot.
Run Claude Code Locally with OtoDock
Deploy OtoDock on your own hardware to run Claude Code and Codex as local AI agents without paying per API call. Full self-hosted setup.
Self-host Claude Code agents with OtoDock
Run Claude's code execution engine on your own hardware. Replace the SaaS with OtoDock—a self-hosted agent framework that keeps your inference and execution local.