aihomelabprivacy

Run Ollama Locally: Skip Google's Cloud, Own Your AI

🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.

Google’s latest AI updates push cloud inference and subscription pricing. Ollama flips that script: run open-source models directly on your hardware, no API key, no monthly bill, no data leaving your network.

This guide gets you chatting with Gemma 4 or any model from the Ollama library in under 10 minutes. Everything runs locally behind your firewall.

Install Ollama

  1. On Linux, run the official installer:

    curl -fsSL https://ollama.com/install.sh | sh
  2. On macOS, use the same command or download the DMG manually from ollama.com/download.

  3. On Windows, run the PowerShell installer:

    irm https://ollama.com/install.ps1 | iex

    Or download OllamaSetup.exe from ollama.com/download.

  4. Verify installation:

    ollama --version

Run Your First Model

  1. Start Ollama and download Gemma 4 (or any model from ollama.com/library):

    ollama run gemma4

    The first run downloads the model—size varies by model, usually 4–40 GB. Be patient.

  2. Once loaded, you’ll see a prompt. Type your question:

    >>> Why is the sky blue?

    The model responds locally. No internet required after download.

  3. Exit the chat with Ctrl+D or type /bye.

Access via REST API

Ollama exposes a local REST API on port 11434 by default. This lets you integrate models into scripts, web apps, or other tools.

  1. Start the Ollama service (runs in background on Linux/macOS):

    ollama serve

    On macOS/Windows, the service starts automatically.

  2. Query the API from another terminal or script:

    curl http://localhost:11434/api/chat -d '{
      "model": "gemma4",
      "messages": [{
        "role": "user",
        "content": "Explain machine learning in one sentence."
      }],
      "stream": false
    }'

Use Ollama with Python

  1. Install the Python client:

    pip install ollama
  2. Write a simple script (chat.py):

    from ollama import chat
    
    response = chat(model='gemma4', messages=[
      {
        'role': 'user',
        'content': 'What is a neural network?',
      },
    ])
    print(response.message.content)
  3. Run it:

    python chat.py

Use Ollama with JavaScript

  1. Install the npm package:

    npm i ollama
  2. Create a script (chat.js):

    import ollama from "ollama";
    
    const response = await ollama.chat({
      model: "gemma4",
      messages: [{ role: "user", content: "What is a neural network?" }],
    });
    console.log(response.message.content);
  3. Run it:

    node chat.js

Connect a Web UI

Ollama is headless by default. For a ChatGPT-like interface, pair it with Open WebUI or another community frontend.

  1. Using Docker (if you have Docker installed):

    docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway ghcr.io/open-webui/open-webui:latest

    Open http://localhost:3000 in your browser.

  2. Or run Ollama and Open WebUI on the same machine without Docker—see the Open WebUI repo for native setup.

  3. Point the UI to your local Ollama instance at http://localhost:11434.

Network Security

Important: By default, Ollama listens only on localhost (127.0.0.1) and is not accessible from other machines. Keep it that way. If you need to expose it to your LAN:

  • Set OLLAMA_HOST=0.0.0.0:11434 as an environment variable only if your machine is behind a router firewall.
  • Never expose port 11434 to the public internet. Always keep Ollama on your LAN or behind a VPN.
  • If you need remote access, use a VPN (WireGuard, Tailscale, OpenVPN) instead of opening the port.

Manage Models

  • List installed models:

    ollama list
  • Pull a specific model without running it:

    ollama pull llama2

    See ollama.com/library for all available models.

  • Remove a model to free disk space:

    ollama rm gemma4

Is It Worth It?

Yes, if you:

  • Want zero API costs after initial hardware investment.
  • Run inference frequently and need predictable billing.
  • Care about privacy and keeping data local.
  • Have a decent GPU or CPU (even a 4-core CPU works for smaller models).

No, if you:

  • Need the absolute best performance or latest proprietary models (GPT-4, Claude).
  • Don’t have spare hardware or don’t want to manage it.
  • Only use AI occasionally (cloud APIs are cheaper per-query).

For most homelabs and self-hosted setups, Ollama is the obvious play. You own the hardware, you own the model, you own the data. Google’s cloud strategy doesn’t apply here.

Related guides

← All guides