Build AI Prompts for Business with Ollama
🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.
You have good models sitting on your hardware right now. Instead of paying per-token to OpenAI or Claude for every business email and sales pitch, run Ollama locally and generate unlimited prompts. No cloud vendor lock-in, no usage limits, no surprise bills.
This guide walks you through setting up Ollama, loading a capable model, and building a simple prompt system for common business tasks—customer outreach, product descriptions, internal comms, whatever you need.
Install Ollama
Ollama runs on macOS, Windows, and Linux. Pick your platform:
macOS:
- Run the installer script:
curl -fsSL https://ollama.com/install.sh | sh
Or download the DMG manually from https://ollama.com/download/Ollama.dmg and drag it to Applications.
Linux:
- Run the installer:
curl -fsSL https://ollama.com/install.sh | sh
For manual install, see the Linux docs.
Windows:
- Run the PowerShell installer:
irm https://ollama.com/install.ps1 | iex
Or download the exe from https://ollama.com/download/OllamaSetup.exe.
- Verify the install by opening a terminal and running:
ollama
You’ll see the Ollama prompt or launcher.
Pull a Business-Ready Model
Ollama’s model library includes Gemma 4, Mistral, Llama, and others. For business prompts, Gemma 4 is solid—fast, capable, and runs on modest hardware.
- Pull the model:
ollama pull gemma4
This downloads the model weights (a few GB depending on the model). On first run, it may take a few minutes.
- Test it interactively:
ollama run gemma4
You’ll get a chat prompt. Type a test prompt—e.g., “Write a short sales email for a SaaS product”—and hit Enter. Type exit to quit.
Set Up a REST API for Programmatic Access
Ollama runs a REST API on localhost:11434 by default. This lets you call it from scripts, web apps, or other tools without the interactive CLI.
- Start Ollama in the background (it usually stays running after install):
ollama serve
On macOS and Windows, Ollama often auto-starts. On Linux, you may need to run it manually or set up a systemd service.
- Test the API with curl:
curl http://localhost:11434/api/chat -d '{
"model": "gemma4",
"messages": [{
"role": "user",
"content": "Write a professional email subject line for a product launch"
}],
"stream": false
}'
You’ll get a JSON response with the model’s output.
Build a Prompt Library Script
Create a simple Python script to manage business prompts. This gives you a repeatable, version-controlled way to generate content.
- Install the Python client:
pip install ollama
- Create a file called
business_prompts.py:
from ollama import chat
# Define reusable prompt templates
prompts = {
"sales_email": "Write a professional cold email to a potential customer interested in {product}. Keep it under 100 words. Focus on value, not features.",
"product_description": "Write a product description for {product} that highlights benefits for {audience}. Use clear, conversational language.",
"internal_memo": "Write an internal memo announcing {announcement} to the team. Keep it concise and actionable.",
"social_post": "Write a LinkedIn post about {topic} that is engaging and professional. Include a call-to-action.",
}
def generate_prompt(prompt_type, **kwargs):
"""Generate content using a prompt template."""
if prompt_type not in prompts:
print(f"Unknown prompt type: {prompt_type}")
return
template = prompts[prompt_type]
filled_prompt = template.format(**kwargs)
response = chat(model='gemma4', messages=[
{
'role': 'user',
'content': filled_prompt,
},
])
return response.message.content
if __name__ == "__main__":
# Example: generate a sales email
result = generate_prompt(
"sales_email",
product="cloud backup software"
)
print("Sales Email:")
print(result)
print("\n" + "="*50 + "\n")
# Example: generate a product description
result = generate_prompt(
"product_description",
product="AI-powered scheduling tool",
audience="small business owners"
)
print("Product Description:")
print(result)
- Run it:
python business_prompts.py
You’ll get generated content. Edit the templates or add new ones as needed.
Keep It Private: Network Security
Critical: Ollama listens on localhost:11434 by default, which is safe for local use. If you expose it to the internet, anyone can use your hardware to run inference.
-
Local only: Never expose port 11434 to the public internet. If you need remote access, use a VPN (WireGuard, Tailscale) to tunnel through securely.
-
Firewall check: Verify Ollama is not listening on
0.0.0.0:
netstat -tuln | grep 11434
You should see 127.0.0.1:11434 (local) or ::1:11434 (IPv6 local), not 0.0.0.0:11434.
- If you need remote access: Set up a VPN first, then access Ollama only through that tunnel. Never open it directly.
Integrate with Web or Automation Tools
Ollama’s REST API works with any tool that can make HTTP requests.
JavaScript (Node.js):
npm install ollama
import { Ollama } from "ollama";
const ollama = new Ollama({ host: "http://localhost:11434" });
const response = await ollama.chat({
model: "gemma4",
messages: [{ role: "user", content: "Write a product tagline for a fitness app" }],
});
console.log(response.message.content);
Bash/curl:
You can call the API directly from shell scripts or automation tools:
curl -X POST http://localhost:11434/api/chat \
-H "Content-Type: application/json" \
-d '{
"model": "gemma4",
"messages": [{"role": "user", "content": "Write a customer testimonial for a productivity app"}],
"stream": false
}' | jq '.message.content'
Benchmark and Tune
Different models have different speeds and quality. Experiment locally:
- Try other models from the library:
ollama pull mistral
ollama pull llama2
-
Test them with the same prompt and compare output quality and speed.
-
Stick with what works for your use case. Smaller models (Gemma 4) are faster; larger ones (Llama 2 70B) are more capable but need more VRAM.
Is It Worth It?
Yes, if you generate business content regularly. You save money on API calls, keep data off third-party servers, and avoid rate limits. The trade-off: you manage the hardware and model updates yourself. For a small team or solo business, running Ollama locally is faster and cheaper than cloud APIs over time. If you generate only a few prompts a month, the API overhead is negligible and cloud might be simpler. For anything in between—daily emails, weekly posts, ongoing customer outreach—Ollama pays for itself in token costs alone within weeks.
Sources
Gear used in this build
* Affiliate links — I earn a small commission at no cost to you. It's gear I use and would genuinely recommend. See the full disclosure.
Related video
New self-hosted AI & homelab shorts, daily.
Subscribe on YouTubeRelated guides
Run Claude Code and Codex Locally with OtoDock
Deploy OtoDock on your server to get local code generation without cloud API costs. Self-hosted alternative to Claude Code and GitHub Copilot.
Run Claude Code Locally with OtoDock
Deploy OtoDock on your own hardware to run Claude Code and Codex as local AI agents without paying per API call. Full self-hosted setup.
Self-host Claude Code agents with OtoDock
Run Claude's code execution engine on your own hardware. Replace the SaaS with OtoDock—a self-hosted agent framework that keeps your inference and execution local.