Run Qwen 80B on Mac with 4.3GB RAM using llama.cpp
🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.
Paying monthly for cloud inference when you own the hardware is leaving money on the table. llama.cpp lets you run Qwen 80B on a Mac with just 4.3GB of RAM through aggressive quantization, or even Qwen 35B on an iPhone. This is a complete local inference setup with no external dependencies.
Download and install llama.cpp
- Visit https://llama.app and follow the installation instructions for macOS, or clone and build from source:
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
make
The build will automatically detect and optimize for Apple Silicon using Metal, ARM NEON, and Accelerate frameworks.
- Verify the installation:
./llama-cli --version
Download a quantized Qwen model
- Pull a GGUF-quantized Qwen model directly from Hugging Face. For 4.3GB memory on a Mac, use a 3-bit or 4-bit quantized variant:
./llama-cli -hf ggml-org/Qwen3.5-0.8B-GGUF
For larger Qwen models (35B, 80B), search Hugging Face for quantized GGUF versions. Look for repos tagged with gguf and check the quantization level (Q3_K_M, Q4_K_M, etc.) to fit your RAM.
- Models download to
~/.cache/huggingface/hubby default. Verify the download completed:
ls -lh ~/.cache/huggingface/hub/
Run inference in the CLI
- Start an interactive chat session:
./llama-cli -hf ggml-org/Qwen3.5-0.8B-GGUF
-
Type your prompt and press Enter. The model runs entirely on your Mac’s GPU via Metal acceleration.
-
Control resource usage with these flags (check the build guide for the full list):
-nsets max tokens to generate-csets context window size-tsets number of CPU threads
Launch an OpenAI-compatible API server
- Start the server to expose inference over HTTP:
./llama-serve -hf ggml-org/Qwen3.5-0.8B-GGUF
The server listens on http://127.0.0.1:8000 by default.
- Test the API from another terminal:
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen",
"messages": [{"role": "user", "content": "What is 2+2?"}],
"temperature": 0.7
}'
- The API is compatible with OpenAI client libraries. Point any existing tools (n8n, Ollama integrations, etc.) to
http://127.0.0.1:8000/v1.
Optimize for your hardware
-
If you’re running on an M-series Mac with less than 8GB unified memory, use a smaller quantization. Check the model card on Hugging Face for RAM requirements per quantization level.
-
For iPhone or iPad, build and run the llama.cpp iOS app. Search the repository for iOS build instructions in
docs/. -
Monitor memory usage while running inference:
top -l 1 | grep llama
If the process gets killed, you’ve exceeded available RAM. Step down to a smaller model or lower quantization.
Keep it local
Do not expose the llama-serve API to the public internet. This runs on your LAN only. If you need remote access, use a VPN tunnel back to your home network or run it behind a reverse proxy with authentication on your private network.
Is it worth it?
Yes, if you have a Mac with at least 8GB of unified memory and you’re currently paying for API inference. One-time setup, zero monthly costs, instant local responses, and your prompts never leave your machine. The quantized models trade a small amount of quality for speed and memory efficiency—acceptable for most tasks like summarization, coding help, and content generation. If you need state-of-the-art accuracy on complex reasoning, cloud inference is still your move.
Sources
Gear used in this build
* Affiliate links — I earn a small commission at no cost to you. It's gear I use and would genuinely recommend. See the full disclosure.
Related video
New self-hosted AI & homelab shorts, daily.
Subscribe on YouTubeRelated guides
Run Claude Code and Codex Locally with OtoDock
Deploy OtoDock on your server to get local code generation without cloud API costs. Self-hosted alternative to Claude Code and GitHub Copilot.
Run Claude Code Locally with OtoDock
Deploy OtoDock on your own hardware to run Claude Code and Codex as local AI agents without paying per API call. Full self-hosted setup.
Self-host Claude Code agents with OtoDock
Run Claude's code execution engine on your own hardware. Replace the SaaS with OtoDock—a self-hosted agent framework that keeps your inference and execution local.