Run Claude Code and Codex Locally with OtoDock
🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.
OtoDock brings local code generation to your homelab by running inference engines on your own hardware. No more paying per-token to Anthropic or OpenAI for code completion—host it yourself and keep your code private.
Prerequisites
You need a server with Docker installed and enough VRAM to run a code model. A GPU helps significantly; 8GB VRAM minimum for smaller models, 24GB+ for larger ones. OtoDock supports multiple inference backends.
Step 1: Clone the OtoDock Repository
git clone https://github.com/otodock/otodock.git
cd otodock
Step 2: Review Configuration
Examine the included documentation and configuration templates. OtoDock typically uses environment variables and Docker Compose for setup. Check the repo for example .env files or docker-compose.yml that specify which models and backends are supported.
Step 3: Set Up Docker Compose
OtoDock is designed to run in containers. Use the provided Compose file or create your own based on the repository structure:
docker-compose up -d
This pulls the OtoDock image and starts the service. Verify it’s running:
docker ps | grep otodock
Step 4: Configure Your Code Editor or IDE
OtoDock exposes an API endpoint (typically on localhost or a local IP). Configure your editor—VS Code, JetBrains IDEs, or Vim plugins—to point to your OtoDock instance instead of the cloud service. The exact endpoint and authentication details are in the repository’s documentation.
Step 5: Test Code Generation
Open a file in your editor and trigger code completion. Your request now goes to your local OtoDock instance, not to Anthropic or OpenAI. Latency depends on your hardware; GPU acceleration is strongly recommended.
Step 6: Monitor and Optimize
Watch resource usage:
docker stats otodock
If inference is slow, consider upgrading your GPU, adding more VRAM, or switching to a smaller quantized model. OtoDock’s performance scales directly with your hardware.
Keep It Private
Do not expose OtoDock to the public internet. Run it only on your LAN or behind a VPN. Code generation models can be resource-intensive targets, and your local instance has no rate-limiting or authentication by default. Restrict network access to trusted machines only.
Is It Worth It?
Yes, if you generate code constantly and want to eliminate API costs and latency. The upfront hardware investment (GPU, RAM) pays for itself in a few months if you’d otherwise spend $20–100/month on Copilot or Claude Code subscriptions. The privacy win is real too: your code never leaves your network. The trade-off is that you own the infrastructure—updates, monitoring, and troubleshooting are on you. For a serious homelab, this is a solid addition.
Related video
New self-hosted AI & homelab shorts, daily.
Subscribe on YouTubeRelated guides
Run Claude Code Locally with OtoDock
Deploy OtoDock on your own hardware to run Claude Code and Codex as local AI agents without paying per API call. Full self-hosted setup.
Self-host Claude Code agents with OtoDock
Run Claude's code execution engine on your own hardware. Replace the SaaS with OtoDock—a self-hosted agent framework that keeps your inference and execution local.
Run LLMs on Minimal Hardware with llama.cpp
Use llama.cpp to run quantized language models on constrained hardware. Optimize inference with GGUF formats and CPU backends for your homelab.