Run GLM 5.2 on AMD MI355X with vLLM
🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.
vLLM is a serving engine built for speed and cost efficiency. It handles the hard parts of LLM inference—memory management, batching, and kernel optimization—so you run models locally without compromise. On AMD MI355X hardware, vLLM’s PagedAttention and HIP kernel support unlock the throughput you need for real workloads.
This guide walks you through deploying Alibaba’s GLM 5.2 model on MI355X with vLLM, targeting 2600+ tokens per second. You’ll set up the runtime, download the model, and serve it over an OpenAI-compatible API on your LAN.
Prerequisites
- AMD MI355X GPU(s) with ROCm support
- Linux host (Ubuntu 22.04 or later recommended)
- Python 3.10+
- 200+ GB disk space for model weights
- Network isolated from the public internet (VPN or LAN only)
1. Install ROCm and vLLM
Start by installing ROCm for your MI355X. Follow AMD’s official ROCm installation guide for your Linux distribution. Verify the installation:
rocm-smi
You should see your MI355X listed with memory info. Then install vLLM using the recommended uv tool:
uv pip install vllm
If you don’t have uv installed, grab it from https://docs.astral.sh/uv/ or fall back to pip:
pip install vllm
2. Download the GLM 5.2 Model
GLM 5.2 is available on Hugging Face. You’ll need git-lfs to fetch the full model weights:
sudo apt-get install git-lfs
git lfs install
Clone the model repository to a local directory:
git clone https://huggingface.co/THUDM/glm-5-2 /path/to/glm-5-2
Replace /path/to/glm-5-2 with a location on a disk with sufficient space. This will take time depending on your network.
3. Start vLLM Server
Launch the vLLM server with the GLM 5.2 model. The --device flag tells vLLM to use HIP (AMD GPU support):
vllm serve /path/to/glm-5-2 \
--device amd \
--dtype float16 \
--max-model-len 4096
Break down the flags:
--device amd: Use AMD GPU via HIP--dtype float16: Run in FP16 precision for speed and memory efficiency--max-model-len 4096: Limit context to 4096 tokens (adjust based on your MI355X memory)
The server will bind to http://127.0.0.1:8000 by default. You’ll see logs about model loading and tensor parallelism.
4. Verify the Server
In another terminal, test the API with a simple completion request:
curl http://127.0.0.1:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-2",
"prompt": "Explain quantum computing in one sentence.",
"max_tokens": 100
}'
You should get a JSON response with generated text. If the request hangs, check that vLLM is still running and no GPU errors occurred.
5. Configure for Local Network Access
If you want other machines on your LAN to access the server, bind to your local IP instead of localhost. First, find your LAN IP:
ip addr | grep "inet "
Note the IP in the 192.168.x.x or 10.x.x.x range. Stop the current vLLM process (Ctrl+C) and restart with:
vllm serve /path/to/glm-5-2 \
--device amd \
--dtype float16 \
--max-model-len 4096 \
--host 192.168.x.x \
--port 8000
Replace 192.168.x.x with your actual LAN IP. Other machines on the network can now query http://192.168.x.x:8000.
Critical: Never expose this port to the internet. Use a firewall rule or VPN to restrict access to trusted networks only:
sudo ufw allow from 192.168.0.0/24 to any port 8000
sudo ufw deny from any to any port 8000
6. Tune for 2600+ Tokens/Second
PagedAttention is already enabled in vLLM by default. To push throughput further on MI355X:
-
Increase
--max-num-seqsto allow more concurrent requests (default is 256):vllm serve /path/to/glm-5-2 \ --device amd \ --dtype float16 \ --max-model-len 4096 \ --max-num-seqs 512 -
Use
--enable-chunked-prefillto split long prompts into chunks, improving batching:vllm serve /path/to/glm-5-2 \ --device amd \ --dtype float16 \ --max-model-len 4096 \ --enable-chunked-prefill -
Monitor GPU utilization with
rocm-smiin another terminal to ensure you’re hitting peak throughput without memory overflow.
7. Integrate with Applications
vLLM’s API is OpenAI-compatible. Any client that speaks the OpenAI API can call your local GLM 5.2 server. Example with Python:
from openai import OpenAI
client = OpenAI(
api_key="not-needed",
base_url="http://192.168.x.x:8000/v1"
)
response = client.completions.create(
model="glm-5-2",
prompt="Write a haiku about AI.",
max_tokens=50
)
print(response.choices[0].text)
8. Run as a Systemd Service (Optional)
To keep vLLM running across reboots, create a systemd service. Save this as /etc/systemd/system/vllm-glm.service:
[Unit]
Description=vLLM GLM 5.2 Server
After=network.target
[Service]
Type=simple
User=your-username
ExecStart=/home/your-username/.local/bin/vllm serve /path/to/glm-5-2 --device amd --dtype float16 --max-model-len 4096 --host 192.168.x.x --port 8000
Restart=on-failure
RestartSec=10
[Install]
WantedBy=multi-user.target
Replace your-username and paths with your actual setup. Then enable and start:
sudo systemctl daemon-reload
sudo systemctl enable vllm-glm
sudo systemctl start vllm-glm
sudo systemctl status vllm-glm
Is It Worth It?
Yes, if you need local inference with real throughput. GLM 5.2 on MI355X hits 2600+ tokens/sec because vLLM’s PagedAttention, continuous batching, and HIP kernel optimization actually work. You get a model that stays on your hardware, zero API costs, and no latency spikes from cloud calls. The setup takes an afternoon, and MI355X hardware is cheaper than equivalent NVIDIA GPUs. Run it behind a firewall on your LAN, and you have a private, fast, cost-effective inference engine.
New self-hosted AI & homelab shorts, daily.
Subscribe on YouTubeRelated guides
Run Claude Code and Codex Locally with OtoDock
Deploy OtoDock on your server to get local code generation without cloud API costs. Self-hosted alternative to Claude Code and GitHub Copilot.
Run Claude Code Locally with OtoDock
Deploy OtoDock on your own hardware to run Claude Code and Codex as local AI agents without paying per API call. Full self-hosted setup.
Self-host Claude Code agents with OtoDock
Run Claude's code execution engine on your own hardware. Replace the SaaS with OtoDock—a self-hosted agent framework that keeps your inference and execution local.