aihomelabprivacy

Run GLM 5.2 on AMD MI355X with vLLM

🤖 Researched and drafted automatically from the official docs, and reviewed before publishing. Commands are taken from the source projects — but always sanity-check before running anything on your own hardware.

vLLM is a serving engine built for speed and cost efficiency. It handles the hard parts of LLM inference—memory management, batching, and kernel optimization—so you run models locally without compromise. On AMD MI355X hardware, vLLM’s PagedAttention and HIP kernel support unlock the throughput you need for real workloads.

This guide walks you through deploying Alibaba’s GLM 5.2 model on MI355X with vLLM, targeting 2600+ tokens per second. You’ll set up the runtime, download the model, and serve it over an OpenAI-compatible API on your LAN.

Prerequisites

  • AMD MI355X GPU(s) with ROCm support
  • Linux host (Ubuntu 22.04 or later recommended)
  • Python 3.10+
  • 200+ GB disk space for model weights
  • Network isolated from the public internet (VPN or LAN only)

1. Install ROCm and vLLM

Start by installing ROCm for your MI355X. Follow AMD’s official ROCm installation guide for your Linux distribution. Verify the installation:

rocm-smi

You should see your MI355X listed with memory info. Then install vLLM using the recommended uv tool:

uv pip install vllm

If you don’t have uv installed, grab it from https://docs.astral.sh/uv/ or fall back to pip:

pip install vllm

2. Download the GLM 5.2 Model

GLM 5.2 is available on Hugging Face. You’ll need git-lfs to fetch the full model weights:

sudo apt-get install git-lfs
git lfs install

Clone the model repository to a local directory:

git clone https://huggingface.co/THUDM/glm-5-2 /path/to/glm-5-2

Replace /path/to/glm-5-2 with a location on a disk with sufficient space. This will take time depending on your network.

3. Start vLLM Server

Launch the vLLM server with the GLM 5.2 model. The --device flag tells vLLM to use HIP (AMD GPU support):

vllm serve /path/to/glm-5-2 \
  --device amd \
  --dtype float16 \
  --max-model-len 4096

Break down the flags:

  • --device amd: Use AMD GPU via HIP
  • --dtype float16: Run in FP16 precision for speed and memory efficiency
  • --max-model-len 4096: Limit context to 4096 tokens (adjust based on your MI355X memory)

The server will bind to http://127.0.0.1:8000 by default. You’ll see logs about model loading and tensor parallelism.

4. Verify the Server

In another terminal, test the API with a simple completion request:

curl http://127.0.0.1:8000/v1/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5-2",
    "prompt": "Explain quantum computing in one sentence.",
    "max_tokens": 100
  }'

You should get a JSON response with generated text. If the request hangs, check that vLLM is still running and no GPU errors occurred.

5. Configure for Local Network Access

If you want other machines on your LAN to access the server, bind to your local IP instead of localhost. First, find your LAN IP:

ip addr | grep "inet "

Note the IP in the 192.168.x.x or 10.x.x.x range. Stop the current vLLM process (Ctrl+C) and restart with:

vllm serve /path/to/glm-5-2 \
  --device amd \
  --dtype float16 \
  --max-model-len 4096 \
  --host 192.168.x.x \
  --port 8000

Replace 192.168.x.x with your actual LAN IP. Other machines on the network can now query http://192.168.x.x:8000.

Critical: Never expose this port to the internet. Use a firewall rule or VPN to restrict access to trusted networks only:

sudo ufw allow from 192.168.0.0/24 to any port 8000
sudo ufw deny from any to any port 8000

6. Tune for 2600+ Tokens/Second

PagedAttention is already enabled in vLLM by default. To push throughput further on MI355X:

  • Increase --max-num-seqs to allow more concurrent requests (default is 256):

    vllm serve /path/to/glm-5-2 \
      --device amd \
      --dtype float16 \
      --max-model-len 4096 \
      --max-num-seqs 512
  • Use --enable-chunked-prefill to split long prompts into chunks, improving batching:

    vllm serve /path/to/glm-5-2 \
      --device amd \
      --dtype float16 \
      --max-model-len 4096 \
      --enable-chunked-prefill
  • Monitor GPU utilization with rocm-smi in another terminal to ensure you’re hitting peak throughput without memory overflow.

7. Integrate with Applications

vLLM’s API is OpenAI-compatible. Any client that speaks the OpenAI API can call your local GLM 5.2 server. Example with Python:

from openai import OpenAI

client = OpenAI(
    api_key="not-needed",
    base_url="http://192.168.x.x:8000/v1"
)

response = client.completions.create(
    model="glm-5-2",
    prompt="Write a haiku about AI.",
    max_tokens=50
)

print(response.choices[0].text)

8. Run as a Systemd Service (Optional)

To keep vLLM running across reboots, create a systemd service. Save this as /etc/systemd/system/vllm-glm.service:

[Unit]
Description=vLLM GLM 5.2 Server
After=network.target

[Service]
Type=simple
User=your-username
ExecStart=/home/your-username/.local/bin/vllm serve /path/to/glm-5-2 --device amd --dtype float16 --max-model-len 4096 --host 192.168.x.x --port 8000
Restart=on-failure
RestartSec=10

[Install]
WantedBy=multi-user.target

Replace your-username and paths with your actual setup. Then enable and start:

sudo systemctl daemon-reload
sudo systemctl enable vllm-glm
sudo systemctl start vllm-glm
sudo systemctl status vllm-glm

Is It Worth It?

Yes, if you need local inference with real throughput. GLM 5.2 on MI355X hits 2600+ tokens/sec because vLLM’s PagedAttention, continuous batching, and HIP kernel optimization actually work. You get a model that stays on your hardware, zero API costs, and no latency spikes from cloud calls. The setup takes an afternoon, and MI355X hardware is cheaper than equivalent NVIDIA GPUs. Run it behind a firewall on your LAN, and you have a private, fast, cost-effective inference engine.

New self-hosted AI & homelab shorts, daily.

Subscribe on YouTube

Related guides

← All guides