# How to Deploy GLM-5 Locally with SGLang: A Complete Setup Guide

> Deploy GLM-5 locally with SGLang. Follow this complete setup guide to install SGLang, download the checkpoint, and run your model efficiently.

- Repository: [Z.ai/GLM-5](https://github.com/zai-org/GLM-5)
- Tags: getting-started
- Published: 2026-06-19

---

**Deploy GLM-5 locally by installing SGLang v0.5.13.post1+, downloading the checkpoint from HuggingFace, and running `sglang serve` with `--model-type glm_moe_dsa`.**

The `zai-org/GLM-5` repository provides official support for serving the GLM-5 series (including the GLM-S variant) using the SGLang inference engine. SGLang exposes the model through an OpenAI-compatible HTTP endpoint, making it straightforward to integrate into existing applications. This guide covers the complete workflow from environment setup to querying the local endpoint.

## Prerequisites

Before deploying GLM-5 with SGLang, ensure your environment meets the following requirements:

- **Python 3.9+** (tested with Python 3.10)
- **Git** and **Git LFS** (for cloning large model files)
- **CUDA 11.8** or compatible GPU driver (optional but recommended for GPU acceleration)
- At least **30 GB** of disk space for the BF16 checkpoint

Verify your GPU setup with:

```bash
nvidia-smi

```

## Install SGLang

SGLang must be installed from source at version **v0.5.13.post1** or later, as earlier versions lack the GLM-5-specific kernels and `glm_moe_dsa` support required for the model.

Clone the repository and install in editable mode:

```bash
git clone https://github.com/sgl-project/sglang.git
cd sglang
git checkout v0.5.13.post1
pip install -e .

```

This installs the inference engine along with its dependencies including PyTorch, Transformers, and FlashAttention.

## Download the GLM-S Checkpoint

GLM-S (the "small" variant of the GLM-5 series) is hosted on HuggingFace under the `zai-org` organization. Use Git LFS to download the full checkpoint:

```bash
mkdir -p models/glm-s
cd models/glm-s
git lfs install
git clone https://huggingface.co/zai-org/GLM-S .

```

For FP8 precision (lower memory footprint), use the FP8 variant instead:

```bash
git clone https://huggingface.co/zai-org/GLM-S-FP8 .

```

Verify the download contains [`config.json`](https://github.com/zai-org/GLM-5/blob/main/config.json), `pytorch_model.bin`, and tokenizer files. The BF16 checkpoint requires approximately 30 GB of storage.

## Launch the SGLang Server

Start the inference server by specifying the model path and the required GLM-5 model type. In the `zai-org/GLM-5` implementation, you must use `--model-type glm_moe_dsa` to enable the correct kernels.

```bash
export GLM_MODEL_PATH=$(pwd)/models/glm-s
export CUDA_VISIBLE_DEVICES=0

sglang serve \
  --model-path $GLM_MODEL_PATH \
  --model-type glm_moe_dsa \
  --dtype bf16 \
  --port 8080 \
  --max-num-batches 4

```

For FP8 checkpoints, change `--dtype` to `fp8`. The server prints a confirmation when ready:

```

[INFO] OpenAI-compatible endpoint listening on http://0.0.0.0:8080/v1

```

## Test the Endpoint

Send requests to the local server using any OpenAI-compatible client. The following Python example queries the GLM-5 model with the optional `reasoning_effort` parameter:

```python
import openai

openai.api_base = "http://127.0.0.1:8080/v1"
openai.api_key = "unused"

response = openai.ChatCompletion.create(
    model="glm-s",
    messages=[{"role": "user", "content": "Explain the difference between BFS and DFS"}],
    temperature=0.7,
    max_tokens=256,
    reasoning_effort="high",
)

print(response.choices[0].message["content"])

```

The server accepts standard chat completion parameters and returns generated text through the OpenAI-compatible JSON format.

## Configure Inference Parameters

GLM-5 supports two runtime knobs that control the reasoning behavior:

- **`reasoning_effort`**: Set to `"max"` (default) or `"high"`. The `"high"` value allocates a larger context budget for speculative decoding, improving answer quality at the cost of increased latency.
- **`enable_thinking`**: Set to `true` or `false`. Disabling thinking (`false`) skips the internal reasoning phase for faster responses.

Pass these as top-level parameters in your request payload alongside standard fields like `temperature` and `max_tokens`.

## Deploy on Ascend NPU

For Ascend hardware deployments, the `zai-org/GLM-5` repository provides specific compatibility through the SGLang binary. Set the Ascend device variable before launching:

```bash
export ASCEND_DEVICE_ID=0
sglang serve \
  --model-path $GLM_MODEL_PATH \
  --model-type glm_moe_dsa \
  --device ascend

```

Detailed Ascend-specific instructions are available in the repository's [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) file.

## Summary

- **Version requirement**: SGLang v0.5.13.post1+ is mandatory for GLM-5 support due to the required `glm_moe_dsa` kernels.
- **Model type**: Always specify `--model-type glm_moe_dsa` when serving GLM-5 checkpoints.
- **Data types**: Use `--dtype bf16` for standard checkpoints or `--dtype fp8` for the FP8 variant.
- **Endpoint**: The server exposes an OpenAI-compatible API at `http://localhost:8080/v1`.
- **Reasoning control**: Adjust `reasoning_effort` (`"max"`/`"high"`) and `enable_thinking` to trade off between speed and response quality.

## Frequently Asked Questions

### What is the minimum SGLang version required for GLM-5?

You need SGLang v0.5.13.post1 or later. According to the `zai-org/GLM-5` README, this version contains the GLM-5-specific kernels and `glm_moe_dsa` support necessary to run the model architecture correctly.

### Can I run GLM-5 without a GPU?

Yes, but inference will be significantly slower. SGLang supports CPU-only deployment, though the `zai-org/GLM-5` documentation recommends CUDA 11.8+ for production use. Remove the `CUDA_VISIBLE_DEVICES` environment variable and ensure you have sufficient system RAM (60+ GB recommended) to load the model weights.

### How do I switch between high-quality and fast inference modes?

Set the `reasoning_effort` parameter in your API requests. Use `"max"` for the highest quality with standard latency, or `"high"` for improved reasoning through speculative decoding. To maximize speed, set `enable_thinking` to `false`, which disables the internal reasoning phase entirely.

### Where can I find the Ascend NPU deployment instructions?

The `zai-org/GLM-5` repository includes Ascend-specific documentation in [`example/ascend.md`](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md). This file contains the exact environment variables and device flags needed for Huawei Ascend hardware, including the required `--device ascend` flag and `ASCEND_DEVICE_ID` configuration.