How to Deploy GLM-5 Locally with SGLang: A Complete Setup Guide
Deploy GLM-5 locally by installing SGLang v0.5.13.post1+, downloading the checkpoint from HuggingFace, and running sglang serve with --model-type glm_moe_dsa.
The zai-org/GLM-5 repository provides official support for serving the GLM-5 series (including the GLM-S variant) using the SGLang inference engine. SGLang exposes the model through an OpenAI-compatible HTTP endpoint, making it straightforward to integrate into existing applications. This guide covers the complete workflow from environment setup to querying the local endpoint.
Prerequisites
Before deploying GLM-5 with SGLang, ensure your environment meets the following requirements:
- Python 3.9+ (tested with Python 3.10)
- Git and Git LFS (for cloning large model files)
- CUDA 11.8 or compatible GPU driver (optional but recommended for GPU acceleration)
- At least 30 GB of disk space for the BF16 checkpoint
Verify your GPU setup with:
nvidia-smi
Install SGLang
SGLang must be installed from source at version v0.5.13.post1 or later, as earlier versions lack the GLM-5-specific kernels and glm_moe_dsa support required for the model.
Clone the repository and install in editable mode:
git clone https://github.com/sgl-project/sglang.git
cd sglang
git checkout v0.5.13.post1
pip install -e .
This installs the inference engine along with its dependencies including PyTorch, Transformers, and FlashAttention.
Download the GLM-S Checkpoint
GLM-S (the "small" variant of the GLM-5 series) is hosted on HuggingFace under the zai-org organization. Use Git LFS to download the full checkpoint:
mkdir -p models/glm-s
cd models/glm-s
git lfs install
git clone https://huggingface.co/zai-org/GLM-S .
For FP8 precision (lower memory footprint), use the FP8 variant instead:
git clone https://huggingface.co/zai-org/GLM-S-FP8 .
Verify the download contains config.json, pytorch_model.bin, and tokenizer files. The BF16 checkpoint requires approximately 30 GB of storage.
Launch the SGLang Server
Start the inference server by specifying the model path and the required GLM-5 model type. In the zai-org/GLM-5 implementation, you must use --model-type glm_moe_dsa to enable the correct kernels.
export GLM_MODEL_PATH=$(pwd)/models/glm-s
export CUDA_VISIBLE_DEVICES=0
sglang serve \
--model-path $GLM_MODEL_PATH \
--model-type glm_moe_dsa \
--dtype bf16 \
--port 8080 \
--max-num-batches 4
For FP8 checkpoints, change --dtype to fp8. The server prints a confirmation when ready:
[INFO] OpenAI-compatible endpoint listening on http://0.0.0.0:8080/v1
Test the Endpoint
Send requests to the local server using any OpenAI-compatible client. The following Python example queries the GLM-5 model with the optional reasoning_effort parameter:
import openai
openai.api_base = "http://127.0.0.1:8080/v1"
openai.api_key = "unused"
response = openai.ChatCompletion.create(
model="glm-s",
messages=[{"role": "user", "content": "Explain the difference between BFS and DFS"}],
temperature=0.7,
max_tokens=256,
reasoning_effort="high",
)
print(response.choices[0].message["content"])
The server accepts standard chat completion parameters and returns generated text through the OpenAI-compatible JSON format.
Configure Inference Parameters
GLM-5 supports two runtime knobs that control the reasoning behavior:
reasoning_effort: Set to"max"(default) or"high". The"high"value allocates a larger context budget for speculative decoding, improving answer quality at the cost of increased latency.enable_thinking: Set totrueorfalse. Disabling thinking (false) skips the internal reasoning phase for faster responses.
Pass these as top-level parameters in your request payload alongside standard fields like temperature and max_tokens.
Deploy on Ascend NPU
For Ascend hardware deployments, the zai-org/GLM-5 repository provides specific compatibility through the SGLang binary. Set the Ascend device variable before launching:
export ASCEND_DEVICE_ID=0
sglang serve \
--model-path $GLM_MODEL_PATH \
--model-type glm_moe_dsa \
--device ascend
Detailed Ascend-specific instructions are available in the repository's example/ascend.md file.
Summary
- Version requirement: SGLang v0.5.13.post1+ is mandatory for GLM-5 support due to the required
glm_moe_dsakernels. - Model type: Always specify
--model-type glm_moe_dsawhen serving GLM-5 checkpoints. - Data types: Use
--dtype bf16for standard checkpoints or--dtype fp8for the FP8 variant. - Endpoint: The server exposes an OpenAI-compatible API at
http://localhost:8080/v1. - Reasoning control: Adjust
reasoning_effort("max"/"high") andenable_thinkingto trade off between speed and response quality.
Frequently Asked Questions
What is the minimum SGLang version required for GLM-5?
You need SGLang v0.5.13.post1 or later. According to the zai-org/GLM-5 README, this version contains the GLM-5-specific kernels and glm_moe_dsa support necessary to run the model architecture correctly.
Can I run GLM-5 without a GPU?
Yes, but inference will be significantly slower. SGLang supports CPU-only deployment, though the zai-org/GLM-5 documentation recommends CUDA 11.8+ for production use. Remove the CUDA_VISIBLE_DEVICES environment variable and ensure you have sufficient system RAM (60+ GB recommended) to load the model weights.
How do I switch between high-quality and fast inference modes?
Set the reasoning_effort parameter in your API requests. Use "max" for the highest quality with standard latency, or "high" for improved reasoning through speculative decoding. To maximize speed, set enable_thinking to false, which disables the internal reasoning phase entirely.
Where can I find the Ascend NPU deployment instructions?
The zai-org/GLM-5 repository includes Ascend-specific documentation in example/ascend.md. This file contains the exact environment variables and device flags needed for Huawei Ascend hardware, including the required --device ascend flag and ASCEND_DEVICE_ID configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →