How to Run GLM 5.2 Models with Supported Routed Quantization Layouts
GLM 5.2 models in ds4 require specific routed-expert quantization layouts—Q2_K, Q4_K, or Q5_K for gate/up tensors and Q2_K, Q4_K, Q5_K, or Q6_K for down tensors—and will fail to load if any other quantization type is detected.
Running GLM 5.2 (also referred to as GLM-Dense-Sparse-Attention) in the antirez/ds4 engine involves understanding its unique mixture-of-experts architecture. Unlike standard dense models, GLM 5.2 combines normal dense tensors with routed MoE (Mixture of Experts) tensors that demand exacting quantization formats. This guide covers how to identify supported layouts, launch models on Metal, CUDA, and ROCm backends, and avoid common configuration errors.
Understanding GLM 5.2 Architecture
Model Family Detection
When ds4 opens a GGUF file, it immediately classifies the model family. In ds4.c, the engine checks for DS4_MODEL_FAMILY_GLM_DSA and branches to GLM-specific shape handling (DS4_SHAPE_GLM52):
if (DS4_MODEL_FAMILY == DS4_MODEL_FAMILY_GLM_DSA)
This classification occurs at line 618 of ds4.c and determines which quantization validation paths activate.
Routed-Expert Quantization Detection
The engine validates routed-expert quantization through ds4_engine_routed_quant_bits, implemented at line 49903 of ds4.c:
int ds4_engine_routed_quant_bits(ds4_engine *e) {
// Scans first layer for gate tensor
// Returns 4 for DS4_TENSOR_Q4_K, otherwise 2 (Q2/K-style)
}
This function automatically detects whether your model uses 2-bit or 4-bit routed experts by inspecting the gate tensor type.
Supported Quantization Layouts
The ds4 README and validation code enforce strict quantization rules for routed experts:
| Tensor Role | Supported Quant Types |
|---|---|
| Gate / Up | Q2_K, Q4_K, Q5_K |
| Down | Q2_K, Q4_K, Q5_K, Q6_K |
Any GLM GGUF using unsupported types (e.g., Q8_0, Q3_K, IQ3_XXS) triggers a fatal error from tensor_expect_routed_expert with "unsupported routed expert tensor type".
Preparing Supported GLM 5.2 Models
Downloading Pre-Quantized Models
The repository provides helper scripts to fetch validated GLM 5.2 models:
# Q2_K routed layout (smallest, fastest)
./download_model.sh glm-antirez-q2
# Q4_K routed layout (better quality, unsupported on GPU-TP)
./download_model.sh glm-antirez-q4
# IQ2_XXS for gate/up, Q2_K for down (experimental quality)
./download_model.sh glm-antirez-iq2xxs
Verify your downloaded model follows the supported layout using the quality-testing fixtures in gguf-tools/quality-testing/README.md.
Running GLM 5.2 on Metal (macOS)
Full Resident Mode
Load the entire model into GPU memory for maximum speed:
./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
--glm-mtp-timing --temp 0
--glm-mtp-timing enables greedy MTP (Multi-Token Prediction) speculation using the MTP block embedded in the main GGUF. --temp 0 forces deterministic greedy decoding required for this mode.
SSD Streaming Mode
For large models on limited RAM, stream routed experts on-demand:
./ds4 -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
--ssd-streaming --ctx 32768
--ssd-streaming keeps non-routed weights resident while caching routed experts dynamically. The engine automatically computes a routed-expert cache budget based on available system memory.
Running GLM 5.2 on CUDA
Single-GPU Configuration
GLM 5.2 on CUDA must not use tensor parallelism. Launch with:
./ds4 --cuda -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
--gpu-devices 0 --power 100
Critical: The --power 100 flag is required—GLM 5.2 ignores lower power settings per the README's "Power" section.
Multi-GPU Configuration
For distributed inference across multiple GPUs, specify devices without the tensor-parallel flag:
./ds4 --cuda -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
--gpu-devices 0,2,4,6,1,3,5,7 --gpu-vram auto
--gpu-vram auto allows the engine to distribute model layers across the specified device list using conventional placement.
Running GLM 5.2 on ROCm (Strix Halo)
AMD Strix Halo APUs use the ROCm backend with similar streaming support:
./ds4 --rocm -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
--ssd-streaming --ctx 4096
Build with make strix-halo to enable the optimized ROCm kernels for this platform.
GLM 5.2 Command-Line Flags Reference
| Flag | Purpose | Backend |
|---|---|---|
--glm-mtp |
Enable GLM-MTP speculative decoding | All |
--glm-mtp-timing |
Enable timing instrumentation for MTP | All |
--ssd-streaming |
Activate routed-expert SSD streaming | Metal, ROCm |
--cuda |
Select CUDA backend | CUDA |
--rocm |
Select ROCm backend | ROCm |
--gpu-devices |
Specify GPU indices | CUDA |
--gpu-vram auto |
Auto-distribute VRAM | CUDA |
--power 100 |
Required power setting | All |
--ctx |
Context window size | All |
Common Pitfalls and Solutions
Unsupported Quantization Layout
Symptom: Fatal error "unsupported routed expert tensor type" during model loading.
Fix: Verify tensor types with gguf-tools or download officially supported models. Gate/up tensors must be Q2_K/Q4_K/Q5_K; down tensors add Q6_K as an option.
Incorrect Tensor-Parallel Flag
Symptom: Error "GLM does not support CUDA-TP" on startup.
Fix: Remove --cuda-tensor-parallel from your command line. GLM 5.2 uses conventional layer placement, not the DeepSeek-Flash TP layout. For Metal, --tensor-parallel is silently ignored.
Power Setting Ignored
Symptom: Model runs at full speed despite lower --power value.
Fix: GLM 5.2 only accepts --power 100. Lower values are intentionally ignored by the engine.
Complete Working Examples
macOS Metal with MTP Speculation
# Download and run with speculative decoding
./download_model.sh glm-antirez-iq2xxs
./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
--glm-mtp-timing --temp 0 --ctx 8192
CUDA Server with 8 GPUs
# Multi-GPU without tensor parallelism
./ds4 --cuda \
-m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
--gpu-devices 0,2,4,6,1,3,5,7 \
--gpu-vram auto \
--power 100 \
--ctx 32768
Low-Memory Streaming Setup
# ROCm with SSD streaming for large context
./ds4 --rocm \
-m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
--ssd-streaming \
--ctx 65536 \
--temp 0.6
Key Source Files
Understanding these files helps debug GLM 5.2 issues:
ds4.c— Core engine withds4_engine_routed_quant_bits()validation and model-family detectionds4_help.c— CLI flag definitions for--glm-mtp,--ssd-streaming, and related optionsds4_server.c— Server-mode GLM handling and tool-call syntaxgguf-tools/quality-testing/README.md— Official quality fixtures and accepted quant layouts
Summary
- GLM 5.2 requires exact quantization: Gate/up tensors must use
Q2_K,Q4_K, orQ5_K; down tensors additionally acceptQ6_K - Tensor parallelism forbidden on CUDA: Omit
--cuda-tensor-parallelentirely; use--gpu-devicesfor multi-GPU - MTP speculation available: Enable with
--glm-mtpor--glm-mtp-timingfor faster generation - SSD streaming for large models: Use
--ssd-streamingto cache routed experts on-demand - Power setting fixed: Always use
--power 100; other values are ignored
Frequently Asked Questions
What happens if I use an unsupported quantization type for GLM 5.2 routed experts?
The engine aborts during model loading with "unsupported routed expert tensor type". The tensor_expect_routed_expert function in ds4.c validates every routed-expert tensor against the allowed set and rejects any Q8_0, Q3_K, or other unsupported types before inference begins.
Can I use tensor parallelism with GLM 5.2 on CUDA?
No. GLM 5.2 explicitly does not support --cuda-tensor-parallel. Supplying this flag causes immediate abort with "GLM does not support CUDA-TP". For multi-GPU inference, list devices with --gpu-devices instead—the engine uses conventional layer placement across the specified GPUs.
What is the difference between --glm-mtp and --glm-mtp-timing?
Both enable the experimental greedy MTP (Multi-Token Prediction) speculation using the MTP block inside the GLM GGUF. --glm-mtp-timing adds performance instrumentation showing speculation acceptance rates and latency. Both require --temp 0 for deterministic greedy decoding.
Does SSD streaming work with all GLM 5.2 backends?
SSD streaming (--ssd-streaming) is available on Metal (macOS) and ROCm (Strix Halo) backends. The CUDA backend handles memory differently and does not use this path. When activated, non-routed weights stay GPU-resident while routed experts are fetched on-demand from storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →