# How to Run GLM 5.2 Models with Supported Routed Quantization Layouts

> Learn how to run GLM 5.2 models with ds4 using supported routed quantization layouts like Q2 K, Q4 K, Q5 K, and Q6 K. Ensure correct tensor quantization for successful model loading.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-05

---

**GLM 5.2 models in ds4 require specific routed-expert quantization layouts—Q2_K, Q4_K, or Q5_K for gate/up tensors and Q2_K, Q4_K, Q5_K, or Q6_K for down tensors—and will fail to load if any other quantization type is detected.**

Running **GLM 5.2** (also referred to as **GLM-Dense-Sparse-Attention**) in the **antirez/ds4** engine involves understanding its unique mixture-of-experts architecture. Unlike standard dense models, GLM 5.2 combines normal dense tensors with **routed MoE (Mixture of Experts) tensors** that demand exacting quantization formats. This guide covers how to identify supported layouts, launch models on Metal, CUDA, and ROCm backends, and avoid common configuration errors.

## Understanding GLM 5.2 Architecture

### Model Family Detection

When ds4 opens a GGUF file, it immediately classifies the model family. In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the engine checks for `DS4_MODEL_FAMILY_GLM_DSA` and branches to GLM-specific shape handling (`DS4_SHAPE_GLM52`):

```c
if (DS4_MODEL_FAMILY == DS4_MODEL_FAMILY_GLM_DSA)

```

This classification occurs at [line 618 of ds4.c](https://github.com/antirez/ds4/blob/main/ds4.c#L618) and determines which quantization validation paths activate.

### Routed-Expert Quantization Detection

The engine validates routed-expert quantization through `ds4_engine_routed_quant_bits`, implemented at [line 49903 of ds4.c](https://github.com/antirez/ds4/blob/main/ds4.c#L49903):

```c
int ds4_engine_routed_quant_bits(ds4_engine *e) {
    // Scans first layer for gate tensor
    // Returns 4 for DS4_TENSOR_Q4_K, otherwise 2 (Q2/K-style)
}

```

This function automatically detects whether your model uses 2-bit or 4-bit routed experts by inspecting the gate tensor type.

### Supported Quantization Layouts

The ds4 README and validation code enforce strict quantization rules for routed experts:

| Tensor Role | Supported Quant Types |
|-------------|----------------------|
| **Gate / Up** | `Q2_K`, `Q4_K`, `Q5_K` |
| **Down** | `Q2_K`, `Q4_K`, `Q5_K`, `Q6_K` |

Any GLM GGUF using unsupported types (e.g., `Q8_0`, `Q3_K`, `IQ3_XXS`) triggers a fatal error from `tensor_expect_routed_expert` with *"unsupported routed expert tensor type"*.

## Preparing Supported GLM 5.2 Models

### Downloading Pre-Quantized Models

The repository provides helper scripts to fetch validated GLM 5.2 models:

```bash

# Q2_K routed layout (smallest, fastest)

./download_model.sh glm-antirez-q2

# Q4_K routed layout (better quality, unsupported on GPU-TP)

./download_model.sh glm-antirez-q4

# IQ2_XXS for gate/up, Q2_K for down (experimental quality)

./download_model.sh glm-antirez-iq2xxs

```

Verify your downloaded model follows the supported layout using the quality-testing fixtures in [`gguf-tools/quality-testing/README.md`](https://github.com/antirez/ds4/blob/main/gguf-tools/quality-testing/README.md).

## Running GLM 5.2 on Metal (macOS)

### Full Resident Mode

Load the entire model into GPU memory for maximum speed:

```bash
./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  --glm-mtp-timing --temp 0

```

**`--glm-mtp-timing`** enables greedy MTP (Multi-Token Prediction) speculation using the MTP block embedded in the main GGUF. **`--temp 0`** forces deterministic greedy decoding required for this mode.

### SSD Streaming Mode

For large models on limited RAM, stream routed experts on-demand:

```bash
./ds4 -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
  --ssd-streaming --ctx 32768

```

**`--ssd-streaming`** keeps non-routed weights resident while caching routed experts dynamically. The engine automatically computes a routed-expert cache budget based on available system memory.

## Running GLM 5.2 on CUDA

### Single-GPU Configuration

GLM 5.2 on CUDA **must not use tensor parallelism**. Launch with:

```bash
./ds4 --cuda -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  --gpu-devices 0 --power 100

```

**Critical:** The **`--power 100`** flag is required—GLM 5.2 ignores lower power settings per the README's "Power" section.

### Multi-GPU Configuration

For distributed inference across multiple GPUs, specify devices without the tensor-parallel flag:

```bash
./ds4 --cuda -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  --gpu-devices 0,2,4,6,1,3,5,7 --gpu-vram auto

```

**`--gpu-vram auto`** allows the engine to distribute model layers across the specified device list using conventional placement.

## Running GLM 5.2 on ROCm (Strix Halo)

AMD Strix Halo APUs use the ROCm backend with similar streaming support:

```bash
./ds4 --rocm -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
  --ssd-streaming --ctx 4096

```

Build with `make strix-halo` to enable the optimized ROCm kernels for this platform.

## GLM 5.2 Command-Line Flags Reference

| Flag | Purpose | Backend |
|------|---------|---------|
| `--glm-mtp` | Enable GLM-MTP speculative decoding | All |
| `--glm-mtp-timing` | Enable timing instrumentation for MTP | All |
| `--ssd-streaming` | Activate routed-expert SSD streaming | Metal, ROCm |
| `--cuda` | Select CUDA backend | CUDA |
| `--rocm` | Select ROCm backend | ROCm |
| `--gpu-devices` | Specify GPU indices | CUDA |
| `--gpu-vram auto` | Auto-distribute VRAM | CUDA |
| `--power 100` | Required power setting | All |
| `--ctx` | Context window size | All |

## Common Pitfalls and Solutions

### Unsupported Quantization Layout

**Symptom:** Fatal error *"unsupported routed expert tensor type"* during model loading.

**Fix:** Verify tensor types with `gguf-tools` or download officially supported models. Gate/up tensors must be `Q2_K`/`Q4_K`/`Q5_K`; down tensors add `Q6_K` as an option.

### Incorrect Tensor-Parallel Flag

**Symptom:** Error *"GLM does not support CUDA-TP"* on startup.

**Fix:** Remove **`--cuda-tensor-parallel`** from your command line. GLM 5.2 uses conventional layer placement, not the DeepSeek-Flash TP layout. For Metal, `--tensor-parallel` is silently ignored.

### Power Setting Ignored

**Symptom:** Model runs at full speed despite lower `--power` value.

**Fix:** GLM 5.2 only accepts `--power 100`. Lower values are intentionally ignored by the engine.

## Complete Working Examples

### macOS Metal with MTP Speculation

```bash

# Download and run with speculative decoding

./download_model.sh glm-antirez-iq2xxs

./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  --glm-mtp-timing --temp 0 --ctx 8192

```

### CUDA Server with 8 GPUs

```bash

# Multi-GPU without tensor parallelism

./ds4 --cuda \
  -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  --gpu-devices 0,2,4,6,1,3,5,7 \
  --gpu-vram auto \
  --power 100 \
  --ctx 32768

```

### Low-Memory Streaming Setup

```bash

# ROCm with SSD streaming for large context

./ds4 --rocm \
  -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
  --ssd-streaming \
  --ctx 65536 \
  --temp 0.6

```

## Key Source Files

Understanding these files helps debug GLM 5.2 issues:

- **[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)** — Core engine with `ds4_engine_routed_quant_bits()` validation and model-family detection
- **[`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c)** — CLI flag definitions for `--glm-mtp`, `--ssd-streaming`, and related options
- **[`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)** — Server-mode GLM handling and tool-call syntax
- **[`gguf-tools/quality-testing/README.md`](https://github.com/antirez/ds4/blob/main/gguf-tools/quality-testing/README.md)** — Official quality fixtures and accepted quant layouts

## Summary

- **GLM 5.2 requires exact quantization:** Gate/up tensors must use `Q2_K`, `Q4_K`, or `Q5_K`; down tensors additionally accept `Q6_K`
- **Tensor parallelism forbidden on CUDA:** Omit `--cuda-tensor-parallel` entirely; use `--gpu-devices` for multi-GPU
- **MTP speculation available:** Enable with `--glm-mtp` or `--glm-mtp-timing` for faster generation
- **SSD streaming for large models:** Use `--ssd-streaming` to cache routed experts on-demand
- **Power setting fixed:** Always use `--power 100`; other values are ignored

## Frequently Asked Questions

### What happens if I use an unsupported quantization type for GLM 5.2 routed experts?

The engine aborts during model loading with *"unsupported routed expert tensor type"*. The `tensor_expect_routed_expert` function in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) validates every routed-expert tensor against the allowed set and rejects any `Q8_0`, `Q3_K`, or other unsupported types before inference begins.

### Can I use tensor parallelism with GLM 5.2 on CUDA?

No. GLM 5.2 explicitly does not support `--cuda-tensor-parallel`. Supplying this flag causes immediate abort with *"GLM does not support CUDA-TP"*. For multi-GPU inference, list devices with `--gpu-devices` instead—the engine uses conventional layer placement across the specified GPUs.

### What is the difference between `--glm-mtp` and `--glm-mtp-timing`?

Both enable the experimental greedy MTP (Multi-Token Prediction) speculation using the MTP block inside the GLM GGUF. `--glm-mtp-timing` adds performance instrumentation showing speculation acceptance rates and latency. Both require `--temp 0` for deterministic greedy decoding.

### Does SSD streaming work with all GLM 5.2 backends?

SSD streaming (`--ssd-streaming`) is available on Metal (macOS) and ROCm (Strix Halo) backends. The CUDA backend handles memory differently and does not use this path. When activated, non-routed weights stay GPU-resident while routed experts are fetched on-demand from storage.