# Differences Between GLM 5.2 Support and DeepSeek V4 Models in ds4: A Technical Comparison

> Compare GLM 5.2 support and DeepSeek V4 models in ds4. Discover DeepSeek V4's full speculative decoding, MTP blocks, and 2-bit MoE quantization versus GLM 5.2's restricted inference.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-09

---

**DeepSeek V4 models in ds4 enable full speculative decoding with embedded MTP blocks and 2-bit routed-MoE quantization, while GLM 5.2 support offers a compatible but restricted inference path lacking Flash-specific optimizations.**

The ds4 inference engine is architecturally built around two distinct model families: **DeepSeek V4** (Flash and PRO variants) and **GLM 5.2**. Understanding the differences between GLM 5.2 support and DeepSeek V4 models in ds4 is essential for selecting the correct quantization layout, command-line flags, and hardware backend for your deployment.

## Model Selection and Architectural Scope

DeepSeek V4 operates as the primary target architecture for ds4. The engine only accepts specially prepared GGUFs listed in the Model Weights section of the repository, specifically Flash and PRO weights that contain a **Multi-Token Prediction (MTP)** block inside the GGUF container. These models follow a strict validation path defined in [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) lines 5-8.

GLM 5.2 support is intentionally limited to a curated set of GGUFs tested against the official 100-case fixture. As documented in [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) lines 48-61, the engine validates these models against a stricter layout schema than DeepSeek V4, rejecting unsupported quantization combinations early in the load sequence.

## Quantization Layout and Memory Formats

The most significant divergence lies in the tensor quantization strategies. DeepSeek V4 implements a **2-bit routed-MoE layout** where expert tensors are aggressively quantized to **IQ2_XXS** (gate and up projections) and **Q2_K** (down projections), while non-expert tensors remain in full precision. This layout is hard-coded through constants like `DS4_GLM_WS_SLOTS` in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 38997-39002).

GLM 5.2 maintains dense tensors on conventional **Q8/F32** paths. Only the routed expert tensors utilize the same low-bit quantizations (Q2_K, Q4_K, Q5_K, Q6_K) as DeepSeek V4. The engine treats these as optional layouts, and any deviation from the supported configurations documented in [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) lines 57-61 triggers an immediate runtime error.

## Speculative Decoding and MTP Capabilities

DeepSeek V4 Flash models embed an **MTP block** that drives *DSpark* speculative decoding. When enabled via `--glm-mtp` or `--glm-mtp-timing`, this block proposes up to five future tokens that the main model later verifies. The implementation resides in `ds4_streaming_hotlist_glm52.inc` and associated inference kernels that read the MTP state directly from the GGUF.

GLM 5.2 does not support external MTP files. The `--glm-mtp` flag only toggles experimental greedy speculation on models containing embedded MTP blocks, but crucial Flash-only features are disabled. Directional steering, `--power` values below 100, `--prefill-chunk`, and external MTP file loading are **unsupported** for GLM 5.2 inference according to [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) lines 78-81.

## KV Cache and Streaming Architecture

DeepSeek V4 supports sophisticated KV-cache checkpointing with a custom `KVC` header format that survives process restarts, implemented in [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c). The engine can stream model weights from SSD when RAM is insufficient, a capability tied to the Flash model architecture.

GLM 5.2 utilizes the same KV-cache machinery but lacks the Flash-specific MTP checkpoint handling. While routed-expert tensors store in the 2-bit layout, the checkpoint format does not preserve the speculative decoding state between sessions.

## Hardware Backend Optimization

DeepSeek V4 kernels are optimized for **Metal** (macOS), **CUDA** (DGX Spark, multi-GPU), and **ROCm** (Strix Halo). The backend-specific implementations in [`ds4_gpu.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu.c) and `ds4_metal.m` contain dedicated routing logic for the 2-bit expert quantization and MTP block processing.

GLM 5.2 runs on identical backends but bypasses the Flash-only optimizations. It executes through generic Q8/F32 kernels for dense operations, using only the shared routed-expert kernels for MoE layers without the specialized memory pipelining found in DeepSeek V4 paths.

## Practical Usage Examples

The following commands demonstrate the functional divergence between the two model families:

```bash

# DeepSeek V4 Flash with 2-bit routed-MoE and speculative decoding

./ds4 -m gguf/deepseek-v4-flash.q2.gguf \
      --glm-mtp-timing --temp 0 \
      -p "Explain quantum entanglement versus classical correlation."

```

```bash

# DeepSeek V4 PRO without MTP, using standard quantization

./ds4 -m gguf/deepseek-v4-pro.q4.gguf \
      --temp 0.7 \
      -p "Write a Rust program that reads a file line-by-line."

```

```bash

# GLM 5.2 with limited flag support (no --power or --prefill-chunk)

./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
      --temp 0.8 \
      -p "Summarize The Little Prince in three sentences."

```

```bash

# This fails: Flash-only features rejected for GLM 5.2

./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS.gguf --power 80 -p "Test"

```

The first two examples exercise the full DeepSeek V4 pipeline including MTP verification and KV-cache streaming. The third demonstrates valid GLM 5.2 invocation with restricted options, while the fourth illustrates the engine's validation logic rejecting incompatible flags.

## Summary

- **DeepSeek V4** defines the core ds4 architecture with 2-bit routed-MoE quantization, embedded MTP blocks for speculative decoding, and full feature parity including directional steering and power limiting.
- **GLM 5.2** provides a compatible inference path using Q8/F32 dense tensors and optional routed-expert quantization, but lacks Flash-specific MTP handling and restricts command-line options.
- Key implementation differences reside in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (model constants), `ds4_streaming_hotlist_glm52.inc` (MTP logic), and [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c) (checkpoint formats).
- Hardware backends support both families, but DeepSeek V4 utilizes optimized kernels absent in the GLM 5.2 execution path.

## Frequently Asked Questions

### Can I use DeepSeek V4 quantization layouts with GLM 5.2 models?

No. While GLM 5.2 supports routed-expert quantization using Q2_K, Q4_K, Q5_K, and Q6_K formats similar to DeepSeek V4, it requires dense tensors to remain in Q8 or F32 precision. The 2-bit IQ2_XXS layout hard-coded for DeepSeek V4 Flash experts in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) is not valid for GLM 5.2 GGUFs, and the engine validates tensor layouts strictly according to the model family detected at load time.

### Why does speculative decoding behave differently on GLM 5.2 versus DeepSeek V4?

DeepSeek V4 Flash models contain an embedded MTP (Multi-Token Prediction) block within the GGUF that enables *DSpark* speculative decoding proposing multiple future tokens. GLM 5.2 lacks this embedded block architecture, so the `--glm-mtp` flag only enables experimental greedy speculation without the verification pipeline. Consequently, GLM 5.2 cannot utilize external MTP files or achieve the same throughput gains as DeepSeek V4 Flash.

### Is the KV-cache checkpoint format compatible between both model families?

Both families use the `KVC` header format defined in [`ds4_kvstore.c`](https://github.com/antirez/ds4/blob/main/ds4_kvstore.c), but DeepSeek V4 checkpoints preserve additional MTP state and speculative decoding contexts that GLM 5.2 checkpoints omit. While GLM 5.2 can resume inference from checkpoints, it does not support the Flash-specific features like streaming SSD fallback or MTP-aware state restoration available to DeepSeek V4 models.

### Which command-line flags are restricted when using GLM 5.2 support?

GLM 5.2 inference rejects directional steering parameters, `--power` settings below 100, `--prefill-chunk` configurations, and external MTP file paths. The engine validates these restrictions at startup according to [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) lines 78-81, allowing only basic inference flags like `--temp` and the experimental `--glm-mtp-timing` for embedded blocks.