Differences Between GLM 5.2 Support and DeepSeek V4 Models in ds4: A Technical Comparison

DeepSeek V4 models in ds4 enable full speculative decoding with embedded MTP blocks and 2-bit routed-MoE quantization, while GLM 5.2 support offers a compatible but restricted inference path lacking Flash-specific optimizations.

The ds4 inference engine is architecturally built around two distinct model families: DeepSeek V4 (Flash and PRO variants) and GLM 5.2. Understanding the differences between GLM 5.2 support and DeepSeek V4 models in ds4 is essential for selecting the correct quantization layout, command-line flags, and hardware backend for your deployment.

Model Selection and Architectural Scope

DeepSeek V4 operates as the primary target architecture for ds4. The engine only accepts specially prepared GGUFs listed in the Model Weights section of the repository, specifically Flash and PRO weights that contain a Multi-Token Prediction (MTP) block inside the GGUF container. These models follow a strict validation path defined in README.md lines 5-8.

GLM 5.2 support is intentionally limited to a curated set of GGUFs tested against the official 100-case fixture. As documented in README.md lines 48-61, the engine validates these models against a stricter layout schema than DeepSeek V4, rejecting unsupported quantization combinations early in the load sequence.

Quantization Layout and Memory Formats

The most significant divergence lies in the tensor quantization strategies. DeepSeek V4 implements a 2-bit routed-MoE layout where expert tensors are aggressively quantized to IQ2_XXS (gate and up projections) and Q2_K (down projections), while non-expert tensors remain in full precision. This layout is hard-coded through constants like DS4_GLM_WS_SLOTS in ds4.c (lines 38997-39002).

GLM 5.2 maintains dense tensors on conventional Q8/F32 paths. Only the routed expert tensors utilize the same low-bit quantizations (Q2_K, Q4_K, Q5_K, Q6_K) as DeepSeek V4. The engine treats these as optional layouts, and any deviation from the supported configurations documented in README.md lines 57-61 triggers an immediate runtime error.

Speculative Decoding and MTP Capabilities

DeepSeek V4 Flash models embed an MTP block that drives DSpark speculative decoding. When enabled via --glm-mtp or --glm-mtp-timing, this block proposes up to five future tokens that the main model later verifies. The implementation resides in ds4_streaming_hotlist_glm52.inc and associated inference kernels that read the MTP state directly from the GGUF.

GLM 5.2 does not support external MTP files. The --glm-mtp flag only toggles experimental greedy speculation on models containing embedded MTP blocks, but crucial Flash-only features are disabled. Directional steering, --power values below 100, --prefill-chunk, and external MTP file loading are unsupported for GLM 5.2 inference according to README.md lines 78-81.

KV Cache and Streaming Architecture

DeepSeek V4 supports sophisticated KV-cache checkpointing with a custom KVC header format that survives process restarts, implemented in ds4_kvstore.c. The engine can stream model weights from SSD when RAM is insufficient, a capability tied to the Flash model architecture.

GLM 5.2 utilizes the same KV-cache machinery but lacks the Flash-specific MTP checkpoint handling. While routed-expert tensors store in the 2-bit layout, the checkpoint format does not preserve the speculative decoding state between sessions.

Hardware Backend Optimization

DeepSeek V4 kernels are optimized for Metal (macOS), CUDA (DGX Spark, multi-GPU), and ROCm (Strix Halo). The backend-specific implementations in ds4_gpu.c and ds4_metal.m contain dedicated routing logic for the 2-bit expert quantization and MTP block processing.

GLM 5.2 runs on identical backends but bypasses the Flash-only optimizations. It executes through generic Q8/F32 kernels for dense operations, using only the shared routed-expert kernels for MoE layers without the specialized memory pipelining found in DeepSeek V4 paths.

Practical Usage Examples

The following commands demonstrate the functional divergence between the two model families:


# DeepSeek V4 Flash with 2-bit routed-MoE and speculative decoding

./ds4 -m gguf/deepseek-v4-flash.q2.gguf \
      --glm-mtp-timing --temp 0 \
      -p "Explain quantum entanglement versus classical correlation."

# DeepSeek V4 PRO without MTP, using standard quantization

./ds4 -m gguf/deepseek-v4-pro.q4.gguf \
      --temp 0.7 \
      -p "Write a Rust program that reads a file line-by-line."

# GLM 5.2 with limited flag support (no --power or --prefill-chunk)

./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
      --temp 0.8 \
      -p "Summarize The Little Prince in three sentences."

# This fails: Flash-only features rejected for GLM 5.2

./ds4 -m gguf/GLM-5.2-UD-IQ2_XXS.gguf --power 80 -p "Test"

The first two examples exercise the full DeepSeek V4 pipeline including MTP verification and KV-cache streaming. The third demonstrates valid GLM 5.2 invocation with restricted options, while the fourth illustrates the engine's validation logic rejecting incompatible flags.

Summary

  • DeepSeek V4 defines the core ds4 architecture with 2-bit routed-MoE quantization, embedded MTP blocks for speculative decoding, and full feature parity including directional steering and power limiting.
  • GLM 5.2 provides a compatible inference path using Q8/F32 dense tensors and optional routed-expert quantization, but lacks Flash-specific MTP handling and restricts command-line options.
  • Key implementation differences reside in ds4.c (model constants), ds4_streaming_hotlist_glm52.inc (MTP logic), and ds4_kvstore.c (checkpoint formats).
  • Hardware backends support both families, but DeepSeek V4 utilizes optimized kernels absent in the GLM 5.2 execution path.

Frequently Asked Questions

Can I use DeepSeek V4 quantization layouts with GLM 5.2 models?

No. While GLM 5.2 supports routed-expert quantization using Q2_K, Q4_K, Q5_K, and Q6_K formats similar to DeepSeek V4, it requires dense tensors to remain in Q8 or F32 precision. The 2-bit IQ2_XXS layout hard-coded for DeepSeek V4 Flash experts in ds4.c is not valid for GLM 5.2 GGUFs, and the engine validates tensor layouts strictly according to the model family detected at load time.

Why does speculative decoding behave differently on GLM 5.2 versus DeepSeek V4?

DeepSeek V4 Flash models contain an embedded MTP (Multi-Token Prediction) block within the GGUF that enables DSpark speculative decoding proposing multiple future tokens. GLM 5.2 lacks this embedded block architecture, so the --glm-mtp flag only enables experimental greedy speculation without the verification pipeline. Consequently, GLM 5.2 cannot utilize external MTP files or achieve the same throughput gains as DeepSeek V4 Flash.

Is the KV-cache checkpoint format compatible between both model families?

Both families use the KVC header format defined in ds4_kvstore.c, but DeepSeek V4 checkpoints preserve additional MTP state and speculative decoding contexts that GLM 5.2 checkpoints omit. While GLM 5.2 can resume inference from checkpoints, it does not support the Flash-specific features like streaming SSD fallback or MTP-aware state restoration available to DeepSeek V4 models.

Which command-line flags are restricted when using GLM 5.2 support?

GLM 5.2 inference rejects directional steering parameters, --power settings below 100, --prefill-chunk configurations, and external MTP file paths. The engine validates these restrictions at startup according to README.md lines 78-81, allowing only basic inference flags like --temp and the experimental --glm-mtp-timing for embedded blocks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →