Supported GGUF Quantization Layouts for Routed MoE Experts in DS4

The DS4 engine explicitly supports seven GGUF quantization layouts for routed Mixture-of-Experts tensors: Q8_0, IQ2_XXS, Q2_K, Q4_K, Q5_K, Q6_K, and MXFP4, rejecting any other type at runtime via the tensor_is_routed_expert_type validator.

The antirez/ds4 repository implements a specialized inference engine for DeepSeek-style MoE models that enforces strict constraints on how routed expert weights are quantized. Unlike standard dense layers, the gate (w₁), up (w₃), and down (w₂) projection tensors for routed experts must use one of the predefined GGUF block formats to ensure compatibility with the engine's memory mapping and kernel dispatch logic.

Complete List of Supported GGUF Layouts

The definitive enumeration appears in a block-format comment near the top of ds4.c (lines 755-762) and is enforced by the validation logic in the same file. Each layout maps to a DS4 tensor constant and serves specific roles within the expert MLP:

  • Q8_0 (DS4_TENSOR_Q8_0): Standard 8-bit quantization used when memory is plentiful; suitable for any expert type.
  • IQ2_XXS (DS4_TENSOR_IQ2_XXS): Ultra-low-bit integer quantization typically reserved for gate (w₁) and up (w₃) projectors.
  • Q2_K (DS4_TENSOR_Q2_K): 2-bit "K-block" quantization optimized for down (w₂) projectors in routed experts.
  • Q4_K (DS4_TENSOR_Q4_K): 4-bit block quantization balancing size and accuracy for general routed experts (high-memory variant).
  • Q5_K (DS4_TENSOR_Q5_K): 5-bit block quantization used specifically by GLM-style routed experts.
  • Q6_K (DS4_TENSOR_Q6_K): 6-bit block quantization used specifically by GLM-style routed experts.
  • MXFP4 (DS4_TENSOR_MXFP4): Mixed-precision format for losslessly repacking packed experts from native checkpoints.

How DS4 Validates Expert Tensor Types

The runtime validation occurs in ds4.c within the tensor_is_routed_expert_type function (around lines 4426-4434). This helper returns true only if the tensor's GGUF type matches one of the seven supported constants listed above.

If an unsupported layout is encountered, the routed_expert_block_bytes helper triggers a fatal error at lines 4444-4446 via its default: case, emitting an "unsupported routed expert tensor type" message and terminating execution. This strict validation prevents silent misinterpretation of quantized weights that could corrupt inference results.

Quantizing Checkpoints with Specific Expert Layouts

The deepseek4-quantize CLI tool in gguf-tools/deepseek4-quantize.c exposes the supported layouts through dedicated flags. You can assign specific quantization schemes to each expert projection:


# Set gate (w1) and up (w3) to IQ2_XXS, down (w2) to Q4_K

deepseek4-quantize \
  --hf path/to/hf_dir \
  --template model_template.gguf \
  --out model_quantized.gguf \
  --routed-w1 iq2_xxs \
  --routed-w2 q4_k \
  --routed-w3 iq2_xxs \
  --n-experts 64

The --routed-w1, --routed-w2, and --routed-w3 flags map directly to the GGUF layout constants. According to the usage text at lines 2650-2653 of deepseek4-quantize.c, these options accept the lowercase string identifiers shown above. Additionally, line 2664 documents MXFP4 as a lossless repack option for preserving packed expert data from native checkpoints.

Programmatic Validation

When working with the DS4 C API, you can verify tensor compatibility before inference:

#include "ds4.h"

bool is_supported = tensor_is_routed_expert_type(tensor->type);
if (!is_supported) {
    ds4_die("unsupported routed expert tensor type");
}

This check mirrors the internal validation performed by the engine when loading GGUF files, ensuring your code fails fast if an incompatible expert tensor is encountered.

Summary

  • Seven layouts supported: Q8_0, IQ2_XXS, Q2_K, Q4_K, Q5_K, Q6_K, and MXFP4 are the only valid GGUF types for routed MoE experts in antirez/ds4.
  • Validation location: The tensor_is_routed_expert_type function in ds4.c (lines 4426-4434) enforces these constraints at runtime.
  • Typical assignments: IQ2_XXS for gate/up projections, Q2_K or Q4_K for down projections, and MXFP4 for lossless repacking.
  • CLI tooling: Use deepseek4-quantize with --routed-w1, --routed-w2, and --routed-w3 flags to select layouts during quantization.
  • Error handling: Unsupported types trigger a fatal error in routed_expert_block_bytes (lines 4444-4446 of ds4.c).

Frequently Asked Questions

What happens if I use an unsupported GGUF quantization layout for routed experts?

The DS4 runtime will terminate with a fatal error. When loading a model, the routed_expert_block_bytes helper checks the tensor type, and if it does not match one of the seven supported layouts (Q8_0, IQ2_XXS, Q2_K, Q4_K, Q5_K, Q6_K, or MXFP4), the default: case at lines 4444-4446 of ds4.c triggers an "unsupported routed expert tensor type" error.

Which quantization layout should I use for gate versus down expert projections?

According to the DS4 source documentation, IQ2_XXS is optimized for gate (w₁) and up (w₃) projectors, while Q2_K is specifically designated for down (w₂) projectors. For higher accuracy at the cost of larger file sizes, Q4_K can be used for any expert type, and GLM-style models typically employ Q5_K or Q6_K.

What is the MXFP4 layout used for?

MXFP4 is a special mixed-precision format defined as DS4_TENSOR_MXFP4 that performs lossless repacking of packed experts from native DeepSeek checkpoints. Unlike other layouts that re-quantize weights, MXFP4 preserves the original packed representation, making it ideal for converting pre-quantized models without precision loss. This is documented in deepseek4-quantize.c at line 2664.

Where are the supported quantization constants defined?

The string-to-constant mappings reside in gguf-tools/quants.h under the ds4q_type enum, while the validation logic and layout descriptions live in ds4.c. The definitive comment block listing all supported GGUF block formats appears at lines 755-762 of ds4.c, and the runtime switch validating these types is implemented around lines 4426-4434 of the same file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →