Quantization Formats Supported by ds4 for DeepSeek and GLM Models
The ds4 runtime supports Q8_0, Q4_K, Q2_K, and IQ2_XXS quantization for DeepSeek models, while GLM models add Q8_K support, all defined in gguf-tools/quants.h and implemented in hardware-specific kernels.
The antirez/ds4 repository ships a compact quantizer designed for efficient inference of large language models through a fixed set of tensor-type identifiers. Understanding which quantization formats ds4 supports for DeepSeek and GLM models is essential for preparing compatible GGUF weights and avoiding loader rejection errors during initialization.
DeepSeek V4 Flash Quantization Support
The ds4 runtime implements four specific quantization formats for DeepSeek V4 Flash models, as explicitly documented in gguf-tools/README.md (lines 73-75).
Q8_0 (8-bit Integer)
Q8_0 provides full-precision integer quantization using 8 bits per value. This format offers the highest numerical fidelity among the supported DeepSeek formats, preserving model accuracy at the cost of larger memory footprints.
Q4_K (4-bit Block Quantization)
Q4_K employs 4-bit block-quantization with per-block scaling factors using the "K" layout. This format balances aggressive model compression with inference accuracy by grouping values into blocks with shared scaling parameters.
Q2_K (2-bit Block Quantization)
Q2_K utilizes 2-bit block-quantization with the "K" layout, enabling extreme compression for deployment scenarios with strict memory constraints. This format reduces storage requirements by 75% compared to 8-bit alternatives.
IQ2_XXS (2-bit Int-Quant for MoE)
IQ2_XXS is an extreme-low-bit format specifically designed for Mixture-of-Experts (MoE) gating mechanisms. According to the gguf-tools documentation, this format handles the expert-gate tensors in DeepSeek architectures where minimal bit-width is critical for routing efficiency.
GLM 5.x Quantization Support
GLM models support the same four formats as DeepSeek plus an additional specialized format implemented in GLM-specific compute kernels.
Core Formats (Q8_0, Q4_K, Q2_K, IQ2_XXS)
GLM 5.x models utilize Q8_0, Q4_K, Q2_K, and IQ2_XXS for general tensor storage and MoE routing, maintaining compatibility with the same quantization pipelines used for DeepSeek architectures.
Q8_K Low-Precision Layout
Unique to GLM support, Q8_K represents a low-precision "Q8-low-QK" layout optimized for GLM attention mechanisms. This format is exposed through dedicated kernels such as glm_q8_* found in metal/flash_attn.metal and rocm/ds4_rocm_q8.cuh, providing hardware-accelerated attention for GLM models.
Additionally, the Q2_K format for GLM is processed through specialized Metal kernels like glm_q2_K_pair_swiglu_simd_f32_impl in metal/moe.metal, handling the SwiGLU activation functions specific to GLM architecture.
Quantization Type Definitions in Source Code
The canonical enumeration of all supported quantization identifiers resides in gguf-tools/quants.h (lines 19-43). This header defines the tensor-type constants used throughout the ds4 runtime to identify quantization schemes during model loading:
/* From gguf-tools/quants.h - lines 19-43 */
#define DS4Q_TYPE_Q8_0 0x01
#define DS4Q_TYPE_Q4_K 0x02
#define DS4Q_TYPE_Q2_K 0x03
#define DS4Q_TYPE_IQ2_XXS 0x04
#define DS4Q_TYPE_Q8_K 0x05 /* GLM-specific extension */
These constants determine how the loader interprets tensor blocks and which dequantization kernels to invoke during inference.
Model Loading Compatibility and Validation
When loading a model, ds4 validates tensor types against the supported subsets for each model family. If a GGUF file contains unsupported formats such as Q5_0, IQ3_S, or BF16, the loader explicitly rejects the model to prevent execution on unoptimized code paths.
For example, attempting to load an unsupported quantization type results in a clear error:
# Attempting to load a Q5_0 quantized DeepSeek model
./ds4 --model deepseek-q5_0.gguf
# Error: Unsupported quantization type Q5_0 for model family DeepSeek
This strict validation ensures that only tested quantization formats execute during inference, preventing numerical instability or kernel execution failures.
Summary
- DeepSeek V4 Flash supports Q8_0, Q4_K, Q2_K, and IQ2_XXS formats as documented in
gguf-tools/README.md. - GLM 5.x adds Q8_K support through specialized kernels in
metal/flash_attn.metalandrocm/ds4_rocm_q8.cuh. - Type identifiers are defined as constants in
gguf-tools/quants.h(lines 19-43), includingDS4Q_TYPE_Q8_0andDS4Q_TYPE_Q4_K. - The loader strictly rejects unsupported formats (e.g., Q5_0, BF16, FP16) to ensure runtime stability.
Frequently Asked Questions
Does ds4 support 3-bit quantization formats like Q3_K?
No, ds4 does not support Q3_K or any 3-bit quantization variants. According to the source in gguf-tools/quants.h, only 2-bit, 4-bit, and 8-bit integer formats are implemented for DeepSeek and GLM model families.
Can I use BF16 or FP16 models with ds4?
No, ds4 does not support BF16 or FP16 floating-point formats for DeepSeek or GLM models. The runtime is optimized specifically for the integer quantization formats listed in the supported subsets, and the loader will reject any floating-point tensor types.
What is the difference between Q2_K and IQ2_XXS?
While both formats use 2-bit storage per value, Q2_K is a general-purpose block-quantization format suitable for standard tensors, whereas IQ2_XXS is specifically optimized for MoE (Mixture-of-Experts) gating tensors with extreme compression requirements and specialized integer quantization patterns.
Where are the quantization constants defined in the ds4 codebase?
The quantization type identifiers are defined in gguf-tools/quants.h (lines 19-43), which assigns integer constants such as DS4Q_TYPE_Q8_0, DS4Q_TYPE_Q4_K, and DS4Q_TYPE_Q8_K. These constants are referenced by the GGUF loader and kernel dispatch mechanisms throughout the Metal and ROCm backend implementations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →