# Quantization Formats Supported by ds4 for DeepSeek and GLM Models

> Discover the quantization formats ds4 supports for DeepSeek and GLM models including Q8_0, Q4_K, Q2_K, IQ2_XXS, and Q8_K. Optimize your model performance.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-07

---

**The ds4 runtime supports Q8_0, Q4_K, Q2_K, and IQ2_XXS quantization for DeepSeek models, while GLM models add Q8_K support, all defined in [`gguf-tools/quants.h`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.h) and implemented in hardware-specific kernels.**

The antirez/ds4 repository ships a compact quantizer designed for efficient inference of large language models through a fixed set of tensor-type identifiers. Understanding which quantization formats ds4 supports for DeepSeek and GLM models is essential for preparing compatible GGUF weights and avoiding loader rejection errors during initialization.

## DeepSeek V4 Flash Quantization Support

The ds4 runtime implements four specific quantization formats for DeepSeek V4 Flash models, as explicitly documented in [`gguf-tools/README.md`](https://github.com/antirez/ds4/blob/main/gguf-tools/README.md) (lines 73-75).

### Q8_0 (8-bit Integer)

**Q8_0** provides full-precision integer quantization using 8 bits per value. This format offers the highest numerical fidelity among the supported DeepSeek formats, preserving model accuracy at the cost of larger memory footprints.

### Q4_K (4-bit Block Quantization)

**Q4_K** employs 4-bit block-quantization with per-block scaling factors using the "K" layout. This format balances aggressive model compression with inference accuracy by grouping values into blocks with shared scaling parameters.

### Q2_K (2-bit Block Quantization)

**Q2_K** utilizes 2-bit block-quantization with the "K" layout, enabling extreme compression for deployment scenarios with strict memory constraints. This format reduces storage requirements by 75% compared to 8-bit alternatives.

### IQ2_XXS (2-bit Int-Quant for MoE)

**IQ2_XXS** is an extreme-low-bit format specifically designed for Mixture-of-Experts (MoE) gating mechanisms. According to the gguf-tools documentation, this format handles the expert-gate tensors in DeepSeek architectures where minimal bit-width is critical for routing efficiency.

## GLM 5.x Quantization Support

GLM models support the same four formats as DeepSeek plus an additional specialized format implemented in GLM-specific compute kernels.

### Core Formats (Q8_0, Q4_K, Q2_K, IQ2_XXS)

GLM 5.x models utilize **Q8_0**, **Q4_K**, **Q2_K**, and **IQ2_XXS** for general tensor storage and MoE routing, maintaining compatibility with the same quantization pipelines used for DeepSeek architectures.

### Q8_K Low-Precision Layout

Unique to GLM support, **Q8_K** represents a low-precision "Q8-low-QK" layout optimized for GLM attention mechanisms. This format is exposed through dedicated kernels such as `glm_q8_*` found in `metal/flash_attn.metal` and `rocm/ds4_rocm_q8.cuh`, providing hardware-accelerated attention for GLM models.

Additionally, the Q2_K format for GLM is processed through specialized Metal kernels like `glm_q2_K_pair_swiglu_simd_f32_impl` in `metal/moe.metal`, handling the SwiGLU activation functions specific to GLM architecture.

## Quantization Type Definitions in Source Code

The canonical enumeration of all supported quantization identifiers resides in [`gguf-tools/quants.h`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.h) (lines 19-43). This header defines the tensor-type constants used throughout the ds4 runtime to identify quantization schemes during model loading:

```c
/* From gguf-tools/quants.h - lines 19-43 */
#define DS4Q_TYPE_Q8_0     0x01
#define DS4Q_TYPE_Q4_K     0x02
#define DS4Q_TYPE_Q2_K     0x03
#define DS4Q_TYPE_IQ2_XXS  0x04
#define DS4Q_TYPE_Q8_K     0x05  /* GLM-specific extension */

```

These constants determine how the loader interprets tensor blocks and which dequantization kernels to invoke during inference.

## Model Loading Compatibility and Validation

When loading a model, ds4 validates tensor types against the supported subsets for each model family. If a GGUF file contains unsupported formats such as `Q5_0`, `IQ3_S`, or `BF16`, the loader explicitly rejects the model to prevent execution on unoptimized code paths.

For example, attempting to load an unsupported quantization type results in a clear error:

```bash

# Attempting to load a Q5_0 quantized DeepSeek model

./ds4 --model deepseek-q5_0.gguf

# Error: Unsupported quantization type Q5_0 for model family DeepSeek

```

This strict validation ensures that only tested quantization formats execute during inference, preventing numerical instability or kernel execution failures.

## Summary

- **DeepSeek V4 Flash** supports **Q8_0**, **Q4_K**, **Q2_K**, and **IQ2_XXS** formats as documented in [`gguf-tools/README.md`](https://github.com/antirez/ds4/blob/main/gguf-tools/README.md).
- **GLM 5.x** adds **Q8_K** support through specialized kernels in `metal/flash_attn.metal` and `rocm/ds4_rocm_q8.cuh`.
- Type identifiers are defined as constants in [`gguf-tools/quants.h`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.h) (lines 19-43), including `DS4Q_TYPE_Q8_0` and `DS4Q_TYPE_Q4_K`.
- The loader strictly rejects unsupported formats (e.g., Q5_0, BF16, FP16) to ensure runtime stability.

## Frequently Asked Questions

### Does ds4 support 3-bit quantization formats like Q3_K?

No, ds4 does not support Q3_K or any 3-bit quantization variants. According to the source in [`gguf-tools/quants.h`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.h), only 2-bit, 4-bit, and 8-bit integer formats are implemented for DeepSeek and GLM model families.

### Can I use BF16 or FP16 models with ds4?

No, ds4 does not support BF16 or FP16 floating-point formats for DeepSeek or GLM models. The runtime is optimized specifically for the integer quantization formats listed in the supported subsets, and the loader will reject any floating-point tensor types.

### What is the difference between Q2_K and IQ2_XXS?

While both formats use 2-bit storage per value, **Q2_K** is a general-purpose block-quantization format suitable for standard tensors, whereas **IQ2_XXS** is specifically optimized for MoE (Mixture-of-Experts) gating tensors with extreme compression requirements and specialized integer quantization patterns.

### Where are the quantization constants defined in the ds4 codebase?

The quantization type identifiers are defined in [`gguf-tools/quants.h`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.h) (lines 19-43), which assigns integer constants such as `DS4Q_TYPE_Q8_0`, `DS4Q_TYPE_Q4_K`, and `DS4Q_TYPE_Q8_K`. These constants are referenced by the GGUF loader and kernel dispatch mechanisms throughout the Metal and ROCm backend implementations.