# Comparing GGUF Quantization Options (Q2_K, Q4_K, MXFP4, IQ2_XXS) in ds4 for Quality vs Memory

> Explore GGUF quantization in ds4 comparing Q2 K, Q4 K, MXFP4, and IQ2 XXS formats. Discover the best balance of quality and memory for your needs.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: comparison
- Published: 2026-08-08

---

**DS4 supports four distinct GGUF quantization formats—Q2_K, Q4_K, MXFP4, and IQ2_XXS—that trade off memory footprint against inference accuracy, with MXFP4 providing lossless compression and Q2_K achieving the smallest size when paired with an imatrix file.**

The `antirez/ds4` inference engine implements specialized quantization schemes for Mixture-of-Experts (MoE) models, allowing significant GPU memory reduction through GGUF quantization options. Understanding the technical differences between these formats is essential for optimizing both model quality and runtime performance on resource-constrained hardware.

## Overview of Supported GGUF Quantization Formats

DS4 defines four main tensor types for expert weights in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), each optimized for different memory and quality requirements:

- **Q2_K** (`DS4_TENSOR_Q2_K`): 2-bit quantization aligned to 128-element blocks
- **Q4_K** (`DS4_TENSOR_Q4_K`): 4-bit quantization aligned to 128-element blocks  
- **MXFP4** (`DS4_TENSOR_MXFP4`): 4-bit mixed-precision floating-point for lossless repacking
- **IQ2_XXS** (`DS4_TENSOR_IQ2_XXS`): 2-bit "int-quant" format with per-tensor scaling

The block size and memory characteristics vary significantly across these implementations. According to the `ds4q_type_traits` table in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c)【https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c#L48-L55】, the storage requirements per block are:

| Format | Block Size | Bytes per Block | Relative Memory |
|--------|------------|-----------------|-----------------|
| **Q2_K** | 128 elements | 84 B | ~½ × Q4_K |
| **Q4_K** | 128 elements | 144 B | ~⅔ × Q8_K |
| **MXFP4** | 32 elements | 17 B | ~¼ × Q8_K |
| **IQ2_XXS** | 128 elements | 66 B | ~¾ × Q8_K |

## Memory Footprint and Block Structure

The `ds4_tensor_block_size()` function returns the storage requirements for each format, calculating memory per token as `block_size * (hidden_dim / QK_K)` where `hidden_dim` typically equals 4096 for DeepSeek-V4-Flash models.

**Q2_K** and **Q4_K** both use `QK_K = 128` alignment, storing quantized weights with scales and minima for each block. The **MXFP4** format uses a smaller `QK_MXFP4 = 32` block size, enabling finer-grained quantization for already-packed expert tensors. **IQ2_XXS** achieves the highest compression by using 66 bytes per 128-element block, though it stores expert *down* projections separately as Q2_K tensors.

## Quality Characteristics and Imatrix Optimization

Quantization quality varies significantly based on the format and whether auxiliary calibration data is provided.

**Q2_K** initially uses a synthetic "weight-energy fallback" algorithm that estimates activation importance based on weight magnitudes alone. However, when an **imatrix** file is supplied via `gguf-tools/imatrix`, the quantizer in [`deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/deepseek4-quantize.c) substitutes real expert activation statistics【https://github.com/antirez/ds4/blob/main/gguf-tools/imatrix/README.md#L70-L78】. This improves accuracy substantially without increasing the tensor size, as the runtime type remains `DS4_TENSOR_Q2_K`.

**Q4_K** provides higher fidelity for routed expert matrices by storing 4 bits per weight, reducing quantization error compared to 2-bit alternatives. **MXFP4** introduces **no dequantization error** because it merely repacks already-quantized tensors from native checkpoints, preserving the original FP16 arithmetic precision while reducing storage overhead.

**IQ2_XXS** applies aggressive compression to gate and up projections while maintaining Q2_K for down projections, placing its overall quality between imatrix-enhanced Q2_K and standard Q4_K.

## Runtime Implementation and Dispatch

The inference engine dispatches to format-specific kernels based on tensor type checks scattered throughout [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c). For example, the expert matrix multiplication validates tensor types before execution:

```c
/* Runtime kernel dispatch (ds4.c) */
if (gate_w->type != DS4_TENSOR_Q4_K || up_w->type != DS4_TENSOR_Q4_K)
    ds4_die("expected Q4_K expert tensors");
...
if (gate_w->type == DS4_TENSOR_Q2_K) {
    // run Q2_K specific matmul
    q2k_matmul(...);
}

```

The quantization logic in [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) selects the appropriate kernel based on the target type enum:

```c
/* Quantizer selection (deepseek4-quantize.c) */
if (type == DS4Q_TYPE_Q2_K) {
    quantize_q2k(...);
} else if (type == DS4Q_TYPE_Q4_K) {
    quantize_q4k(...);
} else if (type == DS4Q_TYPE_MXFP4) {
    // MXFP4 is only repacked, no dequant needed
    repack_mxfp4(...);
} else if (type == DS4Q_TYPE_IQ2_XXS) {
    quantize_iq2_xxs(...);
}

```

## Practical Usage Examples

Select your GGUF quantization format by specifying the appropriate model file when launching the inference server. The following commands demonstrate each format with identical context and prompt parameters:

```bash

# Q2_K with imatrix (recommended quality/size trade-off)

./ds4 -m gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
     --ctx 8192 -p "Explain quantization trade-offs."

# Pure Q4_K (high quality, moderate memory)

./ds4 -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
     --ctx 8192 -p "Summarize the article."

# MXFP4 (lossless expert repacking)

./ds4 -m gguf/DeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
     --ctx 8192 -p "Translate to French."

# IQ2_XXS (maximum compression)

./ds4 -m gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf \
     --ctx 8192 -p "Write a short poem."

```

## Summary

- **Q2_K with imatrix** provides the best balance of minimal memory footprint (~½ of Q4_K) and acceptable quality for production deployments
- **Q4_K** offers superior accuracy for quality-critical applications at ~1.7× the memory cost of Q2_K
- **MXFP4** enables lossless compression of pre-quantized experts, introducing zero accuracy degradation while reducing storage to ~25% of Q8_K
- **IQ2_XXS** achieves the highest compression (75% of Q8_K) but requires careful validation due to higher quantization error on gate/up projections
- The **imatrix** calibration mechanism is the only method to improve Q2_K quality without increasing tensor size or changing the runtime format

## Frequently Asked Questions

### Which GGUF quantization option offers the smallest memory footprint in ds4?

**IQ2_XXS** provides the highest compression, using only 66 bytes per 128-element block compared to 84 bytes for Q2_K. However, Q2_K achieves better practical compression for the complete model when paired with an imatrix file, as it uniformly quantizes all expert tensors while IQ2_XXS falls back to Q2_K for down projections.

### Does MXFP4 quantization affect model accuracy?

No. According to the implementation in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c), MXFP4 performs **lossless repacking** of already-packed expert tensors from native checkpoints. The `repack_mxfp4()` function stores weights in 17-byte blocks without introducing dequantization error, maintaining identical inference accuracy to the original FP16 checkpoint.

### What is an imatrix file and why does it improve Q2_K quality?

An imatrix (importance matrix) file contains real activation statistics collected from calibration data. By default, Q2_K uses synthetic weight-energy estimates to determine quantization scales. When an imatrix is supplied to [`deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/deepseek4-quantize.c), the quantizer weights errors by actual expert activation frequencies, significantly reducing perplexity without changing the 2-bit storage format or block structure.

### Can I mix different quantization formats within the same model?

Yes. The ds4 architecture specifically supports heterogeneous quantization schemes where attention projections, shared experts, and routed experts use different formats. The runtime validates tensor types at load time through checks like `gate_w->type != DS4_TENSOR_Q4_K`, ensuring each kernel receives the expected quantization layout.