Comparing GGUF Quantization Options (Q2_K, Q4_K, MXFP4, IQ2_XXS) in ds4 for Quality vs Memory
DS4 supports four distinct GGUF quantization formats—Q2_K, Q4_K, MXFP4, and IQ2_XXS—that trade off memory footprint against inference accuracy, with MXFP4 providing lossless compression and Q2_K achieving the smallest size when paired with an imatrix file.
The antirez/ds4 inference engine implements specialized quantization schemes for Mixture-of-Experts (MoE) models, allowing significant GPU memory reduction through GGUF quantization options. Understanding the technical differences between these formats is essential for optimizing both model quality and runtime performance on resource-constrained hardware.
Overview of Supported GGUF Quantization Formats
DS4 defines four main tensor types for expert weights in ds4.c, each optimized for different memory and quality requirements:
- Q2_K (
DS4_TENSOR_Q2_K): 2-bit quantization aligned to 128-element blocks - Q4_K (
DS4_TENSOR_Q4_K): 4-bit quantization aligned to 128-element blocks - MXFP4 (
DS4_TENSOR_MXFP4): 4-bit mixed-precision floating-point for lossless repacking - IQ2_XXS (
DS4_TENSOR_IQ2_XXS): 2-bit "int-quant" format with per-tensor scaling
The block size and memory characteristics vary significantly across these implementations. According to the ds4q_type_traits table in gguf-tools/quants.c【https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c#L48-L55】, the storage requirements per block are:
| Format | Block Size | Bytes per Block | Relative Memory |
|---|---|---|---|
| Q2_K | 128 elements | 84 B | ~½ × Q4_K |
| Q4_K | 128 elements | 144 B | ~⅔ × Q8_K |
| MXFP4 | 32 elements | 17 B | ~¼ × Q8_K |
| IQ2_XXS | 128 elements | 66 B | ~¾ × Q8_K |
Memory Footprint and Block Structure
The ds4_tensor_block_size() function returns the storage requirements for each format, calculating memory per token as block_size * (hidden_dim / QK_K) where hidden_dim typically equals 4096 for DeepSeek-V4-Flash models.
Q2_K and Q4_K both use QK_K = 128 alignment, storing quantized weights with scales and minima for each block. The MXFP4 format uses a smaller QK_MXFP4 = 32 block size, enabling finer-grained quantization for already-packed expert tensors. IQ2_XXS achieves the highest compression by using 66 bytes per 128-element block, though it stores expert down projections separately as Q2_K tensors.
Quality Characteristics and Imatrix Optimization
Quantization quality varies significantly based on the format and whether auxiliary calibration data is provided.
Q2_K initially uses a synthetic "weight-energy fallback" algorithm that estimates activation importance based on weight magnitudes alone. However, when an imatrix file is supplied via gguf-tools/imatrix, the quantizer in deepseek4-quantize.c substitutes real expert activation statistics【https://github.com/antirez/ds4/blob/main/gguf-tools/imatrix/README.md#L70-L78】. This improves accuracy substantially without increasing the tensor size, as the runtime type remains DS4_TENSOR_Q2_K.
Q4_K provides higher fidelity for routed expert matrices by storing 4 bits per weight, reducing quantization error compared to 2-bit alternatives. MXFP4 introduces no dequantization error because it merely repacks already-quantized tensors from native checkpoints, preserving the original FP16 arithmetic precision while reducing storage overhead.
IQ2_XXS applies aggressive compression to gate and up projections while maintaining Q2_K for down projections, placing its overall quality between imatrix-enhanced Q2_K and standard Q4_K.
Runtime Implementation and Dispatch
The inference engine dispatches to format-specific kernels based on tensor type checks scattered throughout ds4.c. For example, the expert matrix multiplication validates tensor types before execution:
/* Runtime kernel dispatch (ds4.c) */
if (gate_w->type != DS4_TENSOR_Q4_K || up_w->type != DS4_TENSOR_Q4_K)
ds4_die("expected Q4_K expert tensors");
...
if (gate_w->type == DS4_TENSOR_Q2_K) {
// run Q2_K specific matmul
q2k_matmul(...);
}
The quantization logic in gguf-tools/deepseek4-quantize.c selects the appropriate kernel based on the target type enum:
/* Quantizer selection (deepseek4-quantize.c) */
if (type == DS4Q_TYPE_Q2_K) {
quantize_q2k(...);
} else if (type == DS4Q_TYPE_Q4_K) {
quantize_q4k(...);
} else if (type == DS4Q_TYPE_MXFP4) {
// MXFP4 is only repacked, no dequant needed
repack_mxfp4(...);
} else if (type == DS4Q_TYPE_IQ2_XXS) {
quantize_iq2_xxs(...);
}
Practical Usage Examples
Select your GGUF quantization format by specifying the appropriate model file when launching the inference server. The following commands demonstrate each format with identical context and prompt parameters:
# Q2_K with imatrix (recommended quality/size trade-off)
./ds4 -m gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
--ctx 8192 -p "Explain quantization trade-offs."
# Pure Q4_K (high quality, moderate memory)
./ds4 -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
--ctx 8192 -p "Summarize the article."
# MXFP4 (lossless expert repacking)
./ds4 -m gguf/DeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
--ctx 8192 -p "Translate to French."
# IQ2_XXS (maximum compression)
./ds4 -m gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf \
--ctx 8192 -p "Write a short poem."
Summary
- Q2_K with imatrix provides the best balance of minimal memory footprint (~½ of Q4_K) and acceptable quality for production deployments
- Q4_K offers superior accuracy for quality-critical applications at ~1.7× the memory cost of Q2_K
- MXFP4 enables lossless compression of pre-quantized experts, introducing zero accuracy degradation while reducing storage to ~25% of Q8_K
- IQ2_XXS achieves the highest compression (75% of Q8_K) but requires careful validation due to higher quantization error on gate/up projections
- The imatrix calibration mechanism is the only method to improve Q2_K quality without increasing tensor size or changing the runtime format
Frequently Asked Questions
Which GGUF quantization option offers the smallest memory footprint in ds4?
IQ2_XXS provides the highest compression, using only 66 bytes per 128-element block compared to 84 bytes for Q2_K. However, Q2_K achieves better practical compression for the complete model when paired with an imatrix file, as it uniformly quantizes all expert tensors while IQ2_XXS falls back to Q2_K for down projections.
Does MXFP4 quantization affect model accuracy?
No. According to the implementation in gguf-tools/quants.c, MXFP4 performs lossless repacking of already-packed expert tensors from native checkpoints. The repack_mxfp4() function stores weights in 17-byte blocks without introducing dequantization error, maintaining identical inference accuracy to the original FP16 checkpoint.
What is an imatrix file and why does it improve Q2_K quality?
An imatrix (importance matrix) file contains real activation statistics collected from calibration data. By default, Q2_K uses synthetic weight-energy estimates to determine quantization scales. When an imatrix is supplied to deepseek4-quantize.c, the quantizer weights errors by actual expert activation frequencies, significantly reducing perplexity without changing the 2-bit storage format or block structure.
Can I mix different quantization formats within the same model?
Yes. The ds4 architecture specifically supports heterogeneous quantization schemes where attention projections, shared experts, and routed experts use different formats. The runtime validates tensor types at load time through checks like gate_w->type != DS4_TENSOR_Q4_K, ensuring each kernel receives the expected quantization layout.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →