DS4 GGUF Quantization Formats: Complete Hardware Compatibility Guide
DS4 supports 20+ GGUF quantization formats ranging from 1-bit to 16-bit precision, with q4_K recommended for modern NVIDIA GPUs, q4_0 for Apple Metal, and q8_0 or bf16 for CPU inference.
The DS4 inference engine by Salvatore Sanfilippo (antirez) implements a comprehensive quantization system for running large language models efficiently across diverse hardware. This guide covers every supported GGUF format, their memory characteristics, and how to select the optimal quantization for your specific GPU or CPU.
Supported GGUF Quantization Formats in DS4
DS4 defines all quantization types in gguf-tools/quants.c through the ds4q_type enumeration. Each format uses a specific block size and packing strategy to balance memory efficiency against computational overhead.
Integer Quantization Formats (4-bit to 8-bit)
| Format | GGUF Type ID | Block Size | Packed | Best Use Case |
|---|---|---|---|---|
| q4_0 | DS4Q_TYPE_Q4_0 |
18 B / 32 values | No | General-purpose CUDA/Metal GPUs |
| q4_1 | DS4Q_TYPE_Q4_1 |
20 B / 32 values | No | Slightly higher accuracy than q4_0 |
| q5_0 | DS4Q_TYPE_Q5_0 |
22 B / 32 values | No | GPUs with moderate VRAM headroom |
| q5_1 | DS4Q_TYPE_Q5_1 |
24 B / 32 values | No | High-VRAM GPUs needing precision |
| q8_0 | DS4Q_TYPE_Q8_0 |
34 B / 32 values | Yes | CPU inference, memory-rich systems |
| q8_1 | DS4Q_TYPE_Q8_1 |
36 B / 32 values | No | Rare: exact 8-bit requirements |
These legacy formats appear at lines 42-47 in gguf-tools/quants.c, with DS4Q_TYPE_Q4_0 serving as the default for many GPU kernels.
K-Quants: Optimized Block Formats
K-quant formats use 64-value blocks with improved quantization algorithms:
| Format | GGUF Type ID | Block Size | Packed | Hardware Target |
|---|---|---|---|---|
| q2_K | DS4Q_TYPE_Q2_K |
84 B / 64 values | Yes | Low-VRAM GPUs, large models |
| q3_K | DS4Q_TYPE_Q3_K |
110 B / 64 values | No | Mid-range GPUs |
| q4_K | DS4Q_TYPE_Q4_K |
144 B / 64 values | Yes | Recommended for CUDA GPUs |
| q5_K | DS4Q_TYPE_Q5_K |
176 B / 64 values | No | High-end GPUs |
| q6_K | DS4Q_TYPE_Q6_K |
210 B / 64 values | No | High-end GPUs |
| q8_K | DS4Q_TYPE_Q8_K |
292 B / 64 values | Yes | Maximum precision 8-bit |
The K-quant enumeration begins at line 50 in gguf-tools/quants.c, with DS4Q_TYPE_Q4_K at line 52 representing the most widely used format for modern GPU inference.
Imatrix Quants: Improved Quantization
| Format | GGUF Type ID | Block Size | Best Use Case |
|---|---|---|---|
| iq2_xxs | DS4Q_TYPE_IQ2_XXS |
66 B / 64 values | Ultra-low VRAM scenarios |
| iq2_xs | DS4Q_TYPE_IQ2_XS |
74 B / 64 values | Better accuracy than iq2_xxs |
| iq3_xxs | DS4Q_TYPE_IQ3_XXS |
98 B / 64 values | Mid-range GPU balance |
| iq1_s | DS4Q_TYPE_IQ1_S |
50 B / 64 values | Small model deployment |
| iq4_xs | DS4Q_TYPE_IQ4_XS |
136 B / 64 values | Accuracy-memory balance |
Floating-Point and Specialized Formats
| Format | GGUF Type ID | Block Size | Hardware Requirement |
|---|---|---|---|
| bf16 | DS4Q_TYPE_BF16 |
2 B / value | BFloat16-capable CPU/GPU |
| mxfp4 | DS4Q_TYPE_MXFP4 |
17 B / 32 values | DeepSeek V4 Flash models |
| nvfp4 | DS4Q_TYPE_NVFP4 |
36 B / 64 values | NVIDIA FP4 hardware |
| q1_0 | DS4Q_TYPE_Q1_0 |
18 B / 128 values | Experimental, minimal precision |
The specialized formats including DS4Q_TYPE_MXFP4 (line 71) and DS4Q_TYPE_NVFP4 appear later in the enumeration, reflecting their newer addition to the GGUF specification.
Hardware-Specific GGUF Format Recommendations
NVIDIA CUDA GPUs (Ampere and Newer)
Recommended format: q4_K
The q4_K format provides optimal throughput on NVIDIA Tensor Cores. In metal/dsv4_misc.metal (lines 1622-1664), similar packed 4-bit kernels demonstrate how block-quantized weights map efficiently to SIMT execution units.
For older CUDA GPUs (Turing, Pascal), use q4_0 instead—its simpler block structure avoids instruction throughput bottlenecks on pre-Ampere architectures.
Apple Silicon (Metal)
Recommended formats: q4_0 or q4_K
DS4 contains dedicated Metal kernels for these formats. The Metal implementation at metal/dsv4_misc.metal handles the memory layout optimizations specific to Apple GPUs, ensuring coalesced memory access patterns for 4-bit quantized tensors.
CPU-Only Inference
Recommended formats: q8_0 or bf16
The packed q8_0 format (34 bytes per 32 values) enables efficient vectorized operations on modern CPUs with AVX-512 or AVX2 support. BFloat16 (bf16) works optimally on Intel Sapphire Rapids or AMD Zen 4 processors with native BFloat16 acceleration.
Ultra-Low VRAM Scenarios (< 4GB)
Recommended formats: iq2_xxs or q2_K
These 2-bit formats achieve the smallest memory footprint. The iq2_xxs imatrix variant (66 bytes per 64 values) uses importance matrix weighting to preserve critical weight precision despite aggressive quantization.
DeepSeek V4 Flash Models
Required format: mxfp4
DeepSeek V4 Flash was pre-quantized with MXFP4. Using this format avoids runtime dequantization overhead and preserves the original quantization-aware training benefits. Specify --target mxfp4 when converting from other formats.
Converting Models to GGUF Quantization Formats
DS4 provides gguf-tools/deepseek4-quantize.c for model conversion. The tool accepts --target with the lowercase underscore variant of each format name.
Basic Conversion Examples
Convert FP16 checkpoint to Q4_K (recommended default):
./gguf-tools/deepseek4-quantize \
--base model-fp16.gguf \
--out model-q4k.gguf \
--target q4_k
Convert to Q8_0 for CPU deployment:
./gguf-tools/deepseek4-quantize \
--base model-fp16.gguf \
--out model-q8_0.gguf \
--target q8_0
Convert to MXFP4 for Flash model compatibility:
./gguf-tools/deepseek4-quantize \
--base model-fp16.gguf \
--out model-mxfp4.gguf \
--target mxfp4
Runtime Format Selection
DS4 automatically selects optimal kernels based on the tensor's type field. In ds4.c (lines 2058-2061), the inference engine dispatches to hardware-specific implementations:
// Tensor type constants used for kernel selection
DS4_TENSOR_Q4_0
DS4_TENSOR_Q4_K
DS4_TENSOR_Q8_0
// ... additional types
You can inspect a model's quantization format:
./ds4_cli -m model.gguf --inspect
Performance and Memory Trade-offs
| Format | Bits/Weight | Relative Speed | VRAM for 7B Model |
|---|---|---|---|
| q4_0 | 4.5 | 1.0x (baseline) | ~4.0 GB |
| q4_K | 4.5 | 0.95x | ~4.0 GB |
| q5_K | 5.5 | 0.85x | ~4.9 GB |
| q8_0 | 8.5 | 0.70x | ~7.5 GB |
| bf16 | 16 | 0.60x | ~14.0 GB |
| iq2_xxs | 2.06 | 0.75x | ~1.8 GB |
Lower bit-width formats reduce memory but may increase dequantization overhead. K-quants (q4_K, q5_K) generally outperform legacy formats at equivalent bit-widths due to superior quantization algorithms.
Compile-Time Default Configuration
Override the default quantization type at build time:
#define DS4_DEFAULT_TENSOR_TYPE DS4_TENSOR_Q4_K
#include "ds4.h"
This macro affects all tensor allocations that don't specify an explicit type. Rebuild after modification:
make clean && make DS4_DEFAULT_TENSOR_TYPE=DS4_TENSOR_Q8_0
Summary
- DS4 supports 20 GGUF quantization formats defined in
gguf-tools/quants.c, from 1-bit experimental formats to 16-bit BFloat16 - q4_K is the recommended default for modern NVIDIA GPUs and Apple Metal, offering optimal speed-memory balance
- q8_0 or bf16 work best for CPU inference where memory is less constrained
- iq2_xxs and q2_K enable large model deployment on GPUs with ≤4GB VRAM
- mxfp4 is required for DeepSeek V4 Flash model compatibility
- Use
deepseek4-quantizewith--targetto convert between formats, matching the lowercase name variants (q4_k,q8_0,mxfp4)
Frequently Asked Questions
What is the difference between q4_0 and q4_K in DS4?
q4_0 uses legacy 32-value blocks with simple scaling, while q4_K employs 64-value blocks with optimized K-quantization that preserves more weight information at the same bit-width. In DS4's Metal and CUDA kernels, q4_K typically achieves better perplexity scores with minimal speed penalty. Choose q4_0 for older GPUs lacking optimized K-quant kernels; otherwise prefer q4_K.
Can I run DS4 with mixed quantization formats within the same model?
DS4 supports per-tensor quantization types stored in the GGUF file header—different layers can use different formats. However, the deepseek4-quantize utility applies a uniform target format by default. Mixed quantization requires manual GGUF editing or custom conversion pipelines not provided in the standard DS4 distribution.
Why does DS4 include both mxfp4 and nvfp4 formats?
MXFP4 (Microscaling FP4) is an open standard used by DeepSeek V4 Flash models, while NVFP4 is NVIDIA's hardware-specific FP4 implementation. They use different exponent/mantissa bit layouts and block structures. Use MXFP4 for Flash model compatibility; NVFP4 only when targeting NVIDIA-specific FP4 acceleration hardware. The formats are not interchangeable—converting between them requires full requantization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →