DS4 GGUF Quantization Formats: Complete Hardware Compatibility Guide

DS4 supports 20+ GGUF quantization formats ranging from 1-bit to 16-bit precision, with q4_K recommended for modern NVIDIA GPUs, q4_0 for Apple Metal, and q8_0 or bf16 for CPU inference.

The DS4 inference engine by Salvatore Sanfilippo (antirez) implements a comprehensive quantization system for running large language models efficiently across diverse hardware. This guide covers every supported GGUF format, their memory characteristics, and how to select the optimal quantization for your specific GPU or CPU.

Supported GGUF Quantization Formats in DS4

DS4 defines all quantization types in gguf-tools/quants.c through the ds4q_type enumeration. Each format uses a specific block size and packing strategy to balance memory efficiency against computational overhead.

Integer Quantization Formats (4-bit to 8-bit)

Format GGUF Type ID Block Size Packed Best Use Case
q4_0 DS4Q_TYPE_Q4_0 18 B / 32 values No General-purpose CUDA/Metal GPUs
q4_1 DS4Q_TYPE_Q4_1 20 B / 32 values No Slightly higher accuracy than q4_0
q5_0 DS4Q_TYPE_Q5_0 22 B / 32 values No GPUs with moderate VRAM headroom
q5_1 DS4Q_TYPE_Q5_1 24 B / 32 values No High-VRAM GPUs needing precision
q8_0 DS4Q_TYPE_Q8_0 34 B / 32 values Yes CPU inference, memory-rich systems
q8_1 DS4Q_TYPE_Q8_1 36 B / 32 values No Rare: exact 8-bit requirements

These legacy formats appear at lines 42-47 in gguf-tools/quants.c, with DS4Q_TYPE_Q4_0 serving as the default for many GPU kernels.

K-Quants: Optimized Block Formats

K-quant formats use 64-value blocks with improved quantization algorithms:

Format GGUF Type ID Block Size Packed Hardware Target
q2_K DS4Q_TYPE_Q2_K 84 B / 64 values Yes Low-VRAM GPUs, large models
q3_K DS4Q_TYPE_Q3_K 110 B / 64 values No Mid-range GPUs
q4_K DS4Q_TYPE_Q4_K 144 B / 64 values Yes Recommended for CUDA GPUs
q5_K DS4Q_TYPE_Q5_K 176 B / 64 values No High-end GPUs
q6_K DS4Q_TYPE_Q6_K 210 B / 64 values No High-end GPUs
q8_K DS4Q_TYPE_Q8_K 292 B / 64 values Yes Maximum precision 8-bit

The K-quant enumeration begins at line 50 in gguf-tools/quants.c, with DS4Q_TYPE_Q4_K at line 52 representing the most widely used format for modern GPU inference.

Imatrix Quants: Improved Quantization

Format GGUF Type ID Block Size Best Use Case
iq2_xxs DS4Q_TYPE_IQ2_XXS 66 B / 64 values Ultra-low VRAM scenarios
iq2_xs DS4Q_TYPE_IQ2_XS 74 B / 64 values Better accuracy than iq2_xxs
iq3_xxs DS4Q_TYPE_IQ3_XXS 98 B / 64 values Mid-range GPU balance
iq1_s DS4Q_TYPE_IQ1_S 50 B / 64 values Small model deployment
iq4_xs DS4Q_TYPE_IQ4_XS 136 B / 64 values Accuracy-memory balance

Floating-Point and Specialized Formats

Format GGUF Type ID Block Size Hardware Requirement
bf16 DS4Q_TYPE_BF16 2 B / value BFloat16-capable CPU/GPU
mxfp4 DS4Q_TYPE_MXFP4 17 B / 32 values DeepSeek V4 Flash models
nvfp4 DS4Q_TYPE_NVFP4 36 B / 64 values NVIDIA FP4 hardware
q1_0 DS4Q_TYPE_Q1_0 18 B / 128 values Experimental, minimal precision

The specialized formats including DS4Q_TYPE_MXFP4 (line 71) and DS4Q_TYPE_NVFP4 appear later in the enumeration, reflecting their newer addition to the GGUF specification.

Hardware-Specific GGUF Format Recommendations

NVIDIA CUDA GPUs (Ampere and Newer)

Recommended format: q4_K

The q4_K format provides optimal throughput on NVIDIA Tensor Cores. In metal/dsv4_misc.metal (lines 1622-1664), similar packed 4-bit kernels demonstrate how block-quantized weights map efficiently to SIMT execution units.

For older CUDA GPUs (Turing, Pascal), use q4_0 instead—its simpler block structure avoids instruction throughput bottlenecks on pre-Ampere architectures.

Apple Silicon (Metal)

Recommended formats: q4_0 or q4_K

DS4 contains dedicated Metal kernels for these formats. The Metal implementation at metal/dsv4_misc.metal handles the memory layout optimizations specific to Apple GPUs, ensuring coalesced memory access patterns for 4-bit quantized tensors.

CPU-Only Inference

Recommended formats: q8_0 or bf16

The packed q8_0 format (34 bytes per 32 values) enables efficient vectorized operations on modern CPUs with AVX-512 or AVX2 support. BFloat16 (bf16) works optimally on Intel Sapphire Rapids or AMD Zen 4 processors with native BFloat16 acceleration.

Ultra-Low VRAM Scenarios (< 4GB)

Recommended formats: iq2_xxs or q2_K

These 2-bit formats achieve the smallest memory footprint. The iq2_xxs imatrix variant (66 bytes per 64 values) uses importance matrix weighting to preserve critical weight precision despite aggressive quantization.

DeepSeek V4 Flash Models

Required format: mxfp4

DeepSeek V4 Flash was pre-quantized with MXFP4. Using this format avoids runtime dequantization overhead and preserves the original quantization-aware training benefits. Specify --target mxfp4 when converting from other formats.

Converting Models to GGUF Quantization Formats

DS4 provides gguf-tools/deepseek4-quantize.c for model conversion. The tool accepts --target with the lowercase underscore variant of each format name.

Basic Conversion Examples

Convert FP16 checkpoint to Q4_K (recommended default):

./gguf-tools/deepseek4-quantize \
    --base model-fp16.gguf \
    --out model-q4k.gguf \
    --target q4_k

Convert to Q8_0 for CPU deployment:

./gguf-tools/deepseek4-quantize \
    --base model-fp16.gguf \
    --out model-q8_0.gguf \
    --target q8_0

Convert to MXFP4 for Flash model compatibility:

./gguf-tools/deepseek4-quantize \
    --base model-fp16.gguf \
    --out model-mxfp4.gguf \
    --target mxfp4

Runtime Format Selection

DS4 automatically selects optimal kernels based on the tensor's type field. In ds4.c (lines 2058-2061), the inference engine dispatches to hardware-specific implementations:

// Tensor type constants used for kernel selection
DS4_TENSOR_Q4_0
DS4_TENSOR_Q4_K
DS4_TENSOR_Q8_0
// ... additional types

You can inspect a model's quantization format:

./ds4_cli -m model.gguf --inspect

Performance and Memory Trade-offs

Format Bits/Weight Relative Speed VRAM for 7B Model
q4_0 4.5 1.0x (baseline) ~4.0 GB
q4_K 4.5 0.95x ~4.0 GB
q5_K 5.5 0.85x ~4.9 GB
q8_0 8.5 0.70x ~7.5 GB
bf16 16 0.60x ~14.0 GB
iq2_xxs 2.06 0.75x ~1.8 GB

Lower bit-width formats reduce memory but may increase dequantization overhead. K-quants (q4_K, q5_K) generally outperform legacy formats at equivalent bit-widths due to superior quantization algorithms.

Compile-Time Default Configuration

Override the default quantization type at build time:

#define DS4_DEFAULT_TENSOR_TYPE DS4_TENSOR_Q4_K
#include "ds4.h"

This macro affects all tensor allocations that don't specify an explicit type. Rebuild after modification:

make clean && make DS4_DEFAULT_TENSOR_TYPE=DS4_TENSOR_Q8_0

Summary

  • DS4 supports 20 GGUF quantization formats defined in gguf-tools/quants.c, from 1-bit experimental formats to 16-bit BFloat16
  • q4_K is the recommended default for modern NVIDIA GPUs and Apple Metal, offering optimal speed-memory balance
  • q8_0 or bf16 work best for CPU inference where memory is less constrained
  • iq2_xxs and q2_K enable large model deployment on GPUs with ≤4GB VRAM
  • mxfp4 is required for DeepSeek V4 Flash model compatibility
  • Use deepseek4-quantize with --target to convert between formats, matching the lowercase name variants (q4_k, q8_0, mxfp4)

Frequently Asked Questions

What is the difference between q4_0 and q4_K in DS4?

q4_0 uses legacy 32-value blocks with simple scaling, while q4_K employs 64-value blocks with optimized K-quantization that preserves more weight information at the same bit-width. In DS4's Metal and CUDA kernels, q4_K typically achieves better perplexity scores with minimal speed penalty. Choose q4_0 for older GPUs lacking optimized K-quant kernels; otherwise prefer q4_K.

Can I run DS4 with mixed quantization formats within the same model?

DS4 supports per-tensor quantization types stored in the GGUF file header—different layers can use different formats. However, the deepseek4-quantize utility applies a uniform target format by default. Mixed quantization requires manual GGUF editing or custom conversion pipelines not provided in the standard DS4 distribution.

Why does DS4 include both mxfp4 and nvfp4 formats?

MXFP4 (Microscaling FP4) is an open standard used by DeepSeek V4 Flash models, while NVFP4 is NVIDIA's hardware-specific FP4 implementation. They use different exponent/mantissa bit layouts and block structures. Use MXFP4 for Flash model compatibility; NVFP4 only when targeting NVIDIA-specific FP4 acceleration hardware. The formats are not interchangeable—converting between them requires full requantization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →