# DS4 GGUF Quantization Formats: Complete Hardware Compatibility Guide

> Explore DS4 GGUF quantization formats. Find the best option for your hardware from 20+ formats. Get recommendations for NVIDIA, Apple Metal, and CPU for optimal performance.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: guide
- Published: 2026-08-05

---

**DS4 supports 20+ GGUF quantization formats ranging from 1-bit to 16-bit precision, with q4_K recommended for modern NVIDIA GPUs, q4_0 for Apple Metal, and q8_0 or bf16 for CPU inference.**

The DS4 inference engine by Salvatore Sanfilippo (antirez) implements a comprehensive quantization system for running large language models efficiently across diverse hardware. This guide covers every supported GGUF format, their memory characteristics, and how to select the optimal quantization for your specific GPU or CPU.

## Supported GGUF Quantization Formats in DS4

DS4 defines all quantization types in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) through the `ds4q_type` enumeration. Each format uses a specific block size and packing strategy to balance memory efficiency against computational overhead.

### Integer Quantization Formats (4-bit to 8-bit)

| Format | GGUF Type ID | Block Size | Packed | Best Use Case |
|--------|-------------|------------|--------|---------------|
| **q4_0** | `DS4Q_TYPE_Q4_0` | 18 B / 32 values | No | General-purpose CUDA/Metal GPUs |
| **q4_1** | `DS4Q_TYPE_Q4_1` | 20 B / 32 values | No | Slightly higher accuracy than q4_0 |
| **q5_0** | `DS4Q_TYPE_Q5_0` | 22 B / 32 values | No | GPUs with moderate VRAM headroom |
| **q5_1** | `DS4Q_TYPE_Q5_1` | 24 B / 32 values | No | High-VRAM GPUs needing precision |
| **q8_0** | `DS4Q_TYPE_Q8_0` | 34 B / 32 values | **Yes** | CPU inference, memory-rich systems |
| **q8_1** | `DS4Q_TYPE_Q8_1` | 36 B / 32 values | No | Rare: exact 8-bit requirements |

These legacy formats appear at lines 42-47 in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c), with `DS4Q_TYPE_Q4_0` serving as the default for many GPU kernels.

### K-Quants: Optimized Block Formats

K-quant formats use 64-value blocks with improved quantization algorithms:

| Format | GGUF Type ID | Block Size | Packed | Hardware Target |
|--------|-------------|------------|--------|-----------------|
| **q2_K** | `DS4Q_TYPE_Q2_K` | 84 B / 64 values | **Yes** | Low-VRAM GPUs, large models |
| **q3_K** | `DS4Q_TYPE_Q3_K` | 110 B / 64 values | No | Mid-range GPUs |
| **q4_K** | `DS4Q_TYPE_Q4_K` | 144 B / 64 values | **Yes** | **Recommended for CUDA GPUs** |
| **q5_K** | `DS4Q_TYPE_Q5_K` | 176 B / 64 values | No | High-end GPUs |
| **q6_K** | `DS4Q_TYPE_Q6_K` | 210 B / 64 values | No | High-end GPUs |
| **q8_K** | `DS4Q_TYPE_Q8_K` | 292 B / 64 values | **Yes** | Maximum precision 8-bit |

The K-quant enumeration begins at line 50 in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c), with `DS4Q_TYPE_Q4_K` at line 52 representing the most widely used format for modern GPU inference.

### Imatrix Quants: Improved Quantization

| Format | GGUF Type ID | Block Size | Best Use Case |
|--------|-------------|------------|---------------|
| **iq2_xxs** | `DS4Q_TYPE_IQ2_XXS` | 66 B / 64 values | Ultra-low VRAM scenarios |
| **iq2_xs** | `DS4Q_TYPE_IQ2_XS` | 74 B / 64 values | Better accuracy than iq2_xxs |
| **iq3_xxs** | `DS4Q_TYPE_IQ3_XXS` | 98 B / 64 values | Mid-range GPU balance |
| **iq1_s** | `DS4Q_TYPE_IQ1_S` | 50 B / 64 values | Small model deployment |
| **iq4_xs** | `DS4Q_TYPE_IQ4_XS` | 136 B / 64 values | Accuracy-memory balance |

### Floating-Point and Specialized Formats

| Format | GGUF Type ID | Block Size | Hardware Requirement |
|--------|-------------|------------|-------------------|
| **bf16** | `DS4Q_TYPE_BF16` | 2 B / value | BFloat16-capable CPU/GPU |
| **mxfp4** | `DS4Q_TYPE_MXFP4` | 17 B / 32 values | DeepSeek V4 Flash models |
| **nvfp4** | `DS4Q_TYPE_NVFP4` | 36 B / 64 values | NVIDIA FP4 hardware |
| **q1_0** | `DS4Q_TYPE_Q1_0` | 18 B / 128 values | Experimental, minimal precision |

The specialized formats including `DS4Q_TYPE_MXFP4` (line 71) and `DS4Q_TYPE_NVFP4` appear later in the enumeration, reflecting their newer addition to the GGUF specification.

## Hardware-Specific GGUF Format Recommendations

### NVIDIA CUDA GPUs (Ampere and Newer)

**Recommended format:** `q4_K`

The q4_K format provides optimal throughput on NVIDIA Tensor Cores. In `metal/dsv4_misc.metal` (lines 1622-1664), similar packed 4-bit kernels demonstrate how block-quantized weights map efficiently to SIMT execution units.

For older CUDA GPUs (Turing, Pascal), use `q4_0` instead—its simpler block structure avoids instruction throughput bottlenecks on pre-Ampere architectures.

### Apple Silicon (Metal)

**Recommended formats:** `q4_0` or `q4_K`

DS4 contains dedicated Metal kernels for these formats. The Metal implementation at `metal/dsv4_misc.metal` handles the memory layout optimizations specific to Apple GPUs, ensuring coalesced memory access patterns for 4-bit quantized tensors.

### CPU-Only Inference

**Recommended formats:** `q8_0` or `bf16`

The packed `q8_0` format (34 bytes per 32 values) enables efficient vectorized operations on modern CPUs with AVX-512 or AVX2 support. BFloat16 (`bf16`) works optimally on Intel Sapphire Rapids or AMD Zen 4 processors with native BFloat16 acceleration.

### Ultra-Low VRAM Scenarios (< 4GB)

**Recommended formats:** `iq2_xxs` or `q2_K`

These 2-bit formats achieve the smallest memory footprint. The `iq2_xxs` imatrix variant (66 bytes per 64 values) uses importance matrix weighting to preserve critical weight precision despite aggressive quantization.

### DeepSeek V4 Flash Models

**Required format:** `mxfp4`

DeepSeek V4 Flash was pre-quantized with MXFP4. Using this format avoids runtime dequantization overhead and preserves the original quantization-aware training benefits. Specify `--target mxfp4` when converting from other formats.

## Converting Models to GGUF Quantization Formats

DS4 provides [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) for model conversion. The tool accepts `--target` with the lowercase underscore variant of each format name.

### Basic Conversion Examples

Convert FP16 checkpoint to Q4_K (recommended default):

```bash
./gguf-tools/deepseek4-quantize \
    --base model-fp16.gguf \
    --out model-q4k.gguf \
    --target q4_k

```

Convert to Q8_0 for CPU deployment:

```bash
./gguf-tools/deepseek4-quantize \
    --base model-fp16.gguf \
    --out model-q8_0.gguf \
    --target q8_0

```

Convert to MXFP4 for Flash model compatibility:

```bash
./gguf-tools/deepseek4-quantize \
    --base model-fp16.gguf \
    --out model-mxfp4.gguf \
    --target mxfp4

```

### Runtime Format Selection

DS4 automatically selects optimal kernels based on the tensor's `type` field. In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 2058-2061), the inference engine dispatches to hardware-specific implementations:

```c
// Tensor type constants used for kernel selection
DS4_TENSOR_Q4_0
DS4_TENSOR_Q4_K
DS4_TENSOR_Q8_0
// ... additional types

```

You can inspect a model's quantization format:

```bash
./ds4_cli -m model.gguf --inspect

```

## Performance and Memory Trade-offs

| Format | Bits/Weight | Relative Speed | VRAM for 7B Model |
|--------|-------------|----------------|-------------------|
| q4_0 | 4.5 | 1.0x (baseline) | ~4.0 GB |
| q4_K | 4.5 | 0.95x | ~4.0 GB |
| q5_K | 5.5 | 0.85x | ~4.9 GB |
| q8_0 | 8.5 | 0.70x | ~7.5 GB |
| bf16 | 16 | 0.60x | ~14.0 GB |
| iq2_xxs | 2.06 | 0.75x | ~1.8 GB |

Lower bit-width formats reduce memory but may increase dequantization overhead. K-quants (`q4_K`, `q5_K`) generally outperform legacy formats at equivalent bit-widths due to superior quantization algorithms.

## Compile-Time Default Configuration

Override the default quantization type at build time:

```c
#define DS4_DEFAULT_TENSOR_TYPE DS4_TENSOR_Q4_K
#include "ds4.h"

```

This macro affects all tensor allocations that don't specify an explicit type. Rebuild after modification:

```bash
make clean && make DS4_DEFAULT_TENSOR_TYPE=DS4_TENSOR_Q8_0

```

## Summary

- **DS4 supports 20 GGUF quantization formats** defined in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c), from 1-bit experimental formats to 16-bit BFloat16
- **q4_K is the recommended default** for modern NVIDIA GPUs and Apple Metal, offering optimal speed-memory balance
- **q8_0 or bf16** work best for CPU inference where memory is less constrained
- **iq2_xxs and q2_K** enable large model deployment on GPUs with ≤4GB VRAM
- **mxfp4 is required** for DeepSeek V4 Flash model compatibility
- Use `deepseek4-quantize` with `--target` to convert between formats, matching the lowercase name variants (`q4_k`, `q8_0`, `mxfp4`)

## Frequently Asked Questions

### What is the difference between q4_0 and q4_K in DS4?

**q4_0** uses legacy 32-value blocks with simple scaling, while **q4_K** employs 64-value blocks with optimized K-quantization that preserves more weight information at the same bit-width. In DS4's Metal and CUDA kernels, q4_K typically achieves better perplexity scores with minimal speed penalty. Choose q4_0 for older GPUs lacking optimized K-quant kernels; otherwise prefer q4_K.

### Can I run DS4 with mixed quantization formats within the same model?

DS4 supports per-tensor quantization types stored in the GGUF file header—different layers can use different formats. However, the `deepseek4-quantize` utility applies a uniform target format by default. Mixed quantization requires manual GGUF editing or custom conversion pipelines not provided in the standard DS4 distribution.

### Why does DS4 include both mxfp4 and nvfp4 formats?

**MXFP4** (Microscaling FP4) is an open standard used by DeepSeek V4 Flash models, while **NVFP4** is NVIDIA's hardware-specific FP4 implementation. They use different exponent/mantissa bit layouts and block structures. Use MXFP4 for Flash model compatibility; NVFP4 only when targeting NVIDIA-specific FP4 acceleration hardware. The formats are not interchangeable—converting between them requires full requantization.