# How to Build Custom GGUF Quantizations Using deepseek4-quantize

> Learn to build custom GGUF quantizations with deepseek4-quantize from the antirez/ds4 repository. Convert safetensors to GGUF models efficiently using Flash quantization.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-09

---

**`deepseek4-quantize` is a standalone C program in the antirez/ds4 repository that converts Hugging Face safetensors into GGUF models while applying DeepSeek V4 Flash quantization recipes, optionally using activation-importance matrices for routed MoE experts.**

The `deepseek4-quantize` utility provides a complete pipeline for creating optimized GGUF files without dependencies on GGML. Located in the `gguf-tools/` directory, this tool handles metadata parsing, safetensors loading, and custom quantization across multiple precision families.

## Architecture Overview

The quantizer operates as a self-contained binary that performs three distinct stages: **metadata extraction**, **weight loading with optional imatrix application**, and **quantization with GGUF writing**. All heavy lifting occurs within the `gguf-tools/` directory, implemented in modular C source files.

Key components include:

- **[`deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/deepseek4-quantize.c)** – The entry point that parses command-line options, loads template GGUF files, reads safetensors, and drives the quantization pipeline.
- **[`quants.c`](https://github.com/antirez/ds4/blob/main/quants.c) and [`quants.h`](https://github.com/antirez/ds4/blob/main/quants.h)** – Implement DS4-specific quantization families including `q8_0`, `q4_K`, `q2_K`, and `iq2_xxs`.
- **Imatrix pipeline** – Optional calibration data that guides per-column scaling for routed MoE experts, generated by the DS4 runtime and consumed when `--imatrix` is supplied.

The tool processes tensors by first reading a template GGUF to establish tensor ordering and target data types. It then loads weights from HF safetensors, applies importance-based scaling (either from a provided imatrix or a synthetic fallback), and writes quantized codes via the appropriate routine in [`quants.c`](https://github.com/antirez/ds4/blob/main/quants.c) (such as `quantize_q4_K` or `quantize_iq2_xxs`).

## Building the Tool

Compile the binary using the provided Makefile from the repository root:

```bash
make -C gguf-tools

```

This command links [`deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/deepseek4-quantize.c) with [`quants.c`](https://github.com/antirez/ds4/blob/main/quants.c) and pulls in `-lm -pthread` flags as specified in `gguf-tools/Makefile`. After compilation, the executable resides at `gguf-tools/deepseek4-quantize`.

## Preparing an Imatrix (Optional but Recommended)

An **imatrix** (importance matrix) is a binary `.dat` file containing activation-importance statistics that significantly improve quantization quality for routed MoE experts. Generate this calibration data using the DS4 runtime:

```bash
python3 gguf-tools/imatrix/dataset/build_ds4_imatrix_dataset.py
./ds4 \
  -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
  --imatrix-dataset gguf-tools/imatrix/dataset/rendered_prompts.txt \
  --imatrix-out gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat \
  --ctx 32768

```

The resulting `.dat` file can be passed to `deepseek4-quantize` via the `--imatrix` flag. If omitted, the tool falls back to a weight-energy heuristic calculated as `importance[col] = Σ_row weight[row][col]²`.

## Creating a Custom Quantized GGUF

The quantization workflow requires specifying source safetensors, a template GGUF for structure, and output parameters. The basic invocation pattern is:

```bash
gguf-tools/deepseek4-quantize \
  --hf <path-to-hf-safetensors> \
  --template <template.gguf> \
  --out <output.gguf> \
  [--imatrix <imatrix.dat>] \
  [--experts <family>] \
  [--routed-w2 <family>] \
  [--attention-proj <family>] \
  [--shared <family>] \
  [--output <family>] \
  [--threads N] \
  [--compare-tensor <tensor-name>]

```

Key parameters include:

- **`--hf`** – Directory containing original Hugging Face safetensors.
- **`--template`** – GGUF supplying tokenizer, vocab, and tensor layout.
- **`--imatrix`** – Path to calibrated importance matrix (`.dat` file).
- **`--experts`**, **`--routed-w2`**, etc. – Override quantization families for specific tensor groups.
- **`--threads`** – Worker thread count for parallel expert quantization.
- **`--compare-tensor`** – Validate single tensor conversion without writing full output.

### Example: Q2 Expert Quantization with Imatrix

Apply 2-bit quantization to experts using a calibration matrix:

```bash
gguf-tools/deepseek4-quantize \
  --hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
  --template gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf \
  --out gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
  --imatrix gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat

```

### Example: Custom Family Overrides

Specify different quantization schemes for routed experts and attention projections:

```bash
gguf-tools/deepseek4-quantize \
  --hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
  --template gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
  --out gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix.gguf \
  --imatrix gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat \
  --experts iq2_xxs \
  --routed-w2 q2_k

```

### Example: Tensor Validation

Verify conversion accuracy for a specific tensor before full quantization:

```bash
gguf-tools/deepseek4-quantize \
  --hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
  --template MODEL.gguf \
  --compare-tensor blk.0.attn_q_a.weight

```

This command prints hash comparisons and exits without generating a full GGUF, enabling quick sanity checks.

## When No Imatrix Is Provided

If the target quantization requires importance vectors (such as `iq2_xxs`) and `--imatrix` is omitted, `deepseek4-quantize` computes a synthetic fallback using the weight-energy heuristic implemented in lines 70-78 of [`deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/deepseek4-quantize.c):

```text
importance[col] = Σ_row weight[row][col]²

```

While this heuristic supports early-stage 2-bit quantization, real activation-based imatrices yield superior quality for production models.

## Quality-Testing Your Quantization

Evaluate custom GGUF output against official model releases using the provided testing utilities:

```bash
python3 gguf-tools/quality-testing/collect_official.py
make -C gguf-tools quality-score
gguf-tools/quality-testing/score_official MODEL.gguf gguf-tools/quality-testing/data/manifest.tsv /tmp/model.tsv 4096
python3 gguf-tools/quality-testing/compare_scores.py /tmp/old.tsv /tmp/new.tsv

```

These scripts compare token-level logits and report similarity scores, allowing iterative refinement of quantization parameters.

## Summary

- **Standalone Architecture**: `deepseek4-quantize` operates independently of GGML, with all logic contained in [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) and [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c).
- **Three-Stage Pipeline**: Extracts metadata from templates, loads weights with optional imatrix scaling, and applies DS4-specific quantizers (`q8_0`, `q4_K`, `q2_K`, `iq2_xxs`).
- **Build Command**: Run `make -C gguf-tools` to compile the binary with threading and math library support.
- **Imatrix Support**: Calibration data improves routed MoE expert quantization; synthetic fallback uses weight-energy squared summation when omitted.
- **Fine-Grained Control**: Override quantization families per tensor group using `--experts`, `--routed-w2`, `--attention-proj`, and related flags.
- **Validation Tools**: Use `--compare-tensor` for spot checks and the quality-testing scripts for full model evaluation.

## Frequently Asked Questions

### What quantization families does deepseek4-quantize support?

The tool implements DeepSeek V4 Flash-specific quantizers in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c): `q8_0` (8-bit), `q4_K` and `q2_K` (K-quantized 4-bit and 2-bit), and `iq2_xxs` (extremely low-bit importance-aware quantization). Each family offers different trade-offs between model size and inference quality.

### Why does my quantized model require an imatrix file?

Routed MoE (Mixture of Experts) tensors benefit from activation-importance statistics to guide per-column scaling during aggressive quantization (particularly 2-bit schemes). Without an imatrix, the tool uses a synthetic heuristic based on weight energy, which produces lower quality results than calibration data generated from real prompts.

### Can I quantize only specific tensor groups?

Yes. The command-line interface provides family override flags including `--experts`, `--routed-w2`, `--attention-proj`, `--shared`, and `--output`. These allow you to apply different quantization schemes to different model components, such as using `iq2_xxs` for routed experts while keeping attention projections at `q8_0`.

### How do I verify that my custom quantization matches the expected output?

Use the `--compare-tensor` flag to regenerate a single tensor and perform byte-level comparison against the template. For comprehensive validation, run the quality-testing scripts ([`collect_official.py`](https://github.com/antirez/ds4/blob/main/collect_official.py), `score_official`, [`compare_scores.py`](https://github.com/antirez/ds4/blob/main/compare_scores.py)) to compare token-level logits between your custom GGUF and official releases.