How to Build Custom GGUF Quantizations Using deepseek4-quantize

deepseek4-quantize is a standalone C program in the antirez/ds4 repository that converts Hugging Face safetensors into GGUF models while applying DeepSeek V4 Flash quantization recipes, optionally using activation-importance matrices for routed MoE experts.

The deepseek4-quantize utility provides a complete pipeline for creating optimized GGUF files without dependencies on GGML. Located in the gguf-tools/ directory, this tool handles metadata parsing, safetensors loading, and custom quantization across multiple precision families.

Architecture Overview

The quantizer operates as a self-contained binary that performs three distinct stages: metadata extraction, weight loading with optional imatrix application, and quantization with GGUF writing. All heavy lifting occurs within the gguf-tools/ directory, implemented in modular C source files.

Key components include:

  • deepseek4-quantize.c – The entry point that parses command-line options, loads template GGUF files, reads safetensors, and drives the quantization pipeline.
  • quants.c and quants.h – Implement DS4-specific quantization families including q8_0, q4_K, q2_K, and iq2_xxs.
  • Imatrix pipeline – Optional calibration data that guides per-column scaling for routed MoE experts, generated by the DS4 runtime and consumed when --imatrix is supplied.

The tool processes tensors by first reading a template GGUF to establish tensor ordering and target data types. It then loads weights from HF safetensors, applies importance-based scaling (either from a provided imatrix or a synthetic fallback), and writes quantized codes via the appropriate routine in quants.c (such as quantize_q4_K or quantize_iq2_xxs).

Building the Tool

Compile the binary using the provided Makefile from the repository root:

make -C gguf-tools

This command links deepseek4-quantize.c with quants.c and pulls in -lm -pthread flags as specified in gguf-tools/Makefile. After compilation, the executable resides at gguf-tools/deepseek4-quantize.

An imatrix (importance matrix) is a binary .dat file containing activation-importance statistics that significantly improve quantization quality for routed MoE experts. Generate this calibration data using the DS4 runtime:

python3 gguf-tools/imatrix/dataset/build_ds4_imatrix_dataset.py
./ds4 \
  -m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
  --imatrix-dataset gguf-tools/imatrix/dataset/rendered_prompts.txt \
  --imatrix-out gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat \
  --ctx 32768

The resulting .dat file can be passed to deepseek4-quantize via the --imatrix flag. If omitted, the tool falls back to a weight-energy heuristic calculated as importance[col] = Σ_row weight[row][col]².

Creating a Custom Quantized GGUF

The quantization workflow requires specifying source safetensors, a template GGUF for structure, and output parameters. The basic invocation pattern is:

gguf-tools/deepseek4-quantize \
  --hf <path-to-hf-safetensors> \
  --template <template.gguf> \
  --out <output.gguf> \
  [--imatrix <imatrix.dat>] \
  [--experts <family>] \
  [--routed-w2 <family>] \
  [--attention-proj <family>] \
  [--shared <family>] \
  [--output <family>] \
  [--threads N] \
  [--compare-tensor <tensor-name>]

Key parameters include:

  • --hf – Directory containing original Hugging Face safetensors.
  • --template – GGUF supplying tokenizer, vocab, and tensor layout.
  • --imatrix – Path to calibrated importance matrix (.dat file).
  • --experts, --routed-w2, etc. – Override quantization families for specific tensor groups.
  • --threads – Worker thread count for parallel expert quantization.
  • --compare-tensor – Validate single tensor conversion without writing full output.

Example: Q2 Expert Quantization with Imatrix

Apply 2-bit quantization to experts using a calibration matrix:

gguf-tools/deepseek4-quantize \
  --hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
  --template gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf \
  --out gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
  --imatrix gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat

Example: Custom Family Overrides

Specify different quantization schemes for routed experts and attention projections:

gguf-tools/deepseek4-quantize \
  --hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
  --template gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
  --out gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix.gguf \
  --imatrix gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat \
  --experts iq2_xxs \
  --routed-w2 q2_k

Example: Tensor Validation

Verify conversion accuracy for a specific tensor before full quantization:

gguf-tools/deepseek4-quantize \
  --hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
  --template MODEL.gguf \
  --compare-tensor blk.0.attn_q_a.weight

This command prints hash comparisons and exits without generating a full GGUF, enabling quick sanity checks.

When No Imatrix Is Provided

If the target quantization requires importance vectors (such as iq2_xxs) and --imatrix is omitted, deepseek4-quantize computes a synthetic fallback using the weight-energy heuristic implemented in lines 70-78 of deepseek4-quantize.c:

importance[col] = Σ_row weight[row][col]²

While this heuristic supports early-stage 2-bit quantization, real activation-based imatrices yield superior quality for production models.

Quality-Testing Your Quantization

Evaluate custom GGUF output against official model releases using the provided testing utilities:

python3 gguf-tools/quality-testing/collect_official.py
make -C gguf-tools quality-score
gguf-tools/quality-testing/score_official MODEL.gguf gguf-tools/quality-testing/data/manifest.tsv /tmp/model.tsv 4096
python3 gguf-tools/quality-testing/compare_scores.py /tmp/old.tsv /tmp/new.tsv

These scripts compare token-level logits and report similarity scores, allowing iterative refinement of quantization parameters.

Summary

  • Standalone Architecture: deepseek4-quantize operates independently of GGML, with all logic contained in gguf-tools/deepseek4-quantize.c and gguf-tools/quants.c.
  • Three-Stage Pipeline: Extracts metadata from templates, loads weights with optional imatrix scaling, and applies DS4-specific quantizers (q8_0, q4_K, q2_K, iq2_xxs).
  • Build Command: Run make -C gguf-tools to compile the binary with threading and math library support.
  • Imatrix Support: Calibration data improves routed MoE expert quantization; synthetic fallback uses weight-energy squared summation when omitted.
  • Fine-Grained Control: Override quantization families per tensor group using --experts, --routed-w2, --attention-proj, and related flags.
  • Validation Tools: Use --compare-tensor for spot checks and the quality-testing scripts for full model evaluation.

Frequently Asked Questions

What quantization families does deepseek4-quantize support?

The tool implements DeepSeek V4 Flash-specific quantizers in gguf-tools/quants.c: q8_0 (8-bit), q4_K and q2_K (K-quantized 4-bit and 2-bit), and iq2_xxs (extremely low-bit importance-aware quantization). Each family offers different trade-offs between model size and inference quality.

Why does my quantized model require an imatrix file?

Routed MoE (Mixture of Experts) tensors benefit from activation-importance statistics to guide per-column scaling during aggressive quantization (particularly 2-bit schemes). Without an imatrix, the tool uses a synthetic heuristic based on weight energy, which produces lower quality results than calibration data generated from real prompts.

Can I quantize only specific tensor groups?

Yes. The command-line interface provides family override flags including --experts, --routed-w2, --attention-proj, --shared, and --output. These allow you to apply different quantization schemes to different model components, such as using iq2_xxs for routed experts while keeping attention projections at q8_0.

How do I verify that my custom quantization matches the expected output?

Use the --compare-tensor flag to regenerate a single tensor and perform byte-level comparison against the template. For comprehensive validation, run the quality-testing scripts (collect_official.py, score_official, compare_scores.py) to compare token-level logits between your custom GGUF and official releases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →