How to Build Custom GGUF Quantizations Using deepseek4-quantize
deepseek4-quantize is a standalone C program in the antirez/ds4 repository that converts Hugging Face safetensors into GGUF models while applying DeepSeek V4 Flash quantization recipes, optionally using activation-importance matrices for routed MoE experts.
The deepseek4-quantize utility provides a complete pipeline for creating optimized GGUF files without dependencies on GGML. Located in the gguf-tools/ directory, this tool handles metadata parsing, safetensors loading, and custom quantization across multiple precision families.
Architecture Overview
The quantizer operates as a self-contained binary that performs three distinct stages: metadata extraction, weight loading with optional imatrix application, and quantization with GGUF writing. All heavy lifting occurs within the gguf-tools/ directory, implemented in modular C source files.
Key components include:
deepseek4-quantize.c– The entry point that parses command-line options, loads template GGUF files, reads safetensors, and drives the quantization pipeline.quants.candquants.h– Implement DS4-specific quantization families includingq8_0,q4_K,q2_K, andiq2_xxs.- Imatrix pipeline – Optional calibration data that guides per-column scaling for routed MoE experts, generated by the DS4 runtime and consumed when
--imatrixis supplied.
The tool processes tensors by first reading a template GGUF to establish tensor ordering and target data types. It then loads weights from HF safetensors, applies importance-based scaling (either from a provided imatrix or a synthetic fallback), and writes quantized codes via the appropriate routine in quants.c (such as quantize_q4_K or quantize_iq2_xxs).
Building the Tool
Compile the binary using the provided Makefile from the repository root:
make -C gguf-tools
This command links deepseek4-quantize.c with quants.c and pulls in -lm -pthread flags as specified in gguf-tools/Makefile. After compilation, the executable resides at gguf-tools/deepseek4-quantize.
Preparing an Imatrix (Optional but Recommended)
An imatrix (importance matrix) is a binary .dat file containing activation-importance statistics that significantly improve quantization quality for routed MoE experts. Generate this calibration data using the DS4 runtime:
python3 gguf-tools/imatrix/dataset/build_ds4_imatrix_dataset.py
./ds4 \
-m gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
--imatrix-dataset gguf-tools/imatrix/dataset/rendered_prompts.txt \
--imatrix-out gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat \
--ctx 32768
The resulting .dat file can be passed to deepseek4-quantize via the --imatrix flag. If omitted, the tool falls back to a weight-energy heuristic calculated as importance[col] = Σ_row weight[row][col]².
Creating a Custom Quantized GGUF
The quantization workflow requires specifying source safetensors, a template GGUF for structure, and output parameters. The basic invocation pattern is:
gguf-tools/deepseek4-quantize \
--hf <path-to-hf-safetensors> \
--template <template.gguf> \
--out <output.gguf> \
[--imatrix <imatrix.dat>] \
[--experts <family>] \
[--routed-w2 <family>] \
[--attention-proj <family>] \
[--shared <family>] \
[--output <family>] \
[--threads N] \
[--compare-tensor <tensor-name>]
Key parameters include:
--hf– Directory containing original Hugging Face safetensors.--template– GGUF supplying tokenizer, vocab, and tensor layout.--imatrix– Path to calibrated importance matrix (.datfile).--experts,--routed-w2, etc. – Override quantization families for specific tensor groups.--threads– Worker thread count for parallel expert quantization.--compare-tensor– Validate single tensor conversion without writing full output.
Example: Q2 Expert Quantization with Imatrix
Apply 2-bit quantization to experts using a calibration matrix:
gguf-tools/deepseek4-quantize \
--hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
--template gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf \
--out gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
--imatrix gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat
Example: Custom Family Overrides
Specify different quantization schemes for routed experts and attention projections:
gguf-tools/deepseek4-quantize \
--hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
--template gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf \
--out gguf/DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix.gguf \
--imatrix gguf/DeepSeek-V4-Flash-chat-v2-routed-moe-ds4.dat \
--experts iq2_xxs \
--routed-w2 q2_k
Example: Tensor Validation
Verify conversion accuracy for a specific tensor before full quantization:
gguf-tools/deepseek4-quantize \
--hf ../deepseek-v4-quants/hf/DeepSeek-V4-Flash \
--template MODEL.gguf \
--compare-tensor blk.0.attn_q_a.weight
This command prints hash comparisons and exits without generating a full GGUF, enabling quick sanity checks.
When No Imatrix Is Provided
If the target quantization requires importance vectors (such as iq2_xxs) and --imatrix is omitted, deepseek4-quantize computes a synthetic fallback using the weight-energy heuristic implemented in lines 70-78 of deepseek4-quantize.c:
importance[col] = Σ_row weight[row][col]²
While this heuristic supports early-stage 2-bit quantization, real activation-based imatrices yield superior quality for production models.
Quality-Testing Your Quantization
Evaluate custom GGUF output against official model releases using the provided testing utilities:
python3 gguf-tools/quality-testing/collect_official.py
make -C gguf-tools quality-score
gguf-tools/quality-testing/score_official MODEL.gguf gguf-tools/quality-testing/data/manifest.tsv /tmp/model.tsv 4096
python3 gguf-tools/quality-testing/compare_scores.py /tmp/old.tsv /tmp/new.tsv
These scripts compare token-level logits and report similarity scores, allowing iterative refinement of quantization parameters.
Summary
- Standalone Architecture:
deepseek4-quantizeoperates independently of GGML, with all logic contained ingguf-tools/deepseek4-quantize.candgguf-tools/quants.c. - Three-Stage Pipeline: Extracts metadata from templates, loads weights with optional imatrix scaling, and applies DS4-specific quantizers (
q8_0,q4_K,q2_K,iq2_xxs). - Build Command: Run
make -C gguf-toolsto compile the binary with threading and math library support. - Imatrix Support: Calibration data improves routed MoE expert quantization; synthetic fallback uses weight-energy squared summation when omitted.
- Fine-Grained Control: Override quantization families per tensor group using
--experts,--routed-w2,--attention-proj, and related flags. - Validation Tools: Use
--compare-tensorfor spot checks and the quality-testing scripts for full model evaluation.
Frequently Asked Questions
What quantization families does deepseek4-quantize support?
The tool implements DeepSeek V4 Flash-specific quantizers in gguf-tools/quants.c: q8_0 (8-bit), q4_K and q2_K (K-quantized 4-bit and 2-bit), and iq2_xxs (extremely low-bit importance-aware quantization). Each family offers different trade-offs between model size and inference quality.
Why does my quantized model require an imatrix file?
Routed MoE (Mixture of Experts) tensors benefit from activation-importance statistics to guide per-column scaling during aggressive quantization (particularly 2-bit schemes). Without an imatrix, the tool uses a synthetic heuristic based on weight energy, which produces lower quality results than calibration data generated from real prompts.
Can I quantize only specific tensor groups?
Yes. The command-line interface provides family override flags including --experts, --routed-w2, --attention-proj, --shared, and --output. These allow you to apply different quantization schemes to different model components, such as using iq2_xxs for routed experts while keeping attention projections at q8_0.
How do I verify that my custom quantization matches the expected output?
Use the --compare-tensor flag to regenerate a single tensor and perform byte-level comparison against the template. For comprehensive validation, run the quality-testing scripts (collect_official.py, score_official, compare_scores.py) to compare token-level logits between your custom GGUF and official releases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →