How to Collect and Use imatrix for Better Quantization Quality in ds4

TLDR: Generate an importance matrix using ds4 --imatrix-out <file> from your original FP16/FP32 weights, then pass that file to deepseek4-quantize --imatrix <file> to guide IQ2_XXS and Q4_K quantization with real per-column importance values instead of synthetic heuristics.

The antirez/ds4 toolkit compresses DeepSeek mixture-of-experts models into efficient GGUF formats, and you can dramatically improve the fidelity of low-bit quantization by collecting and using an importance matrix (imatrix). This binary data structure stores column-wise importance weights derived from the original floating-point parameters, allowing the quantizer to allocate precision where it matters most.

Understanding the Importance Matrix

An imatrix is a binary file containing per-expert, column-wise importance values extracted from the original FP16 or FP32 weights. When present, the quantization kernels for IQ2_XXS and Q4_K formats use these values to determine which components require higher precision during compression. Without an imatrix, the code falls back to a synthetic heuristic based on weight energy, as noted in the comments in gguf-tools/deepseek4-quantize.c (lines 14-20).

The file follows the legacy llama.cpp .dat format: a header containing the entry count, followed by tensor name strings and float arrays of importance values. The loader imatrix_load in gguf-tools/deepseek4-quantize.c (lines 903-960) parses this structure.

Step 1: Collecting the imatrix File

Run the ds4 collector with the --imatrix-out flag to extract importance values from the original weights:

ds4 --model model.safetensors --imatrix-out my_model.imatrix ...

This writes a binary file containing:

  • A header with the number of entries
  • Tensor names paired with float arrays representing column-wise importance
  • Optional dataset and chunk metadata

The collector analyzes the activation patterns and weight magnitudes to determine which columns carry the most signal energy, creating a profile that the quantizer references during compression.

Step 2: Quantizing with imatrix

Feed the collected matrix into the quantization process using the --imatrix flag:

deepseek4-quantize --model model.safetensors \
    --imatrix my_model.imatrix \
    --out-gguf quantized.gguf ...

When imatrix_load reads the file, it makes the importance data available to the low-bit quantization routines. In gguf-tools/quants.c (lines 1096-1132), the function ds4q_quantize_iq2_xxs accepts an imatrix pointer:

bool ds4q_quantize_iq2_xxs(const float *src, ...,
    const float *imatrix) { ... }

If imatrix is non-NULL, the routine uses the supplied importance vectors to guide bit allocation; when NULL, it falls back to the synthetic heuristic.

Enforcing Strict imatrix Matching

By default, if a tensor lacks a matching entry in the matrix, the quantizer falls back to the synthetic heuristic. To force an abort instead, use --imatrix-strict, which sets p.imatrix_strict = true in ds4.c (lines 2758-2761), ensuring every tensor receives explicit importance data.

Step 3: Embedding imatrix into GGUF Files

To make the matrix portable, embed it directly into the quantized GGUF output. When you provide --imatrix during quantization, the function write_imatrix_kvs in gguf-tools/deepseek4-quantize.c (lines 1649-1650) writes additional key-value pairs such as "quantize.imatrix.file" into the GGUF header. Downstream tools can then retrieve the matrix using imatrix_find without requiring external files.

Complete Workflow Example

Follow this sequence to collect and use imatrix for optimal quantization quality:

  1. Extract the importance matrix:

    ds4 --model model.safetensors --imatrix-out model.imatrix --output-dir ./tmp
  2. Quantize with imatrix guidance:

    deepseek4-quantize --model model.safetensors \
        --imatrix model.imatrix \
        --out-gguf model_q4.gguf \
        --type Q4_K
  3. (Optional) Enforce strict matching for aggressive quantization:

    deepseek4-quantize --model model.safetensors \
        --imatrix model.imatrix \
        --imatrix-strict \
        --out-gguf model_iq2.gguf \
        --type IQ2_XXS
  4. (Optional) Load a GGUF with embedded imatrix—the loader automatically calls imatrix_load if the KV pairs are present.

Key Implementation Files

File Role
gguf-tools/deepseek4-quantize.c Parses --imatrix and --imatrix-out flags; implements imatrix_load (lines 903-960) and write_imatrix_kvs (lines 1649-1650); handles strict mode logic (lines 14-20).
gguf-tools/quants.c Contains low-bit quantization kernels like ds4q_quantize_iq2_xxs (lines 1096-1132) that consume the imatrix pointer.
ds4.c Defines command-line argument handling for --imatrix and --imatrix-strict (lines 2758-2761), mapping them to p.imatrix_file and p.imatrix_strict.

Summary

  • Collect an importance matrix using ds4 --imatrix-out <path> to capture per-column weight importance from the original FP16/FP32 model.
  • Pass the matrix to deepseek4-quantize via --imatrix <path> to improve IQ2_XXS and Q4_K quantization quality with real data instead of synthetic heuristics.
  • Use --imatrix-strict to abort quantization if any tensor lacks importance data, ensuring complete coverage.
  • Embed the matrix into GGUF files automatically by including --imatrix during quantization, storing metadata via write_imatrix_kvs for downstream retrieval.

Frequently Asked Questions

What file format does ds4 use for imatrix files?

ds4 uses the legacy llama.cpp binary .dat format. The file contains a header with entry counts, followed by tensor name strings and float arrays representing column-wise importance values. The loader imatrix_load in gguf-tools/deepseek4-quantize.c (lines 903-960) parses this structure.

Which quantization types benefit most from imatrix in ds4?

The aggressive IQ2_XXS format sees the most significant quality improvements, as implemented in ds4q_quantize_iq2_xxs in gguf-tools/quants.c (lines 1096-1132). Q4_K also benefits from importance-guided bit allocation, though the impact is less dramatic than with 2-bit quantization where precision is scarce.

What happens if a tensor is missing from the imatrix file?

By default, the quantizer falls back to a synthetic heuristic based on weight energy, as noted in the comments at lines 14-20 of deepseek4-quantize.c. If you pass --imatrix-strict, the tool aborts instead, setting p.imatrix_strict = true in ds4.c (lines 2758-2761) to ensure no tensor is quantized without explicit importance data.

Can I embed the imatrix directly into the quantized GGUF file?

Yes. When you provide --imatrix during quantization, the function write_imatrix_kvs (lines 1649-1650 in deepseek4-quantize.c) writes KV pairs like "quantize.imatrix.file" into the GGUF header. Downstream applications can then retrieve the matrix using imatrix_find without requiring external files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →