# How to Collect and Use imatrix for Better Quantization Quality in ds4

> Enhance ds4 quantization quality by collecting and using imatrix. Learn how to generate an importance matrix and apply it for improved per-column quantization accuracy.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-09

---

**TLDR:** Generate an importance matrix using `ds4 --imatrix-out <file>` from your original FP16/FP32 weights, then pass that file to `deepseek4-quantize --imatrix <file>` to guide IQ2_XXS and Q4_K quantization with real per-column importance values instead of synthetic heuristics.

The antirez/ds4 toolkit compresses DeepSeek mixture-of-experts models into efficient GGUF formats, and you can dramatically improve the fidelity of low-bit quantization by collecting and using an **importance matrix** (imatrix). This binary data structure stores column-wise importance weights derived from the original floating-point parameters, allowing the quantizer to allocate precision where it matters most.

## Understanding the Importance Matrix

An **imatrix** is a binary file containing per-expert, column-wise importance values extracted from the original FP16 or FP32 weights. When present, the quantization kernels for **IQ2_XXS** and **Q4_K** formats use these values to determine which components require higher precision during compression. Without an imatrix, the code falls back to a synthetic heuristic based on weight energy, as noted in the comments in [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) (lines 14-20).

The file follows the legacy llama.cpp `.dat` format: a header containing the entry count, followed by tensor name strings and float arrays of importance values. The loader `imatrix_load` in [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) (lines 903-960) parses this structure.

## Step 1: Collecting the imatrix File

Run the ds4 collector with the `--imatrix-out` flag to extract importance values from the original weights:

```bash
ds4 --model model.safetensors --imatrix-out my_model.imatrix ...

```

This writes a binary file containing:
- A header with the number of entries
- Tensor names paired with float arrays representing column-wise importance
- Optional dataset and chunk metadata

The collector analyzes the activation patterns and weight magnitudes to determine which columns carry the most signal energy, creating a profile that the quantizer references during compression.

## Step 2: Quantizing with imatrix

Feed the collected matrix into the quantization process using the `--imatrix` flag:

```bash
deepseek4-quantize --model model.safetensors \
    --imatrix my_model.imatrix \
    --out-gguf quantized.gguf ...

```

When `imatrix_load` reads the file, it makes the importance data available to the low-bit quantization routines. In [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) (lines 1096-1132), the function `ds4q_quantize_iq2_xxs` accepts an `imatrix` pointer:

```c
bool ds4q_quantize_iq2_xxs(const float *src, ...,
    const float *imatrix) { ... }

```

If `imatrix` is non-NULL, the routine uses the supplied importance vectors to guide bit allocation; when NULL, it falls back to the synthetic heuristic.

### Enforcing Strict imatrix Matching

By default, if a tensor lacks a matching entry in the matrix, the quantizer falls back to the synthetic heuristic. To force an abort instead, use `--imatrix-strict`, which sets `p.imatrix_strict = true` in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 2758-2761), ensuring every tensor receives explicit importance data.

## Step 3: Embedding imatrix into GGUF Files

To make the matrix portable, embed it directly into the quantized GGUF output. When you provide `--imatrix` during quantization, the function `write_imatrix_kvs` in [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) (lines 1649-1650) writes additional key-value pairs such as `"quantize.imatrix.file"` into the GGUF header. Downstream tools can then retrieve the matrix using `imatrix_find` without requiring external files.

## Complete Workflow Example

Follow this sequence to collect and use imatrix for optimal quantization quality:

1. **Extract the importance matrix:**
   ```bash
   ds4 --model model.safetensors --imatrix-out model.imatrix --output-dir ./tmp
   ```

2. **Quantize with imatrix guidance:**
   ```bash
   deepseek4-quantize --model model.safetensors \
       --imatrix model.imatrix \
       --out-gguf model_q4.gguf \
       --type Q4_K
   ```

3. **(Optional) Enforce strict matching for aggressive quantization:**
   ```bash
   deepseek4-quantize --model model.safetensors \
       --imatrix model.imatrix \
       --imatrix-strict \
       --out-gguf model_iq2.gguf \
       --type IQ2_XXS
   ```

4. **(Optional) Load a GGUF with embedded imatrix**—the loader automatically calls `imatrix_load` if the KV pairs are present.

## Key Implementation Files

| File | Role |
|------|------|
| [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) | Parses `--imatrix` and `--imatrix-out` flags; implements `imatrix_load` (lines 903-960) and `write_imatrix_kvs` (lines 1649-1650); handles strict mode logic (lines 14-20). |
| [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) | Contains low-bit quantization kernels like `ds4q_quantize_iq2_xxs` (lines 1096-1132) that consume the `imatrix` pointer. |
| [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) | Defines command-line argument handling for `--imatrix` and `--imatrix-strict` (lines 2758-2761), mapping them to `p.imatrix_file` and `p.imatrix_strict`. |

## Summary

- Collect an importance matrix using `ds4 --imatrix-out <path>` to capture per-column weight importance from the original FP16/FP32 model.
- Pass the matrix to `deepseek4-quantize` via `--imatrix <path>` to improve IQ2_XXS and Q4_K quantization quality with real data instead of synthetic heuristics.
- Use `--imatrix-strict` to abort quantization if any tensor lacks importance data, ensuring complete coverage.
- Embed the matrix into GGUF files automatically by including `--imatrix` during quantization, storing metadata via `write_imatrix_kvs` for downstream retrieval.

## Frequently Asked Questions

### What file format does ds4 use for imatrix files?

ds4 uses the legacy llama.cpp binary `.dat` format. The file contains a header with entry counts, followed by tensor name strings and float arrays representing column-wise importance values. The loader `imatrix_load` in [`gguf-tools/deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/deepseek4-quantize.c) (lines 903-960) parses this structure.

### Which quantization types benefit most from imatrix in ds4?

The aggressive **IQ2_XXS** format sees the most significant quality improvements, as implemented in `ds4q_quantize_iq2_xxs` in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) (lines 1096-1132). **Q4_K** also benefits from importance-guided bit allocation, though the impact is less dramatic than with 2-bit quantization where precision is scarce.

### What happens if a tensor is missing from the imatrix file?

By default, the quantizer falls back to a synthetic heuristic based on weight energy, as noted in the comments at lines 14-20 of [`deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/deepseek4-quantize.c). If you pass `--imatrix-strict`, the tool aborts instead, setting `p.imatrix_strict = true` in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 2758-2761) to ensure no tensor is quantized without explicit importance data.

### Can I embed the imatrix directly into the quantized GGUF file?

Yes. When you provide `--imatrix` during quantization, the function `write_imatrix_kvs` (lines 1649-1650 in [`deepseek4-quantize.c`](https://github.com/antirez/ds4/blob/main/deepseek4-quantize.c)) writes KV pairs like `"quantize.imatrix.file"` into the GGUF header. Downstream applications can then retrieve the matrix using `imatrix_find` without requiring external files.