# IQ2_XXS vs Q2_K vs Q4_K Quantization: When to Use Each in DS4

> Compare IQ2_XXS, Q2_K, and Q4_K quantizations in DS4. Understand trade-offs in memory, computation, and quality to choose the best format for your needs.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-04

---

**IQ2_XXS, Q2_K, and Q4_K are GGUF quantization formats in DS4 that trade off memory footprint, computational overhead, and model quality through different bit-widths, block structures, and imatrix requirements.**

DS4 (the DeepSpeed-for-4-bit inference engine by antirez) implements three specialized quantization schemes for deploying mixture-of-experts (MoE) models efficiently. These formats target different matrix roles within expert layers, with IQ2_XXS optimized for extreme compression via imatrix reconstruction, Q2_K balancing speed and size for direct kernel execution, and Q4_K prioritizing accuracy for quality-sensitive layers.

## Quantization Format Comparison

| Format | Bit-width | Block Size | Storage per Block | Requires imatrix | Primary Use Case |
|--------|-----------|------------|-------------------|------------------|----------------|
| **IQ2_XXS** | 2 bits (×8 groups) | 256 values | 66 bytes | **Yes** | Gating/up-projection matrices in MoE layers |
| **Q2_K** | 2 bits | 256 values | 84 bytes | No | Down-projection matrices |
| **Q4_K** | 4 bits | 256 values | 144 bytes | No | High-memory or quality-critical experts |

The 66-byte blocks of **IQ2_XXS** achieve the smallest memory footprint by splitting each 256-value row into eight 32-value groups, each with independent scaling. **Q2_K** uses a simpler single-scale approach at 84 bytes for faster kernel execution. **Q4_K** doubles the precision to 4 bits at 144 bytes for substantially improved output quality.

## Architectural Deep Dive

### IQ2_XXS: Group-Wise Extreme Compression

In [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c), **IQ2_XXS** is flagged with `requires_imatrix = true`, indicating it needs an auxiliary importance matrix for full reconstruction during dequantization. This imatrix stores per-group scaling factors that the runtime must load before matrix-vector operations.

The Metal kernels in `metal/moe.metal` implement this via `kernel_mul_mv_iq2_xxs_*` functions, which perform the extra indirection step. As noted in the source:

> "The imatrix overhead is offset by the smallest memory footprint."

This design makes IQ2_XXS optimal for scenarios where **dozens of expert matrices** must reside simultaneously in GPU/CPU memory, such as routing gates and up-projection layers in large MoE architectures.

### Q2_K: Direct Kernel Execution

**Q2_K** derives from the GGML/llama.cpp quantization lineage. Each 256-value block stores a single scale plus 2-bit quantized values, occupying 84 bytes with **no imatrix requirement** (`requires_imatrix = false` in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c)).

The corresponding `kernel_mul_mv_q2_K_*` kernels in `metal/moe.metal` execute dequantization and multiplication in one pass. This eliminates the memory bandwidth and latency overhead of imatrix lookups, delivering **higher throughput** than IQ2_XXS while maintaining 2-bit compression.

### Q4_K: Quality-Optimized Quantization

**Q4_K** mirrors Q2_K's block structure but allocates 4 bits per value, yielding 144 bytes per 256-value block. The `kernel_mul_mv_q4_K_*` kernels handle this expanded precision directly.

The additional bits significantly reduce quantization error, making Q4_K the preferred choice for **high-memory model variants** or any layer where generation quality outweighs raw compression ratios.

## Selection Guidelines by Use Case

| Scenario | Recommended Format | Rationale |
|----------|------------------|-----------|
| Routing/gating matrices (many small experts) | **IQ2_XXS** | Maximum compression tolerates imatrix overhead when expert count is high |
| Up-projection with immediate fp16/f32 activation | **IQ2_XXS** (or Q2_K for speed) | Imatrix cost is amortized across many experts; memory wins dominate |
| Down-projection (final expert output) | **Q2_K** | Direct kernels maximize throughput for the critical output path |
| Quality-sensitive or high-memory experts | **Q4_K** | 4-bit precision preserves model fidelity with modest overhead |
| Production MoE deployment | **Hybrid** (IQ2_XXS + Q2_K + Q4_K) | Layer-specific quantization optimizes the full memory/quality/speed tradeoff |

The hybrid approach is explicitly supported in [`gguf-tools/mixed/splice_mixed_expert_layers_gguf.py`](https://github.com/antirez/ds4/blob/main/gguf-tools/mixed/splice_mixed_expert_layers_gguf.py), which constructs GGUF files with mixed quantization schemes tailored to each matrix role.

## Practical Usage Examples

Convert a checkpoint with IQ2_XXS for up-projection and Q2_K for down-projection:

```bash
ds4 convert -i model_fp16.gguf -o model_iq2_xxs.gguf \
    --quant-up iq2_xxs --quant-down q2_K

```

Create a three-tier mixed model with Q4_K for high-memory experts:

```bash
ds4 convert -i model_fp16.gguf -o model_mixed.gguf \
    --quant-up iq2_xxs \
    --quant-down q2_K \
    --quant-highmem q4_K

```

Load with explicit quantization selection via Python:

```python
import ds4

model = ds4.load(
    "model_iq2_xxs.gguf",
    quant_up="iq2_xxs",
    quant_down="q2_K"
)

```

## Key Implementation Files

- [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) — Block layouts, byte sizes, and `requires_imatrix` flags
- [`gguf-tools/quants.h`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.h) — Enumeration of quantization types
- `metal/moe.metal` — Metal kernels: `kernel_mul_mv_iq2_xxs_*`, `kernel_mul_mv_q2_K_*`, `kernel_mul_mv_q4_K_*`
- [`gguf-tools/mixed/splice_mixed_expert_layers_gguf.py`](https://github.com/antirez/ds4/blob/main/gguf-tools/mixed/splice_mixed_expert_layers_gguf.py) — Mixed-expert GGUF construction
- [`tests/test_q4k_dot.c`](https://github.com/antirez/ds4/blob/main/tests/test_q4k_dot.c) & [`tests/test_engine_correctness.c`](https://github.com/antirez/ds4/blob/main/tests/test_engine_correctness.c) — Validation suites for quantization correctness

## Summary

- **IQ2_XXS** delivers maximum compression (66 bytes/block) via group-wise quantization with imatrix reconstruction—ideal for many-expert scenarios where memory is the binding constraint.
- **Q2_K** provides direct kernel execution at 84 bytes/block, balancing speed and size for down-projection layers where imatrix overhead would be prohibitive.
- **Q4_K** offers superior fidelity at 144 bytes/block for quality-critical experts, with straightforward kernel implementation matching Q2_K's execution pattern.
- The DS4 codebase supports **hybrid quantization strategies**, assigning each format to the matrix role where its tradeoffs are optimal.

## Frequently Asked Questions

### What is an imatrix and why does IQ2_XXS require one?

An **imatrix** (importance matrix) is an auxiliary data structure storing per-group scaling factors that reconstruct full-precision values from the compressed 2-bit representation. IQ2_XXS needs this because it splits each 256-value block into eight independently-scaled 32-value groups, achieving smaller storage at the cost of extra memory bandwidth during inference. The `requires_imatrix` flag in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) enforces this dependency.

### Can I use Q2_K or Q4_K without an imatrix for all layers?

**Yes**—both Q2_K and Q4_K operate without imatrix overhead. Their kernels (`kernel_mul_mv_q2_K_*` and `kernel_mul_mv_q4_K_*` in `metal/moe.metal`) perform dequantization directly from the compact block representation. This makes them simpler to deploy and often faster for individual matrix operations, though IQ2_XXS may still win for very large expert counts due to its smaller memory footprint.

### When should I prefer Q4_K over the 2-bit formats?

Choose **Q4_K** when output quality metrics (perplexity, downstream task accuracy) degrade unacceptably with 2-bit quantization, or when deploying **high-memory model variants** that explicitly prioritize fidelity. The 4-bit format is also safer for smaller models where expert count is low and the memory savings of IQ2_XXS are less critical.

### Does DS4 support mixing all three formats in one model?

**Yes**—the [`splice_mixed_expert_layers_gguf.py`](https://github.com/antirez/ds4/blob/main/splice_mixed_expert_layers_gguf.py) script in `gguf-tools/mixed/` constructs GGUF files with heterogeneous quantization. Typical configurations use IQ2_XXS for gates and up-projections, Q2_K for down-projections, and Q4_K for designated high-memory experts. The Metal runtime dispatches to the appropriate kernel for each layer's format.