# Quantization Options and Memory Requirements for DeepSeek V4 PRO vs Flash Models in ds4

> Explore quantization options and memory needs for DeepSeek V4 PRO vs Flash models in ds4. Understand RAM usage for low-bit and mixed precision to optimize your deployments.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: performance
- Published: 2026-08-09

---

**DeepSeek V4 Flash supports low-bit quantizations including `q4_K`, `q2_K`, and `iq2_xxs` requiring approximately 30–40 GB of GPU memory for 1-million-token contexts, while the PRO variant relies on mixed FP4+FP8 precision and demands roughly 70–90 GB due to its larger 49 billion active parameters compared to Flash’s 13 billion.**

The `antirez/ds4` repository provides a specialized inference engine for the DeepSeek V4 architecture, implementing distinct quantization strategies and memory management schemes for the **Flash** (efficiency-optimized) and **PRO** (capacity-optimized) model families. Both variants share a 1-million-token context window and identical KV-cache compression logic, yet differ significantly in their supported weight formats and resulting hardware requirements.

## Quantization Formats and Supported Precisions

The ds4 codebase handles model weights through GGUF files, with each model family targeting different quantization schemes based on their parameter scale and intended deployment scenarios.

### Flash Model Quantization

The Flash variant is engineered for extreme efficiency and supports a **mixed FP4+FP8 precision** baseline alongside dedicated low-bit quantizers. According to [`MODEL_CARD.md`](https://github.com/antirez/ds4/blob/main/MODEL_CARD.md) (lines 78–84), Flash stores MoE expert weights in FP4 while retaining FP8 for remaining tensors. Additionally, the [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) file (lines 42–55) implements a quantization façade that enables Flash-specific recipes including:

- **`q8_0`** – 8-bit integer quantization
- **`q4_K`** – 4-bit block quantization with K-quantization
- **`q2_K`** – 2-bit block quantization for aggressive compression
- **`iq2_xxs`** – IQ2-XXS ultra-low-bit format

These formats allow Flash to maintain inference quality while minimizing VRAM occupancy, making it suitable for consumer and mid-range server GPUs.

### PRO Model Quantization

The PRO variant utilizes the same **mixed FP4+FP8** storage scheme as Flash but does not implement the low-bit quantizer façade found in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c). With 1.6 trillion total parameters and 49 billion active parameters per token, PRO GGUF files retain the mixed-precision format to balance file size against the substantial computational requirements of the larger expert matrices. The repository does not currently provide dedicated low-bit recipes (such as `q4_K` or `iq2_xxs`) for the PRO family, relying instead on the baseline mixed-precision representation.

## Memory Architecture and GPU Requirements

Both models leverage identical KV-cache compression constants defined in the core engine, yet their divergent active parameter counts create distinct memory footprints during inference.

### Shared KV-Cache Implementation

The ds4 engine implements a compressed KV-cache architecture governed by constants in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 55–60). Both Flash and PRO utilize:

- **Raw sliding-window KV** for the most recent 128 tokens
- **Compressed KV rows** employing alternating ratio-4 and ratio-128 compression layers
- **Indexer configuration** with 64 heads, 128-dimensional head size, and top-512 entry retrieval

This shared implementation means the per-token memory overhead for context storage remains consistent across both variants, with costs dominated by the raw 128-token window and the top-k indexer rather than the full parameter count.

### Empirical GPU Memory Footprint

Despite sharing KV-cache logic, the active parameter disparity creates substantially different hardware requirements:

| Model | Active Parameters | Typical GPU Memory (1M tokens) |
|-------|------------------|-------------------------------|
| **Flash** | 13 billion | ~30–40 GB |
| **PRO** | 49 billion | ~70–90 GB |

The **Flash** model achieves its lower footprint through both its smaller 13B active parameter set and the availability of low-bit quantization formats (`q2_K`, `iq2_xxs`) that further reduce weight storage. The **PRO** model, while benefiting from identical KV-cache compression, requires roughly double the VRAM due to its 4× larger active weight matrix, even when utilizing the same FP4+FP8 mixed precision.

## Loading and Converting Models in ds4

The repository provides straightforward mechanisms for loading pre-quantized models and converting base weights to Flash-compatible formats.

To load a Flash model with ultra-low-bit quantization:

```c
/* Load DeepSeek-V4-Flash with IQ2-XXS quantization */
const char *model_path = "gguf/DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf";
ds4_load_model(model_path);

```

To load a PRO model with standard mixed precision:

```c
/* Load DeepSeek-V4-PRO (FP4+FP8 mixed) */
const char *model_path = "gguf/DeepSeek-V4-Pro-0731-FP4+FP8.gguf";
ds4_load_model(model_path);

```

For Flash-specific conversion from FP16 to low-bit formats using the quantizer façade:

```bash

# Convert to q4_K using gguf-tools/quants.c implementation

./gguf-tools/quantize \
    --input DeepSeek-V4-Flash-Base-fp16.gguf \
    --output DeepSeek-V4-Flash-Base-q4_K.gguf \
    --type q4_K

```

Running 1-million-token inference with the Flash variant:

```bash
DS4_MODEL=gguf/DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
    ./ds4_cli -t 1 -p 1048576 -e "Explain quantum computing in simple terms."

```

## Summary

- **Flash** supports **mixed FP4+FP8** plus low-bit formats (`q8_0`, `q4_K`, `q2_K`, `iq2_xxs`) via [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c), while **PRO** uses only the mixed FP4+FP8 baseline.
- **Active parameters** differ significantly: 13B for Flash versus 49B for PRO, directly impacting memory requirements.
- Both models share identical **KV-cache constants** defined in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 55–60), implementing sliding-window (128 tokens) and ratio-based compression.
- **GPU memory** requirements scale to approximately **30–40 GB for Flash** and **70–90 GB for PRO** when processing 1-million-token contexts.
- The [`MODEL_CARD.md`](https://github.com/antirez/ds4/blob/main/MODEL_CARD.md) (lines 78–84) and [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) source files provide authoritative specifications for precision formats and memory architecture.

## Frequently Asked Questions

### What is the difference between total and active parameters in ds4?

**Total parameters** represent the complete weight count stored in the GGUF file—284 billion for Flash and 1.6 trillion for PRO—while **active parameters** indicate the subset actually utilized during a single forward pass. Flash activates 13 billion parameters per token, whereas PRO activates 49 billion, explaining the disparate memory requirements despite both using MoE (Mixture of Experts) architectures.

### Why does Flash support more quantization formats than PRO?

The Flash model includes a dedicated **quantization façade** in [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) that implements low-bit block quantizers (`q4_K`, `q2_K`, `iq2_xxs`) specifically optimized for its smaller expert matrices. The PRO variant’s substantially larger active parameter set (49B) would suffer unacceptable accuracy degradation with these aggressive compressions, so the repository maintains only the mixed FP4+FP8 format for PRO to preserve output quality.

### How does the KV cache compression work in both models?

Both Flash and PRO implement an identical compressed KV cache defined in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 55–60). The system maintains a **raw sliding window** of 128 recent tokens while compressing older KV rows through alternating **ratio-4** and **ratio-128** compression layers. A top-512 indexer with 64 heads of 128 dimensions each manages retrieval, ensuring that the per-token memory cost remains constant regardless of the 1-million-token context length.

### Can I convert the PRO model to low-bit quantization like `q4_K`?

**No.** The [`gguf-tools/quants.c`](https://github.com/antirez/ds4/blob/main/gguf-tools/quants.c) implementation specifically targets the Flash model family’s architecture and expert weight distributions. Attempting to apply these low-bit quantizers to PRO would require modifying the quantization façade to handle the larger 49B active parameter matrices, which is not supported in the current ds4 codebase. PRO models should be run using the standard mixed FP4+FP8 GGUF files provided in the repository.