Quantization Options and Memory Requirements for DeepSeek V4 PRO vs Flash Models in ds4
DeepSeek V4 Flash supports low-bit quantizations including q4_K, q2_K, and iq2_xxs requiring approximately 30–40 GB of GPU memory for 1-million-token contexts, while the PRO variant relies on mixed FP4+FP8 precision and demands roughly 70–90 GB due to its larger 49 billion active parameters compared to Flash’s 13 billion.
The antirez/ds4 repository provides a specialized inference engine for the DeepSeek V4 architecture, implementing distinct quantization strategies and memory management schemes for the Flash (efficiency-optimized) and PRO (capacity-optimized) model families. Both variants share a 1-million-token context window and identical KV-cache compression logic, yet differ significantly in their supported weight formats and resulting hardware requirements.
Quantization Formats and Supported Precisions
The ds4 codebase handles model weights through GGUF files, with each model family targeting different quantization schemes based on their parameter scale and intended deployment scenarios.
Flash Model Quantization
The Flash variant is engineered for extreme efficiency and supports a mixed FP4+FP8 precision baseline alongside dedicated low-bit quantizers. According to MODEL_CARD.md (lines 78–84), Flash stores MoE expert weights in FP4 while retaining FP8 for remaining tensors. Additionally, the gguf-tools/quants.c file (lines 42–55) implements a quantization façade that enables Flash-specific recipes including:
q8_0– 8-bit integer quantizationq4_K– 4-bit block quantization with K-quantizationq2_K– 2-bit block quantization for aggressive compressioniq2_xxs– IQ2-XXS ultra-low-bit format
These formats allow Flash to maintain inference quality while minimizing VRAM occupancy, making it suitable for consumer and mid-range server GPUs.
PRO Model Quantization
The PRO variant utilizes the same mixed FP4+FP8 storage scheme as Flash but does not implement the low-bit quantizer façade found in gguf-tools/quants.c. With 1.6 trillion total parameters and 49 billion active parameters per token, PRO GGUF files retain the mixed-precision format to balance file size against the substantial computational requirements of the larger expert matrices. The repository does not currently provide dedicated low-bit recipes (such as q4_K or iq2_xxs) for the PRO family, relying instead on the baseline mixed-precision representation.
Memory Architecture and GPU Requirements
Both models leverage identical KV-cache compression constants defined in the core engine, yet their divergent active parameter counts create distinct memory footprints during inference.
Shared KV-Cache Implementation
The ds4 engine implements a compressed KV-cache architecture governed by constants in ds4.c (lines 55–60). Both Flash and PRO utilize:
- Raw sliding-window KV for the most recent 128 tokens
- Compressed KV rows employing alternating ratio-4 and ratio-128 compression layers
- Indexer configuration with 64 heads, 128-dimensional head size, and top-512 entry retrieval
This shared implementation means the per-token memory overhead for context storage remains consistent across both variants, with costs dominated by the raw 128-token window and the top-k indexer rather than the full parameter count.
Empirical GPU Memory Footprint
Despite sharing KV-cache logic, the active parameter disparity creates substantially different hardware requirements:
| Model | Active Parameters | Typical GPU Memory (1M tokens) |
|---|---|---|
| Flash | 13 billion | ~30–40 GB |
| PRO | 49 billion | ~70–90 GB |
The Flash model achieves its lower footprint through both its smaller 13B active parameter set and the availability of low-bit quantization formats (q2_K, iq2_xxs) that further reduce weight storage. The PRO model, while benefiting from identical KV-cache compression, requires roughly double the VRAM due to its 4× larger active weight matrix, even when utilizing the same FP4+FP8 mixed precision.
Loading and Converting Models in ds4
The repository provides straightforward mechanisms for loading pre-quantized models and converting base weights to Flash-compatible formats.
To load a Flash model with ultra-low-bit quantization:
/* Load DeepSeek-V4-Flash with IQ2-XXS quantization */
const char *model_path = "gguf/DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf";
ds4_load_model(model_path);
To load a PRO model with standard mixed precision:
/* Load DeepSeek-V4-PRO (FP4+FP8 mixed) */
const char *model_path = "gguf/DeepSeek-V4-Pro-0731-FP4+FP8.gguf";
ds4_load_model(model_path);
For Flash-specific conversion from FP16 to low-bit formats using the quantizer façade:
# Convert to q4_K using gguf-tools/quants.c implementation
./gguf-tools/quantize \
--input DeepSeek-V4-Flash-Base-fp16.gguf \
--output DeepSeek-V4-Flash-Base-q4_K.gguf \
--type q4_K
Running 1-million-token inference with the Flash variant:
DS4_MODEL=gguf/DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf \
./ds4_cli -t 1 -p 1048576 -e "Explain quantum computing in simple terms."
Summary
- Flash supports mixed FP4+FP8 plus low-bit formats (
q8_0,q4_K,q2_K,iq2_xxs) viagguf-tools/quants.c, while PRO uses only the mixed FP4+FP8 baseline. - Active parameters differ significantly: 13B for Flash versus 49B for PRO, directly impacting memory requirements.
- Both models share identical KV-cache constants defined in
ds4.c(lines 55–60), implementing sliding-window (128 tokens) and ratio-based compression. - GPU memory requirements scale to approximately 30–40 GB for Flash and 70–90 GB for PRO when processing 1-million-token contexts.
- The
MODEL_CARD.md(lines 78–84) andds4.csource files provide authoritative specifications for precision formats and memory architecture.
Frequently Asked Questions
What is the difference between total and active parameters in ds4?
Total parameters represent the complete weight count stored in the GGUF file—284 billion for Flash and 1.6 trillion for PRO—while active parameters indicate the subset actually utilized during a single forward pass. Flash activates 13 billion parameters per token, whereas PRO activates 49 billion, explaining the disparate memory requirements despite both using MoE (Mixture of Experts) architectures.
Why does Flash support more quantization formats than PRO?
The Flash model includes a dedicated quantization façade in gguf-tools/quants.c that implements low-bit block quantizers (q4_K, q2_K, iq2_xxs) specifically optimized for its smaller expert matrices. The PRO variant’s substantially larger active parameter set (49B) would suffer unacceptable accuracy degradation with these aggressive compressions, so the repository maintains only the mixed FP4+FP8 format for PRO to preserve output quality.
How does the KV cache compression work in both models?
Both Flash and PRO implement an identical compressed KV cache defined in ds4.c (lines 55–60). The system maintains a raw sliding window of 128 recent tokens while compressing older KV rows through alternating ratio-4 and ratio-128 compression layers. A top-512 indexer with 64 heads of 128 dimensions each manages retrieval, ensuring that the per-token memory cost remains constant regardless of the 1-million-token context length.
Can I convert the PRO model to low-bit quantization like q4_K?
No. The gguf-tools/quants.c implementation specifically targets the Flash model family’s architecture and expert weight distributions. Attempting to apply these low-bit quantizers to PRO would require modifying the quantization façade to handle the larger 49B active parameter matrices, which is not supported in the current ds4 codebase. PRO models should be run using the standard mixed FP4+FP8 GGUF files provided in the repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →