# DeepSeek V4 Flash vs PRO Model Support in DS4: Key Differences Explained

> Understand DeepSeek V4 Flash vs PRO model support in DS4. Learn when to use Flash (default) and PRO (high-memory/distributed) for optimal performance.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-05

---

**The DeepSeek V4 Flash model is DS4's primary, default target, while PRO support exists only for high-memory or distributed setups requiring explicit configuration and split-layer GGUF files.**

DS4 (DwarfStar), the open-source inference engine from Salvatore Sanfilippo (`antirez/ds4`), is specifically engineered around two tiers of DeepSeek V4 model support. Understanding the distinction between **Flash** and **PRO** deployments is critical for selecting the right configuration for your hardware.

## What Is DeepSeek V4 Flash in DS4?

Flash is the **default and optimized** model path that DS4 targets.

- The repository ships with `./ds4flash.gguf` as the canonical reference model
- Standard inference binaries—`ds4` and `ds4-server`—use this model automatically
- All kernel optimizations, quantization paths, and speculative decoding implementations in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) are built around the Flash layout

Flash models run on typical high-end laptops and workstations without special configuration. In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the core inference engine implements **Flash-first kernels and KV cache handling**, making this the zero-friction path for most users.

## What Is DeepSeek V4 PRO Model Support?

PRO support is **optional and hardware-intensive**, designed for very different deployment scenarios.

### Hardware Requirements

- **Minimum**: Approximately 512 GB RAM
- **Typical setup**: Multi-machine or multi-GPU distributed deployment

### Model Distribution

PRO GGUFs are **split across layer ranges** rather than shipped as single files. The [`download_model.sh`](https://github.com/antirez/ds4/blob/main/download_model.sh) script provides explicit targets for these splits:

- `pro-q4-layers00-30` — first half of layers
- `pro-q4-layers31-output` — remaining layers through output

This split architecture enables distributed inference where different machines handle different model segments.

## Running Flash vs PRO: Practical Examples

### Default Flash Model (Single-GGUF)

```bash

# Run with the built-in Flash model

./ds4 -m ds4flash.gguf --temp 0.7 "Explain quantum entanglement in simple terms."

```

The `-m` flag is optional when pointing to the default filename; [`run-nvidia-tp-server.sh`](https://github.com/antirez/ds4/blob/main/run-nvidia-tp-server.sh) demonstrates this by selecting Flash automatically when no override is specified.

### PRO Model (Single Machine, Quantized)

```bash

# Download PRO q2-imatrix variant

./download_model.sh pro-q2-imatrix

# Explicit model selection required

./ds4 -m gguf/deepseek-v4-pro.q2.gguf --temp 0.7 "Explain quantum entanglement in simple terms."

```

Note the **mandatory `-m <path>`**—PRO models never run as defaults.

### Distributed PRO Setup (Two-Machine Split)

```bash

# Machine 1: layers 0-30

./download_model.sh pro-q4-layers00-30
./ds4-server -m gguf/pro-q4-layers00-30.gguf --port 8000

# Machine 2: layers 31-output

./download_model.sh pro-q4-layers31-output
./ds4-server -m gguf/pro-q4-layers31-output.gguf --port 8001

```

This architecture allows PRO models to operate across memory boundaries that would exceed single-machine capacity.

## Source Code Evidence

| File | Relevance |
|------|-----------|
| [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) | Core inference engine with Flash-optimized kernels and KV cache management |
| [`run-nvidia-tp-server.sh`](https://github.com/antirez/ds4/blob/main/run-nvidia-tp-server.sh) | Demonstrates Flash-as-default startup behavior |
| [`download_model.sh`](https://github.com/antirez/ds4/blob/main/download_model.sh) | Provides both Flash and PRO (including split-layer) download targets |
| [`gguf-tools/quality-testing/README.md`](https://github.com/antirez/ds4/blob/main/gguf-tools/quality-testing/README.md) | Documents continuation vectors for benchmarking both model tiers |

The README at lines 5-8 establishes Flash as primary, while lines 119-131 detail the PRO split-layer distribution system.

## Summary

- **Flash** = single GGUF, default target, optimized kernels, runs on standard high-end hardware
- **PRO** = split-layer GGUFs, explicit selection only, requires ≥512 GB RAM or distributed setup
- Both use identical quantization schemes (q2, q4 variants with imatrix), but PRO scales through **horizontal distribution** rather than vertical compression

## Frequently Asked Questions

### Can I run PRO models on a single GPU?

Only if that GPU has sufficient VRAM for your chosen quantization level. The `pro-q2-imatrix` variant reduces memory requirements but still typically exceeds consumer GPU capacity. Most PRO deployments use **CPU RAM** (≥512 GB) or **multi-node GPU** configurations.

### Why is Flash the default rather than the smallest model?

DS4 prioritizes **quality-per-parameter** over raw size. Flash represents the optimal balance of inference speed, memory footprint, and output quality that the engineered kernels in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) can exploit. Smaller models would leave optimization headroom unused.

### Do Flash and PRO produce identical outputs?

No. The [`gguf-tools/quality-testing/README.md`](https://github.com/antirez/ds4/blob/main/gguf-tools/quality-testing/README.md) documents distinct **continuation vectors** used to benchmark each tier. PRO models maintain more parameters and different layer distributions, yielding measurably different (typically higher-capability) generation characteristics at the cost of deployment complexity.

### Is speculative decoding available for PRO models?

The speculative decoding paths in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) are **Flash-optimized**. PRO models may run with standard autoregressive generation; check the latest [`README.md`](https://github.com/antirez/ds4/blob/main/README.md) for any updates on speculative support for distributed PRO configurations.