DeepSeek V4 Flash vs PRO Model Support in DS4: Key Differences Explained

The DeepSeek V4 Flash model is DS4's primary, default target, while PRO support exists only for high-memory or distributed setups requiring explicit configuration and split-layer GGUF files.

DS4 (DwarfStar), the open-source inference engine from Salvatore Sanfilippo (antirez/ds4), is specifically engineered around two tiers of DeepSeek V4 model support. Understanding the distinction between Flash and PRO deployments is critical for selecting the right configuration for your hardware.

What Is DeepSeek V4 Flash in DS4?

Flash is the default and optimized model path that DS4 targets.

  • The repository ships with ./ds4flash.gguf as the canonical reference model
  • Standard inference binaries—ds4 and ds4-server—use this model automatically
  • All kernel optimizations, quantization paths, and speculative decoding implementations in ds4.c are built around the Flash layout

Flash models run on typical high-end laptops and workstations without special configuration. In ds4.c, the core inference engine implements Flash-first kernels and KV cache handling, making this the zero-friction path for most users.

What Is DeepSeek V4 PRO Model Support?

PRO support is optional and hardware-intensive, designed for very different deployment scenarios.

Hardware Requirements

  • Minimum: Approximately 512 GB RAM
  • Typical setup: Multi-machine or multi-GPU distributed deployment

Model Distribution

PRO GGUFs are split across layer ranges rather than shipped as single files. The download_model.sh script provides explicit targets for these splits:

  • pro-q4-layers00-30 — first half of layers
  • pro-q4-layers31-output — remaining layers through output

This split architecture enables distributed inference where different machines handle different model segments.

Running Flash vs PRO: Practical Examples

Default Flash Model (Single-GGUF)


# Run with the built-in Flash model

./ds4 -m ds4flash.gguf --temp 0.7 "Explain quantum entanglement in simple terms."

The -m flag is optional when pointing to the default filename; run-nvidia-tp-server.sh demonstrates this by selecting Flash automatically when no override is specified.

PRO Model (Single Machine, Quantized)


# Download PRO q2-imatrix variant

./download_model.sh pro-q2-imatrix

# Explicit model selection required

./ds4 -m gguf/deepseek-v4-pro.q2.gguf --temp 0.7 "Explain quantum entanglement in simple terms."

Note the mandatory -m <path>—PRO models never run as defaults.

Distributed PRO Setup (Two-Machine Split)


# Machine 1: layers 0-30

./download_model.sh pro-q4-layers00-30
./ds4-server -m gguf/pro-q4-layers00-30.gguf --port 8000

# Machine 2: layers 31-output

./download_model.sh pro-q4-layers31-output
./ds4-server -m gguf/pro-q4-layers31-output.gguf --port 8001

This architecture allows PRO models to operate across memory boundaries that would exceed single-machine capacity.

Source Code Evidence

File Relevance
ds4.c Core inference engine with Flash-optimized kernels and KV cache management
run-nvidia-tp-server.sh Demonstrates Flash-as-default startup behavior
download_model.sh Provides both Flash and PRO (including split-layer) download targets
gguf-tools/quality-testing/README.md Documents continuation vectors for benchmarking both model tiers

The README at lines 5-8 establishes Flash as primary, while lines 119-131 detail the PRO split-layer distribution system.

Summary

  • Flash = single GGUF, default target, optimized kernels, runs on standard high-end hardware
  • PRO = split-layer GGUFs, explicit selection only, requires ≥512 GB RAM or distributed setup
  • Both use identical quantization schemes (q2, q4 variants with imatrix), but PRO scales through horizontal distribution rather than vertical compression

Frequently Asked Questions

Can I run PRO models on a single GPU?

Only if that GPU has sufficient VRAM for your chosen quantization level. The pro-q2-imatrix variant reduces memory requirements but still typically exceeds consumer GPU capacity. Most PRO deployments use CPU RAM (≥512 GB) or multi-node GPU configurations.

Why is Flash the default rather than the smallest model?

DS4 prioritizes quality-per-parameter over raw size. Flash represents the optimal balance of inference speed, memory footprint, and output quality that the engineered kernels in ds4.c can exploit. Smaller models would leave optimization headroom unused.

Do Flash and PRO produce identical outputs?

No. The gguf-tools/quality-testing/README.md documents distinct continuation vectors used to benchmark each tier. PRO models maintain more parameters and different layer distributions, yielding measurably different (typically higher-capability) generation characteristics at the cost of deployment complexity.

Is speculative decoding available for PRO models?

The speculative decoding paths in ds4.c are Flash-optimized. PRO models may run with standard autoregressive generation; check the latest README.md for any updates on speculative support for distributed PRO configurations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →