DeepSeek-v4-Flash-DSpark MAX_NUM_SEQS: Maximum Concurrency Explained

The DeepSeek-v4-Flash-DSpark inference engine supports up to 16 concurrent sequences when configured with --max-num-seqs 16, with a conservative default of 6 slots to balance memory usage and throughput.

DeepSeek-v4-Flash-DSpark implements sequence-level parallelism through a configurable concurrency limit controlled by MAX_NUM_SEQS. This parameter determines how many requests the vLLM-based scheduler can process simultaneously, directly impacting throughput and GPU memory utilization. Understanding how to tune this value is essential for optimizing inference performance on multi-GPU DGX deployments.

How MAX_NUM_SEQS Controls Concurrency

The MAX_NUM_SEQS parameter defines the number of active sequence slots available in the KV cache manager. Each concurrent sequence requires dedicated GPU memory for key-value tensors, making this setting a critical trade-off between throughput and memory consumption.

In validate-dspark-config.sh, the default value is explicitly set:


# Lines 149-153 in validate-dspark-config.sh

if [ -z "$MAX_NUM_SEQS" ]; then
    echo "MAX_NUM_SEQS not set, using default: 6"
    MAX_NUM_SEQS=6
else
    echo "MAX_NUM_SEQS: $MAX_NUM_SEQS"
fi

This fallback ensures predictable memory usage out of the box, preventing out-of-memory errors on standard configurations.

Default vs. Maximum Supported Values

Configuration Value Source
Default 6 validate-dspark-config.sh (lines 149-153)
Documented maximum example 16 README.md (lines 534-580)
Hard upper bound None enforced Runtime-limited by GPU memory

The repository does not implement a hardcoded maximum. The practical ceiling emerges from KV cache capacity: each additional sequence multiplies memory requirements by sequence length, hidden dimension, and number of layers.

Launching with Custom Concurrency

The start-tp3.sh launcher accepts --max-num-seqs as a passthrough argument to the underlying vLLM engine (lines 14-20):


# Default launch with 6 concurrent sequences

./start-tp3.sh

# Moderate concurrency increase to 12

./start-tp3.sh --max-num-seqs 12

# Maximum documented configuration (16 slots)

./start-tp3.sh --max-num-seqs 16

The launcher script forwards this flag directly to vLLM's AsyncLLMEngine, which initializes the scheduler with the specified capacity.

Environment Variable Configuration

You can also set MAX_NUM_SEQS via environment variable before launch:

import os
import subprocess

# Configure 12 concurrent sequences

os.environ["MAX_NUM_SEQS"] = "12"

# Launch the inference server

subprocess.run(["./start-tp3.sh", "--tensor-parallel-size", "3"])

This approach integrates cleanly with containerized deployments and orchestration systems where command-line flags are less convenient.

Memory Considerations for High Concurrency

Raising MAX_NUM_SEQS linearly increases KV cache allocation. For DeepSeek-v4's architecture, estimate required memory per sequence:

  • Per-token KV cache: 2 × num_layers × hidden_size × num_kv_heads × precision_bytes
  • Total cache: max_seq_length × max_num_seqs × per_token_size

On 2×DGX configurations with H100 GPUs, the documented maximum of 16 sequences assumes typical sequence lengths (4K-8K tokens). For longer contexts or batching scenarios, reduce concurrency proportionally.

Code Reference: Configuration Propagation

The configuration flows through three key files as implemented in MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark:

  1. validate-dspark-config.sh — Validates and defaults the setting
  2. start-tp3.sh — Accepts and forwards the CLI flag
  3. vLLM engine initialization — Receives via EngineArgs

No source-level enforcement caps the value; vLLM's scheduler will attempt to allocate requested slots until GPU memory exhaustion triggers runtime errors.

Summary

  • Default concurrency: 6 sequences, defined in validate-dspark-config.sh for safe baseline operation
  • Demonstrated maximum: 16 sequences, shown in official documentation examples
  • Configuration methods: --max-num-seqs CLI flag or MAX_NUM_SEQS environment variable
  • Practical limit: GPU memory capacity for KV cache, not a software-enforced cap
  • Tuning recommendation: Start at 6, profile memory usage, increase incrementally to maximize throughput without OOM errors

Frequently Asked Questions

What happens if I set MAX_NUM_SEQS too high?

The vLLM engine will fail to initialize or crash with a CUDA out-of-memory error during the first forward pass. No automatic fallback reduces concurrency; you must restart with a lower value. Profile memory with nvidia-smi or PyTorch memory stats to find your hardware's practical limit.

Does higher MAX_NUM_SEQS always improve throughput?

Up to a point. Additional concurrent sequences improve GPU utilization through better memory bandwidth hiding, but diminishing returns occur when overhead from context switching and memory fragmentation outweighs gains. Benchmark with your specific prompt length distribution to identify the optimal setting.

Can different GPU configurations support different maximums?

Yes. The 2×DGX reference hardware (16×H100 80GB) supports 16 sequences at typical lengths. Smaller deployments should scale proportionally: single-node A100 systems often stabilize around 8-12 sequences for full-context workloads.

Is MAX_NUM_SEQS the same as batch size?

Conceptually related but operationally distinct. Batch size refers to tokens processed together in a single forward pass, while MAX_NUM_SEQS bounds how many independent sequences remain active in the scheduler. Continuous batching in vLLM dynamically packs tokens across these slots, so effective throughput depends on both parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →