What Is the Significance of `VLLM_USE_BREAKABLE_CUDAGRAPH=0` in DeepSeek-v4-Flash-DSpark?

Setting VLLM_USE_BREAKABLE_CUDAGRAPH=0 disables breakable CUDA graphs in the vLLM inference engine, forcing traditional eager kernel execution that trades throughput for improved debuggability and broader hardware compatibility.

The environment variable VLLM_USE_BREAKABLE_CUDAGRAPH controls a critical performance feature in the MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository's patched vLLM implementation. When set to 0, it prevents the engine from capturing reusable CUDA graphs, which is essential for debugging memory issues or running on hardware with strict kernel scheduling requirements. Understanding this toggle helps operators balance between maximum inference speed and operational stability when serving the DeepSeek-v4-Flash model on DGX-Spark clusters.

How Breakable CUDA Graphs Optimize Inference

The Graph Capture Mechanism

By default, modern vLLM engines use breakable CUDA graphs to eliminate CPU launch overhead. When enabled, the engine records kernel sequences into a reusable graph object that can be replayed with minimal CPU intervention. As implemented in the custom vLLM patch located at vllm_patch_gb10/vllm_gb10_hybrid_nvfp4/config.py, this feature reads the VLLM_USE_BREAKABLE_CUDAGRAPH flag during engine initialization to decide whether to enter graph-capture mode.

Dynamic Shape Handling

Unlike static CUDA graphs, breakable graphs support dynamic batch sizes and sequence lengths without requiring a full graph rebuild. This is particularly important for the DSpark execution engine, which handles heterogeneous request batches. The configuration parser in the repository checks for truthy values (1, true, or O) to enable this capability, as documented in docs/ENVS.md.

What Happens When You Set VLLM_USE_BREAKABLE_CUDAGRAPH=0

Reversion to Eager Execution

Setting the variable to 0 forces vLLM to launch kernels individually via the CUDA driver API. While this increases per-token latency by 10–30% in benchmark tests (see scripts/benchmark-0731.py), it eliminates the memory overhead associated with graph storage and allows the host CPU to interleave other tasks between kernel launches.

Enhanced Debugging and Compatibility

Eager mode provides clearer stack traces and deterministic error reporting when kernels fail. This is invaluable when integrating patches such as patches/hotfix-vllm-rope-swa-fix.py, which modify attention kernels. Additionally, some legacy GPU drivers or virtualized environments (e.g., certain DGX-Spark partitions) do not support graph capture, making =0 a requirement for basic functionality.

Configuration and Implementation Details

The repository reads this environment variable in the configuration module vllm_patch_gb10/vllm_gb10_hybrid_nvfp4/config.py. The parser treats any value other than 0, false, or empty string as enabling the feature. According to the source code analysis, this behavior ensures backward compatibility while allowing explicit opt-out for troubleshooting.


# Example: Explicitly disabling breakable CUDA graphs for debugging

import os

os.environ["VLLM_USE_BREAKABLE_CUDAGRAPH"] = "0"

from vllm import LLMEngine, SamplingParams

# The engine will now run in eager mode, ignoring graph optimizations

engine = LLMEngine(
    model="deepseek-v4-flash",
    tokenizer_path="/models/deepseek-v4",
    dtype="float16"
)

# Each generate call launches kernels individually

params = SamplingParams(temperature=0.7, max_tokens=128)
output = engine.generate("Explain quantum computing.", sampling_params=params)

# Shell example for containerized deployments

export VLLM_USE_BREAKABLE_CUDAGRAPH=0
python -m vllm.entrypoints.api_server \
    --model /models/deepseek-v4-flash \
    --tensor-parallel-size 8

Summary

  • Performance trade-off: Setting VLLM_USE_BREAKABLE_CUDAGRAPH=0 disables CUDA graph caching, increasing latency but improving compatibility.
  • Debug workflow: Eager execution mode provides clearer error messages when integrating patches like hotfix-vllm-rope-swa-fix.py.
  • Configuration location: The flag is parsed in vllm_patch_gb10/vllm_gb10_hybrid_nvfp4/config.py and documented in docs/ENVS.md.
  • Use case: Required for legacy drivers, virtualized GPUs, or when profiling individual kernel performance via benchmark-0731.py.

Frequently Asked Questions

What is a breakable CUDA graph?

A breakable CUDA graph is a captured sequence of GPU kernels that can be replayed with varying input shapes without full recompilation. In MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark, this allows vLLM to handle dynamic batching efficiently while minimizing CPU overhead between inference steps.

When should I set VLLM_USE_BREAKABLE_CUDAGRAPH=0?

Disable breakable graphs when debugging kernel errors, running on unsupported hardware drivers, or deploying in virtualized DGX-Spark environments where graph capture fails. The setting is also useful when using external profiling tools that require explicit kernel boundaries.

Does this variable affect multi-GPU tensor parallelism?

Yes. When set to 0, each GPU in a tensor-parallel group launches its kernels independently rather than through synchronized graph replay. This can slightly desynchronize workers but prevents hangs caused by graph metadata mismatches across GPUs, as noted in the repository's DSpark integration notes.

Where is this documented in the repository?

The variable is formally documented in docs/ENVS.md, which lists all environment variables supported by the DeepSeek-v4-Flash-DSpark distribution. Implementation details appear in vllm_patch_gb10/vllm_gb10_hybrid_nvfp4/config.py, where the parser converts the string value into a boolean configuration flag.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →