MAX_NUM_BATCHED_TOKENS Default Value in DeepSeek-v4-Flash-DSpark: Configuration Guide

The default value for MAX_NUM_BATCHED_TOKENS is 8192 tokens, defining the maximum number of tokens processed in a single pre-fill batch unless explicitly overridden by environment variables.

The MAX_NUM_BATCHED_TOKENS parameter controls batch size limits during inference in the MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository. This setting directly impacts GPU memory allocation and throughput on DGX infrastructure, with the system automatically falling back to 8192 when the variable remains unset.

Where the 8192 Default is Defined

Launcher Script Fallback

In start-deepseek-v4-flash-dspark.sh, the effective value resolves using bash parameter expansion with the fallback syntax ${MAX_NUM_BATCHED_TOKENS:-8192} at lines 1274-1280. This ensures the deployment script prints and utilizes 8192 when the environment variable remains undefined.

Validation and Environment Configuration

The validate-dspark-config.sh script employs the identical default expression at lines 150-151 to verify configuration validity before startup. Additionally, the .env.dspark.example file explicitly sets MAX_NUM_BATCHED_TOKENS=8192 at lines 229-230, establishing the baseline for production deployments.

How the Default Value is Applied

The 8192-token limit propagates through the vLLM scheduler configuration. When the Python scheduler initializes, it reads the resolved value from environment variables, defaulting to 8192 if unspecified. This parameter directly constrains the pre-fill batch size, determining how many input tokens the model processes simultaneously during the attention computation phase.

Configuration Examples

Using the Default (8192)


# Execute launcher without explicit configuration

./start-deepseek-v4-flash-dspark.sh

# Output indicates: max batched tokens: 8192

Override for Larger Batches


# Process longer sequences at the cost of increased VRAM

export MAX_NUM_BATCHED_TOKENS=16384
./start-deepseek-v4-flash-dspark.sh

Python Access Pattern


# Accessing the effective value within the scheduler configuration

from vllm.config import SchedulerConfig

scheduler_config = SchedulerConfig(...)
max_batch = scheduler_config.max_num_batched_tokens  # Defaults to 8192

Summary

  • The default is 8192 tokens, defined via bash parameter expansion in start-deepseek-v4-flash-dspark.sh and validate-dspark-config.sh.
  • Three locations confirm this default: the launcher script (lines 1274-1280), validation script (lines 150-151), and example environment file (lines 229-230).
  • Override capability exists by setting the environment variable before executing deployment scripts.
  • Memory impact scales linearly with this value; higher limits require proportionally more GPU memory during the pre-fill phase.

Frequently Asked Questions

What happens if MAX_NUM_BATCHED_TOKENS exceeds available GPU memory?

The vLLM scheduler will attempt to allocate the requested batch size, potentially triggering CUDA out-of-memory errors during the pre-fill phase. The system does not automatically clamp this value to available VRAM, making the 8192 default a conservative starting point for DGX-2 deployments.

Can I set MAX_NUM_BATCHED_TOKENS lower than 8192 for memory-constrained environments?

Yes. Setting the variable to values like 4096 or 2048 reduces peak memory consumption at the cost of throughput. Modify MAX_NUM_BATCHED_TOKENS in your .env.dspark file or export it before running start-deepseek-v4-flash-dspark.sh.

Does this default apply to both pre-fill and decode phases?

No. The MAX_NUM_BATCHED_TOKENS parameter specifically governs the pre-fill batch size—the initial forward pass that processes input prompt tokens. Decode-phase batching follows separate scheduling logic managed by max_num_seqs and related vLLM configuration parameters.

Where should I verify the effective value at runtime?

Check the startup logs generated by start-deepseek-v4-flash-dspark.sh, which echoes the resolved value using the ${MAX_NUM_BATCHED_TOKENS:-8192} expansion. Alternatively, inspect the scheduler configuration object in Python to confirm the loaded integer value matches your expectations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →