Optimal Prefill Chunk Sizes for Different Context Lengths in DS4

DS4 uses adaptive prefill chunk sizing: 2048 tokens for CUDA Tensor-Parallel by default, prompt length for short inputs under 4096, 4096 tokens for long inputs (8192 for PRO variants), with full override support via CLI flags and environment variables.

The DS4 inference engine dynamically selects prefill chunk sizes—the number of tokens processed in a single pre-fill pass—based on hardware configuration, model variant, and prompt length. Understanding these defaults helps optimize throughput and memory utilization for your specific deployment.

How DS4 Determines Prefill Chunk Size

The chunk selection logic resides in two core functions in ds4.c: ds4_effective_prefill_chunk (lines 12153‑12157) and ds4_prefill_cap_for_prompt (lines 12159‑12177). These implement a priority-ordered decision tree:

Scenario Resulting Chunk Size Implementation
Explicit user override (--prefill-chunk N or DS4_METAL_PREFILL_CHUNK) User-specified value requested_chunk != 0 branch in ds4_effective_prefill_chunk
CUDA Tensor-Parallel enabled, no override 2048 tokens DS4_CUDA_TP_DEFAULT_PREFILL_CHUNK constant, line 12155
No TP, no override, prompt ≤ 4096 Prompt length (full prompt) Direct cap to prompt size, lines 12159‑12166
No TP, no override, prompt > 4096 4096 tokens (8192 for PRO) Model variant check at lines 12175‑12177
Metal backend Same as above, or env override DS4_METAL_PREFILL_CHUNK check, lines 12166‑12175

Default Prefill Chunk Sizes by Context Length

Short Prompts (≤ 4096 tokens)

Without CUDA Tensor-Parallel, DS4 processes the entire prompt in one chunk:

// From ds4_prefill_cap_for_prompt (ds4.c, lines 12159-12166)
if (prompt_len <= 4096) {
    return prompt_len;  // Chunk equals prompt length
}

This minimizes kernel launch overhead for conversational and short-context workloads.

Long Prompts (> 4096 tokens)

For prompts exceeding 4096 tokens, DS4 caps the chunk to prevent memory pressure:

// From ds4_prefill_cap_for_prompt (ds4.c, lines 12175-12177)
if (variant == DS4_VARIANT_PRO) return 8192;
return 4096;

Standard variants: 4096-token chunks PRO variants: 8192-token chunks (higher memory capacity required)

CUDA Tensor-Parallel Deployments

TP splits computation across GPUs, introducing synchronization points. DS4 defaults to 2048 tokens to balance parallelism efficiency against communication overhead:

// ds4_effective_prefill_chunk (ds4.c, line 12153-12157)
static uint32_t ds4_effective_prefill_chunk(bool cuda_tensor_parallel,
                                            uint32_t requested_chunk) {
    if (requested_chunk != 0) return requested_chunk;
    return cuda_tensor_parallel ? DS4_TP_DEFAULT_CHUNK : 0;
}

Configuring Prefill Chunk Sizes

Method 1: CLI Flag (All Backends)

Force a specific chunk size regardless of prompt length:


# Force 4096-token chunks

ds4-server --prefill-chunk 4096

# Use automatic defaults

ds4-server --prefill-chunk 0

The value passes directly to ds4_effective_prefill_chunk as requested_chunk, bypassing all automatic logic.

Method 2: Environment Variable (Metal Only)

export DS4_METAL_PREFILL_CHUNK=8192
ds4-server --backend metal

This is evaluated in ds4_prefill_cap_for_prompt before other heuristics, making it the highest-priority override for Metal deployments.

Method 3: Tensor-Parallel with Default


# 4-GPU tensor parallel, uses 2048-token default

ds4-server --cuda-tensor-parallel --gpus 0,1,2,3 --prefill-chunk 0

Unit test validation in test_engine_mgpu_placement.c (lines 551‑555) confirms this 2048-token default and verifies explicit overrides propagate correctly.

Performance Considerations

  • Smaller chunks (2048): Lower peak memory, better TP scaling, more kernel launches
  • Larger chunks (4096/8192): Higher throughput for long contexts, increased memory pressure
  • Prompt-length matching (≤ 4096, no TP): Minimal overhead for short inputs

The PRO variant's 8192-token ceiling for long prompts reflects its architectural support for extended context windows and larger activation caching.

Summary

  • DS4 selects 2048 tokens for CUDA Tensor-Parallel by default
  • Prompt-length chunks are used for short inputs (≤ 4096) without TP
  • 4096 tokens (8192 for PRO) cap long-input processing without TP
  • Explicit overrides via --prefill-chunk or DS4_METAL_PREFILL_CHUNK take precedence
  • Core logic lives in ds4.c functions ds4_effective_prefill_chunk and ds4_prefill_cap_for_prompt

Frequently Asked Questions

What is a prefill chunk in DS4?

A prefill chunk is the number of input tokens DS4 processes in a single forward pass during the context-encoding phase. Smaller chunks reduce memory but increase kernel launches; larger chunks improve throughput for long sequences. The engine automatically sizes these based on hardware and prompt characteristics, or accepts explicit user configuration.

When should I override the default prefill chunk size?

Override when you observe memory pressure (reduce chunk size) or want to maximize throughput on high-memory systems (increase chunk size). The 2048-token TP default is conservative; scaling to 4096 may improve performance on NVLink-connected GPUs. Always benchmark with your specific model and sequence length distribution.

Why does the PRO variant use 8192 tokens instead of 4096?

The PRO variant is architected for extended context windows and larger cache allocations. The doubled chunk ceiling in ds4_prefill_cap_for_prompt (line 12176) leverages this capacity to reduce iteration count during prefill of long documents, trading memory for latency reduction.

How do I verify which chunk size is actually being used?

DS4 does not currently log the active prefill chunk at startup. To confirm behavior, you can instrument ds4_prefill_cap_for_prompt locally or trace the return value through a debugger. The unit test in test_engine_mgpu_placement.c demonstrates the expected values for standard configurations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →