# Optimal Prefill Chunk Sizes for Different Context Lengths in DS4

> Discover optimal prefill chunk sizes for DS4 context lengths. Learn default settings and customization options for efficient processing with CUDA Tensor-Parallel.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: performance
- Published: 2026-08-05

---

**DS4 uses adaptive prefill chunk sizing: 2048 tokens for CUDA Tensor-Parallel by default, prompt length for short inputs under 4096, 4096 tokens for long inputs (8192 for PRO variants), with full override support via CLI flags and environment variables.**

The DS4 inference engine dynamically selects **prefill chunk sizes**—the number of tokens processed in a single pre-fill pass—based on hardware configuration, model variant, and prompt length. Understanding these defaults helps optimize throughput and memory utilization for your specific deployment.

## How DS4 Determines Prefill Chunk Size

The chunk selection logic resides in two core functions in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c): `ds4_effective_prefill_chunk` (lines 12153‑12157) and `ds4_prefill_cap_for_prompt` (lines 12159‑12177). These implement a priority-ordered decision tree:

| Scenario | Resulting Chunk Size | Implementation |
|----------|---------------------|----------------|
| **Explicit user override** (`--prefill-chunk N` or `DS4_METAL_PREFILL_CHUNK`) | User-specified value | `requested_chunk != 0` branch in `ds4_effective_prefill_chunk` |
| **CUDA Tensor-Parallel enabled**, no override | **2048 tokens** | `DS4_CUDA_TP_DEFAULT_PREFILL_CHUNK` constant, line 12155 |
| **No TP, no override, prompt ≤ 4096** | **Prompt length** (full prompt) | Direct cap to prompt size, lines 12159‑12166 |
| **No TP, no override, prompt > 4096** | **4096 tokens** (8192 for PRO) | Model variant check at lines 12175‑12177 |
| **Metal backend** | Same as above, or env override | `DS4_METAL_PREFILL_CHUNK` check, lines 12166‑12175 |

## Default Prefill Chunk Sizes by Context Length

### Short Prompts (≤ 4096 tokens)

Without CUDA Tensor-Parallel, DS4 processes the entire prompt in one chunk:

```c
// From ds4_prefill_cap_for_prompt (ds4.c, lines 12159-12166)
if (prompt_len <= 4096) {
    return prompt_len;  // Chunk equals prompt length
}

```

This minimizes kernel launch overhead for conversational and short-context workloads.

### Long Prompts (> 4096 tokens)

For prompts exceeding 4096 tokens, DS4 caps the chunk to prevent memory pressure:

```c
// From ds4_prefill_cap_for_prompt (ds4.c, lines 12175-12177)
if (variant == DS4_VARIANT_PRO) return 8192;
return 4096;

```

**Standard variants**: 4096-token chunks
**PRO variants**: 8192-token chunks (higher memory capacity required)

### CUDA Tensor-Parallel Deployments

TP splits computation across GPUs, introducing synchronization points. DS4 defaults to **2048 tokens** to balance parallelism efficiency against communication overhead:

```c
// ds4_effective_prefill_chunk (ds4.c, line 12153-12157)
static uint32_t ds4_effective_prefill_chunk(bool cuda_tensor_parallel,
                                            uint32_t requested_chunk) {
    if (requested_chunk != 0) return requested_chunk;
    return cuda_tensor_parallel ? DS4_TP_DEFAULT_CHUNK : 0;
}

```

## Configuring Prefill Chunk Sizes

### Method 1: CLI Flag (All Backends)

Force a specific chunk size regardless of prompt length:

```bash

# Force 4096-token chunks

ds4-server --prefill-chunk 4096

# Use automatic defaults

ds4-server --prefill-chunk 0

```

The value passes directly to `ds4_effective_prefill_chunk` as `requested_chunk`, bypassing all automatic logic.

### Method 2: Environment Variable (Metal Only)

```bash
export DS4_METAL_PREFILL_CHUNK=8192
ds4-server --backend metal

```

This is evaluated in `ds4_prefill_cap_for_prompt` before other heuristics, making it the highest-priority override for Metal deployments.

### Method 3: Tensor-Parallel with Default

```bash

# 4-GPU tensor parallel, uses 2048-token default

ds4-server --cuda-tensor-parallel --gpus 0,1,2,3 --prefill-chunk 0

```

Unit test validation in [`test_engine_mgpu_placement.c`](https://github.com/antirez/ds4/blob/main/test_engine_mgpu_placement.c) (lines 551‑555) confirms this 2048-token default and verifies explicit overrides propagate correctly.

## Performance Considerations

- **Smaller chunks** (2048): Lower peak memory, better TP scaling, more kernel launches
- **Larger chunks** (4096/8192): Higher throughput for long contexts, increased memory pressure
- **Prompt-length matching** (≤ 4096, no TP): Minimal overhead for short inputs

The PRO variant's 8192-token ceiling for long prompts reflects its architectural support for extended context windows and larger activation caching.

## Summary

- DS4 selects **2048 tokens** for CUDA Tensor-Parallel by default
- **Prompt-length chunks** are used for short inputs (≤ 4096) without TP
- **4096 tokens** (8192 for PRO) cap long-input processing without TP
- **Explicit overrides** via `--prefill-chunk` or `DS4_METAL_PREFILL_CHUNK` take precedence
- Core logic lives in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) functions `ds4_effective_prefill_chunk` and `ds4_prefill_cap_for_prompt`

## Frequently Asked Questions

### What is a prefill chunk in DS4?

A **prefill chunk** is the number of input tokens DS4 processes in a single forward pass during the context-encoding phase. Smaller chunks reduce memory but increase kernel launches; larger chunks improve throughput for long sequences. The engine automatically sizes these based on hardware and prompt characteristics, or accepts explicit user configuration.

### When should I override the default prefill chunk size?

Override when you observe memory pressure (reduce chunk size) or want to maximize throughput on high-memory systems (increase chunk size). The 2048-token TP default is conservative; scaling to 4096 may improve performance on NVLink-connected GPUs. Always benchmark with your specific model and sequence length distribution.

### Why does the PRO variant use 8192 tokens instead of 4096?

The **PRO variant** is architected for extended context windows and larger cache allocations. The doubled chunk ceiling in `ds4_prefill_cap_for_prompt` (line 12176) leverages this capacity to reduce iteration count during prefill of long documents, trading memory for latency reduction.

### How do I verify which chunk size is actually being used?

DS4 does not currently log the active prefill chunk at startup. To confirm behavior, you can instrument `ds4_prefill_cap_for_prompt` locally or trace the return value through a debugger. The unit test in [`test_engine_mgpu_placement.c`](https://github.com/antirez/ds4/blob/main/test_engine_mgpu_placement.c) demonstrates the expected values for standard configurations.