# How Prefill Chunking Works with the Metal Graph Backend in ds4

> Explore prefill chunking in ds4 Metal graph backend. Learn how it processes long prompts in token chunks, minimizing kernel launches and preserving KV-cache consistency.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-09

---

**The Metal graph backend in ds4 processes long prompts in configurable token chunks (default 4096) using a single reusable range-capable graph, minimizing kernel launches while preserving KV-cache checkpoint consistency.**

The ds4 inference engine by antirez implements a sophisticated prefill strategy for Apple's Metal backend that avoids processing entire contexts in a single GPU dispatch. Instead of building unique computational graphs for every sequence length, the system evaluates prompts in discrete chunks using a compiled graph that handles arbitrary token sub-ranges. This design balances memory efficiency with computational throughput, enabling the engine to manage contexts exceeding 128K tokens without excessive GPU memory pressure or host-side overhead.

## Default Chunk Size and Configuration

The Metal backend splits prefill operations into **4096-token chunks** by default, a value documented as the canonical setting for the project's benchmark suite in [`README.md`](https://github.com/antirez/ds4/blob/main/README.md). This default represents a calibrated compromise between kernel launch overhead and peak GPU memory utilization during the prefill phase.

Users override this behavior via the `--prefill-chunk N` command-line argument parsed in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c). For example, `--prefill-chunk 2048` reduces the working set size for each graph execution, proving useful when targeting specific KV-cache checkpoint layouts or constrained memory environments. The chosen chunk size directly influences token processing granularity and the resulting logit-storage path layout.

## Range-Capable Graph Architecture

Rather than constructing new per-layer dispatch graphs for each chunk, the implementation in `ds4_metal.m` builds **one range-capable "layer-major" graph** capable of evaluating any contiguous token sub-range. This compiled Metal graph remains resident in GPU memory and is reused across all chunks in a prefill session.

The graph construction logic manages absolute compressor and indexer boundaries internally, ensuring the same computational structure processes tokens 0-4095, 4096-8191, and subsequent ranges without recompilation. This approach eliminates dynamic graph building overhead while maintaining consistent floating-point reduction order across chunk boundaries.

## The Chunk Loop Implementation

The core orchestration logic resides in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), specifically within `ds4_session_sync`. The engine iterates over the prompt length in steps defined by `engine->prefill_chunk`, executing the Metal graph for each slice and checkpointing the resulting KV state.

The loop follows this pattern:

```c
uint32_t chunk = engine->prefill_chunk;          // 4096 by default
for (size_t pos = 0; pos < prompt_len; pos += chunk) {
    size_t cur_len = min(chunk, prompt_len - pos);
    ds4_metal_execute(engine, &prompt[pos], cur_len);   // runs the same graph
    ds4_kv_checkpoint(engine, pos + cur_len);           // store KV state
}

```

Each iteration calls `ds4_metal_execute` with appropriate start and end token indices, allowing the range-capable graph to process the current window. After GPU work completes, `ds4_kv_checkpoint` persists the intermediate key-value cache state, making it available for subsequent decode steps or additional prefill chunks.

## KV Cache Checkpointing and Memory Layout

Between chunks, the system writes KV state to an in-memory cache according to checkpointing logic in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c). Checkpoint placement depends directly on the configured chunk size; smaller chunks create more frequent checkpoints, while larger chunks reduce storage granularity but increase memory required for a single prefill window.

This checkpointing strategy ensures decode steps can immediately reuse prefilled state without recomputing attention over the entire context. The chunk size parameter therefore serves dual purposes: it bounds GPU memory allocation for intermediate activations and determines the granularity of recoverable inference state.

## Performance and Determinism Benefits

Prefill chunking with the Metal backend provides three critical advantages:

- **Reduced Kernel Launch Overhead**: Larger chunks minimize host-GPU synchronizations and command buffer submissions required to process long contexts.
- **Deterministic Execution**: Reusing the same compiled graph across all chunks ensures consistent floating-point reduction ordering, eliminating non-determinism from dynamic graph construction.
- **Predictable Memory Accounting**: The chunk size establishes a hard upper bound on KV memory required for any single prefill operation, preventing out-of-memory errors during context ingestion.

## Practical Usage Examples

Run inference with the default 4096-token chunk size:

```bash
./ds4 -m gguf/DeepSeek-V4-Flash-Q4K.gguf

```

Explicitly configure a 2048-token chunk for stricter checkpoint compatibility:

```bash
./ds4 -m gguf/DeepSeek-V4-Flash-Q4K.gguf --prefill-chunk 2048

```

Benchmark per-chunk throughput using the dedicated test utility:

```bash
./ds4-bench \
  -m gguf/DeepSeek-V4-Flash-Q4K.gguf \
  --ctx-start 0 \
  --ctx-max 65536 \
  --step-incr 4096 \
  --prefill-chunk 4096 \
  --debug

```

The test suite in [`tests/test_metal_session_batch.c`](https://github.com/antirez/ds4/blob/main/tests/test_metal_session_batch.c) verifies correct behavior of Metal prefill chunking, including mixed-prefill scenarios, while [`tests/test_engine_mgpu_placement.c`](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_placement.c) contains helper functions like `ds4_test_planner_prefill_cap` that compute effective chunk sizes for various backend configurations.

## Summary

- The Metal backend in ds4 processes prompts in **4096-token chunks by default**, configurable via `--prefill-chunk` parsed in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c).
- A **single range-capable graph** constructed in `ds4_metal.m` handles all chunk evaluations without recompilation, preserving absolute compressor and indexer boundaries.
- The chunk loop in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) orchestrates execution through `ds4_metal_execute` and persists state via `ds4_kv_checkpoint` after each iteration.
- Chunk size directly impacts **memory usage**, **checkpoint granularity**, and **kernel launch overhead** during prefill.
- This architecture maintains **deterministic floating-point behavior** across arbitrarily long contexts while keeping GPU-resident graphs simple and reusable.

## Frequently Asked Questions

### What is the default prefill chunk size in ds4's Metal backend?

The default prefill chunk size is **4096 tokens**. This value is hardcoded as the canonical setting for the benchmark suite and represents the standard granularity for processing prompts with the Metal graph backend according to the source documentation.

### How does changing the chunk size affect KV-cache checkpointing?

Modifying the chunk size with `--prefill-chunk` changes where the engine places KV-cache checkpoints during prefill. Smaller chunks create more frequent checkpoints, increasing storage granularity but reducing the memory required for any single graph execution. Larger chunks decrease checkpoint frequency but require proportionally more GPU memory per iteration.

### Why does ds4 reuse the same Metal graph for every chunk instead of building new ones?

The system constructs a single **range-capable "layer-major" graph** to eliminate compilation overhead and ensure deterministic execution. Reusing this graph across all chunks prevents floating-point non-determinism that could arise from dynamic graph construction while minimizing host-GPU synchronization points throughout the prefill phase.

### Where is the prefill chunking logic implemented in the source code?

The chunking loop and KV checkpointing reside in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) within functions like `ds4_session_sync`, while the reusable Metal graph execution and range-capable graph construction are implemented in `ds4_metal.m`. Command-line parsing for the `--prefill-chunk` flag occurs in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c).