How Prefill Chunking Works with the Metal Graph Backend in ds4
The Metal graph backend in ds4 processes long prompts in configurable token chunks (default 4096) using a single reusable range-capable graph, minimizing kernel launches while preserving KV-cache checkpoint consistency.
The ds4 inference engine by antirez implements a sophisticated prefill strategy for Apple's Metal backend that avoids processing entire contexts in a single GPU dispatch. Instead of building unique computational graphs for every sequence length, the system evaluates prompts in discrete chunks using a compiled graph that handles arbitrary token sub-ranges. This design balances memory efficiency with computational throughput, enabling the engine to manage contexts exceeding 128K tokens without excessive GPU memory pressure or host-side overhead.
Default Chunk Size and Configuration
The Metal backend splits prefill operations into 4096-token chunks by default, a value documented as the canonical setting for the project's benchmark suite in README.md. This default represents a calibrated compromise between kernel launch overhead and peak GPU memory utilization during the prefill phase.
Users override this behavior via the --prefill-chunk N command-line argument parsed in ds4_cli.c. For example, --prefill-chunk 2048 reduces the working set size for each graph execution, proving useful when targeting specific KV-cache checkpoint layouts or constrained memory environments. The chosen chunk size directly influences token processing granularity and the resulting logit-storage path layout.
Range-Capable Graph Architecture
Rather than constructing new per-layer dispatch graphs for each chunk, the implementation in ds4_metal.m builds one range-capable "layer-major" graph capable of evaluating any contiguous token sub-range. This compiled Metal graph remains resident in GPU memory and is reused across all chunks in a prefill session.
The graph construction logic manages absolute compressor and indexer boundaries internally, ensuring the same computational structure processes tokens 0-4095, 4096-8191, and subsequent ranges without recompilation. This approach eliminates dynamic graph building overhead while maintaining consistent floating-point reduction order across chunk boundaries.
The Chunk Loop Implementation
The core orchestration logic resides in ds4.c, specifically within ds4_session_sync. The engine iterates over the prompt length in steps defined by engine->prefill_chunk, executing the Metal graph for each slice and checkpointing the resulting KV state.
The loop follows this pattern:
uint32_t chunk = engine->prefill_chunk; // 4096 by default
for (size_t pos = 0; pos < prompt_len; pos += chunk) {
size_t cur_len = min(chunk, prompt_len - pos);
ds4_metal_execute(engine, &prompt[pos], cur_len); // runs the same graph
ds4_kv_checkpoint(engine, pos + cur_len); // store KV state
}
Each iteration calls ds4_metal_execute with appropriate start and end token indices, allowing the range-capable graph to process the current window. After GPU work completes, ds4_kv_checkpoint persists the intermediate key-value cache state, making it available for subsequent decode steps or additional prefill chunks.
KV Cache Checkpointing and Memory Layout
Between chunks, the system writes KV state to an in-memory cache according to checkpointing logic in ds4.c. Checkpoint placement depends directly on the configured chunk size; smaller chunks create more frequent checkpoints, while larger chunks reduce storage granularity but increase memory required for a single prefill window.
This checkpointing strategy ensures decode steps can immediately reuse prefilled state without recomputing attention over the entire context. The chunk size parameter therefore serves dual purposes: it bounds GPU memory allocation for intermediate activations and determines the granularity of recoverable inference state.
Performance and Determinism Benefits
Prefill chunking with the Metal backend provides three critical advantages:
- Reduced Kernel Launch Overhead: Larger chunks minimize host-GPU synchronizations and command buffer submissions required to process long contexts.
- Deterministic Execution: Reusing the same compiled graph across all chunks ensures consistent floating-point reduction ordering, eliminating non-determinism from dynamic graph construction.
- Predictable Memory Accounting: The chunk size establishes a hard upper bound on KV memory required for any single prefill operation, preventing out-of-memory errors during context ingestion.
Practical Usage Examples
Run inference with the default 4096-token chunk size:
./ds4 -m gguf/DeepSeek-V4-Flash-Q4K.gguf
Explicitly configure a 2048-token chunk for stricter checkpoint compatibility:
./ds4 -m gguf/DeepSeek-V4-Flash-Q4K.gguf --prefill-chunk 2048
Benchmark per-chunk throughput using the dedicated test utility:
./ds4-bench \
-m gguf/DeepSeek-V4-Flash-Q4K.gguf \
--ctx-start 0 \
--ctx-max 65536 \
--step-incr 4096 \
--prefill-chunk 4096 \
--debug
The test suite in tests/test_metal_session_batch.c verifies correct behavior of Metal prefill chunking, including mixed-prefill scenarios, while tests/test_engine_mgpu_placement.c contains helper functions like ds4_test_planner_prefill_cap that compute effective chunk sizes for various backend configurations.
Summary
- The Metal backend in ds4 processes prompts in 4096-token chunks by default, configurable via
--prefill-chunkparsed inds4_cli.c. - A single range-capable graph constructed in
ds4_metal.mhandles all chunk evaluations without recompilation, preserving absolute compressor and indexer boundaries. - The chunk loop in
ds4.corchestrates execution throughds4_metal_executeand persists state viads4_kv_checkpointafter each iteration. - Chunk size directly impacts memory usage, checkpoint granularity, and kernel launch overhead during prefill.
- This architecture maintains deterministic floating-point behavior across arbitrarily long contexts while keeping GPU-resident graphs simple and reusable.
Frequently Asked Questions
What is the default prefill chunk size in ds4's Metal backend?
The default prefill chunk size is 4096 tokens. This value is hardcoded as the canonical setting for the benchmark suite and represents the standard granularity for processing prompts with the Metal graph backend according to the source documentation.
How does changing the chunk size affect KV-cache checkpointing?
Modifying the chunk size with --prefill-chunk changes where the engine places KV-cache checkpoints during prefill. Smaller chunks create more frequent checkpoints, increasing storage granularity but reducing the memory required for any single graph execution. Larger chunks decrease checkpoint frequency but require proportionally more GPU memory per iteration.
Why does ds4 reuse the same Metal graph for every chunk instead of building new ones?
The system constructs a single range-capable "layer-major" graph to eliminate compilation overhead and ensure deterministic execution. Reusing this graph across all chunks prevents floating-point non-determinism that could arise from dynamic graph construction while minimizing host-GPU synchronization points throughout the prefill phase.
Where is the prefill chunking logic implemented in the source code?
The chunking loop and KV checkpointing reside in ds4.c within functions like ds4_session_sync, while the reusable Metal graph execution and range-capable graph construction are implemented in ds4_metal.m. Command-line parsing for the --prefill-chunk flag occurs in ds4_cli.c.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →