Metal Graph Execution Modes in ds4: Decode vs Batch vs No-Copy

The ds4 inference engine supports three distinct Metal graph execution modes—Decode (single-token), Batch (prefill), and No-Copy (raw)—that differ in how GPU tiers are selected and whether hidden-state tensors are copied between tiers.

The Metal graph execution modes in antirez/ds4 determine how the inference engine drives whole-model Metal graphs on macOS and Apple Silicon. These modes control memory traffic and device switching behavior by choosing which hidden-state buffers are transferred when the active GPU tier changes. Understanding these modes is essential for optimizing throughput during both interactive generation and bulk prefill operations.

The Three Metal Graph Execution Modes

The core logic resides in ds4.c, where three specialized functions handle tier transitions. Each function corresponds to a specific execution strategy tailored to different generation phases.

Decode Mode (Single-Token)

Decode mode is optimized for normal token-by-token generation. When the engine processes one token at a time, it calls metal_graph_set_active_tier_decode() to manage device switches.

This function performs three critical operations:

  • Switches the current GPU device if necessary
  • Copies the current hidden-state tensor (cur_hc) from the previously active tier to the new tier
  • Updates the g->active_tier field to reflect the new active device

The hidden state transferred is minimal—approximately HC × EMBD floats—keeping memory traffic low during interactive generation. According to the source code at lines 15452–15457, this mode is explicitly designed for "decode (one token at a time)" scenarios where latency matters more than throughput.

Batch Mode (Prefill)

Batch mode accelerates prefill and chunked generation where multiple tokens are processed simultaneously. Instead of copying the current single-token state, the engine calls metal_graph_set_active_tier_batch().

This function handles larger tensor transfers:

  • Switches the GPU device when needed
  • Copies the batch-prefill hidden-state (batch_cur_hc) sized chunk_tokens × HC × EMBD from the old tier to the new tier
  • Updates g->active_tier

As noted at lines 15455–15461 in ds4.c, this mode supports "chunked prefill" operations. By transferring the entire chunk's hidden state at once, the engine avoids per-token device switches during prefill, dramatically improving throughput when processing long contexts.

No-Copy Mode (Raw)

No-Copy mode eliminates tensor transfers entirely. When the graph is allocated with raw capacity or when buffers are being prepared without immediate computation, the engine invokes metal_graph_set_active_tier_no_copy().

This function (implemented at lines 15818–15827) only switches the device and updates g->active_tier—no tensor copy is performed. This mode is useful when allocating raw caches or when the caller ensures data already resides on the target tier, minimizing overhead during graph preparation phases.

How Execution Modes Are Selected

The ds4 engine selects execution modes through both command-line interfaces and environment variables.

CLI Flags

In ds4_cli.c (lines 87–100), three boolean flags trigger specific test harnesses that exercise the different modes:

  • --metal-graph-test → Invokes ds4_engine_metal_graph_test() for decode-mode validation
  • --metal-graph-full-test → Invokes ds4_engine_metal_graph_full_test() for full-graph decode testing
  • --metal-graph-prompt-test → Invokes ds4_engine_metal_graph_prompt_test() for prompt-logits verification

These entry points ultimately call internal helpers such as metal_graph_decode_test(), metal_graph_first_token_full_test(), and metal_graph_prompt_logits_test() defined in ds4.c.

Environment Variables

Internal allocation logic responds to variables like DS4_METAL_PREFILL_CHUNK and DS4_METAL_GRAPH_RAW_CAP. While these influence whether the graph runs in batch-prefill or raw-allocation configurations, the actual execution mode (decode vs. batch vs. no-copy) is determined by the specific tier-switching functions called at runtime.

Practical Code Examples

To run validation tests for specific execution modes, use the public API functions:

/* Run a decode-only sanity check (single-token mode) */
int rc = ds4_engine_metal_graph_test(engine, &prompt);

/* Run a full-graph test walking through the whole model in decode mode */
int rc = ds4_engine_metal_graph_full_test(engine, &prompt);

/* Run a prompt-logits test for verifying logits on a given context */
int rc = ds4_engine_metal_graph_prompt_test(engine, &prompt, ctx_size);

For programmatic control over tier switching, call the low-level helpers directly:

/* Switch to tier 2 for decode (single-token) - copies cur_hc */
metal_graph_set_active_tier_decode(g, 2);

/* Switch to tier 1 for batch pre-fill with 2048-token chunk - copies batch_cur_hc */
metal_graph_set_active_tier_batch(g, 1, 2048);

/* Switch to tier 3 without any hidden-state copy */
metal_graph_set_active_tier_no_copy(g, 3);

These functions are implemented in ds4.c at lines 15470–15516 and 15818–15827, respectively.

Summary

  • Decode mode minimizes latency for single-token generation by copying only the current hidden state (cur_hc) between GPU tiers.
  • Batch mode maximizes prefill throughput by copying entire chunks of hidden state (batch_cur_hc) sized by chunk_tokens × HC × EMBD, avoiding per-token synchronization.
  • No-Copy mode eliminates tensor transfers entirely, useful for raw buffer allocation and scenarios where data residency is already guaranteed.
  • The modes are selected via CLI flags (--metal-graph-test, --metal-graph-full-test, --metal-graph-prompt-test) in ds4_cli.c or programmatically through the metal_graph_set_active_tier_* family of functions.

Frequently Asked Questions

What is the difference between decode and batch mode in ds4 Metal graphs?

Decode mode transfers only the current hidden state (cur_hc) sized HC × EMBD, optimized for single-token generation latency. Batch mode transfers the entire prefill buffer (batch_cur_hc) sized chunk_tokens × HC × EMBD, optimized for processing multiple tokens simultaneously during prefill. Both functions update g->active_tier but differ in memory traffic and synchronization patterns.

When should I use no-copy mode?

Use no-copy mode when allocating raw graph capacity or when preparing buffers without immediate computation. Since metal_graph_set_active_tier_no_copy() performs no tensor copies, it avoids unnecessary memory traffic when the hidden state either doesn't exist yet or already resides on the target GPU tier.

How do I test specific Metal graph execution modes?

Pass the corresponding CLI flag when running ds4: --metal-graph-test for decode mode, --metal-graph-full-test for full-graph decode validation, or --metal-graph-prompt-test for prompt-logits verification. These flags are parsed in ds4_cli.c and invoke the respective test harnesses in ds4.c.

Where is the tier-switching logic implemented?

The core tier-switching logic is implemented in ds4.c. The functions metal_graph_set_active_tier_decode() (lines 15452–15457), metal_graph_set_active_tier_batch() (lines 15455–15461), and metal_graph_set_active_tier_no_copy() (lines 15818–15827) handle device switching and optional tensor copying between GPU tiers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →