# Metal Graph Execution Modes in ds4: Decode vs Batch vs No-Copy

> Understand ds4 Metal graph execution modes: Decode, Batch, and No-Copy. Learn how GPU tiers and hidden-state tensors impact performance in this inference engine.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-08

---

**The ds4 inference engine supports three distinct Metal graph execution modes—Decode (single-token), Batch (prefill), and No-Copy (raw)—that differ in how GPU tiers are selected and whether hidden-state tensors are copied between tiers.**

The **Metal graph execution modes** in antirez/ds4 determine how the inference engine drives whole-model Metal graphs on macOS and Apple Silicon. These modes control memory traffic and device switching behavior by choosing which hidden-state buffers are transferred when the active GPU tier changes. Understanding these modes is essential for optimizing throughput during both interactive generation and bulk prefill operations.

## The Three Metal Graph Execution Modes

The core logic resides in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), where three specialized functions handle tier transitions. Each function corresponds to a specific execution strategy tailored to different generation phases.

### Decode Mode (Single-Token)

**Decode mode** is optimized for normal token-by-token generation. When the engine processes one token at a time, it calls `metal_graph_set_active_tier_decode()` to manage device switches.

This function performs three critical operations:
- Switches the current GPU device if necessary
- Copies the **current hidden-state tensor** (`cur_hc`) from the previously active tier to the new tier
- Updates the `g->active_tier` field to reflect the new active device

The hidden state transferred is minimal—approximately `HC × EMBD` floats—keeping memory traffic low during interactive generation. According to the source code at lines **15452–15457**, this mode is explicitly designed for "decode (one token at a time)" scenarios where latency matters more than throughput.

### Batch Mode (Prefill)

**Batch mode** accelerates prefill and chunked generation where multiple tokens are processed simultaneously. Instead of copying the current single-token state, the engine calls `metal_graph_set_active_tier_batch()`.

This function handles larger tensor transfers:
- Switches the GPU device when needed
- Copies the **batch-prefill hidden-state** (`batch_cur_hc`) sized `chunk_tokens × HC × EMBD` from the old tier to the new tier
- Updates `g->active_tier`

As noted at lines **15455–15461** in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), this mode supports "chunked prefill" operations. By transferring the entire chunk's hidden state at once, the engine avoids per-token device switches during prefill, dramatically improving throughput when processing long contexts.

### No-Copy Mode (Raw)

**No-Copy mode** eliminates tensor transfers entirely. When the graph is allocated with raw capacity or when buffers are being prepared without immediate computation, the engine invokes `metal_graph_set_active_tier_no_copy()`.

This function (implemented at lines **15818–15827**) only switches the device and updates `g->active_tier`—**no tensor copy** is performed. This mode is useful when allocating raw caches or when the caller ensures data already resides on the target tier, minimizing overhead during graph preparation phases.

## How Execution Modes Are Selected

The ds4 engine selects execution modes through both command-line interfaces and environment variables.

### CLI Flags

In [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) (lines **87–100**), three boolean flags trigger specific test harnesses that exercise the different modes:

- `--metal-graph-test` → Invokes `ds4_engine_metal_graph_test()` for decode-mode validation
- `--metal-graph-full-test` → Invokes `ds4_engine_metal_graph_full_test()` for full-graph decode testing
- `--metal-graph-prompt-test` → Invokes `ds4_engine_metal_graph_prompt_test()` for prompt-logits verification

These entry points ultimately call internal helpers such as `metal_graph_decode_test()`, `metal_graph_first_token_full_test()`, and `metal_graph_prompt_logits_test()` defined in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).

### Environment Variables

Internal allocation logic responds to variables like `DS4_METAL_PREFILL_CHUNK` and `DS4_METAL_GRAPH_RAW_CAP`. While these influence whether the graph runs in batch-prefill or raw-allocation configurations, the actual execution mode (decode vs. batch vs. no-copy) is determined by the specific tier-switching functions called at runtime.

## Practical Code Examples

To run validation tests for specific execution modes, use the public API functions:

```c
/* Run a decode-only sanity check (single-token mode) */
int rc = ds4_engine_metal_graph_test(engine, &prompt);

/* Run a full-graph test walking through the whole model in decode mode */
int rc = ds4_engine_metal_graph_full_test(engine, &prompt);

/* Run a prompt-logits test for verifying logits on a given context */
int rc = ds4_engine_metal_graph_prompt_test(engine, &prompt, ctx_size);

```

For programmatic control over tier switching, call the low-level helpers directly:

```c
/* Switch to tier 2 for decode (single-token) - copies cur_hc */
metal_graph_set_active_tier_decode(g, 2);

/* Switch to tier 1 for batch pre-fill with 2048-token chunk - copies batch_cur_hc */
metal_graph_set_active_tier_batch(g, 1, 2048);

/* Switch to tier 3 without any hidden-state copy */
metal_graph_set_active_tier_no_copy(g, 3);

```

These functions are implemented in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at lines **15470–15516** and **15818–15827**, respectively.

## Summary

- **Decode mode** minimizes latency for single-token generation by copying only the current hidden state (`cur_hc`) between GPU tiers.
- **Batch mode** maximizes prefill throughput by copying entire chunks of hidden state (`batch_cur_hc`) sized by `chunk_tokens × HC × EMBD`, avoiding per-token synchronization.
- **No-Copy mode** eliminates tensor transfers entirely, useful for raw buffer allocation and scenarios where data residency is already guaranteed.
- The modes are selected via CLI flags (`--metal-graph-test`, `--metal-graph-full-test`, `--metal-graph-prompt-test`) in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) or programmatically through the `metal_graph_set_active_tier_*` family of functions.

## Frequently Asked Questions

### What is the difference between decode and batch mode in ds4 Metal graphs?

Decode mode transfers only the current hidden state (`cur_hc`) sized `HC × EMBD`, optimized for single-token generation latency. Batch mode transfers the entire prefill buffer (`batch_cur_hc`) sized `chunk_tokens × HC × EMBD`, optimized for processing multiple tokens simultaneously during prefill. Both functions update `g->active_tier` but differ in memory traffic and synchronization patterns.

### When should I use no-copy mode?

Use no-copy mode when allocating raw graph capacity or when preparing buffers without immediate computation. Since `metal_graph_set_active_tier_no_copy()` performs no tensor copies, it avoids unnecessary memory traffic when the hidden state either doesn't exist yet or already resides on the target GPU tier.

### How do I test specific Metal graph execution modes?

Pass the corresponding CLI flag when running ds4: `--metal-graph-test` for decode mode, `--metal-graph-full-test` for full-graph decode validation, or `--metal-graph-prompt-test` for prompt-logits verification. These flags are parsed in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c) and invoke the respective test harnesses in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).

### Where is the tier-switching logic implemented?

The core tier-switching logic is implemented in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c). The functions `metal_graph_set_active_tier_decode()` (lines 15452–15457), `metal_graph_set_active_tier_batch()` (lines 15455–15461), and `metal_graph_set_active_tier_no_copy()` (lines 15818–15827) handle device switching and optional tensor copying between GPU tiers.