# Running Batched Inference with Multiple Concurrent Sessions in DS4: A Deep Dive

> Discover how to run batched inference with multiple concurrent sessions in DS4 using ds4_sessions_eval_batch. Maximize GPU throughput and reduce overhead.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-08

---

**DS4 supports session-level batching through `ds4_sessions_eval_batch`, enabling multiple independent generation sessions to be evaluated together in a single GPU/Metal command epoch, reducing kernel launch overhead and maximizing throughput.**

DS4 (the "deep-seek-4" inference engine) implements a sophisticated **session-level batching** system that allows many independent generation sessions to share a single GPU command buffer. This architecture significantly reduces per-token latency on both Metal and CUDA backends by amortizing kernel launch costs across concurrent requests.

## Core Batching Architecture in ds4.c

The batching implementation lives primarily in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), where the engine aggregates decode requests from multiple sessions into unified GPU work units.

### The Entry Point: ds4_sessions_eval_batch

The generic entry point `ds4_sessions_eval_batch` serves as the primary interface used by the CLI ([`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c)), server ([`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)), and test harnesses. This function accepts an array of `ds4_decode_item` structures, each containing a session pointer and the token to process.

In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) (lines 61387-61441), the function validates inputs and dispatches to backend-specific implementations based on the engine's `backend` field. The signature follows this pattern:

```c
int ds4_sessions_eval_batch(ds4_decode_item *items, int count, char *err, size_t errlen);

```

### Backend Dispatch Logic

The dispatch logic branches immediately based on the active backend:

- **Metal**: Routes to `ds4_sessions_eval_batch_metal`
- **CUDA**: Routes to `ds4_sessions_eval_batch_cuda`

This ensures that the high-level batching API remains backend-agnostic while allowing each platform to optimize command encoding for its specific hardware characteristics.

## Metal Backend Implementation

The Metal implementation demonstrates DS4's approach to maximizing GPU utilization through two distinct execution paths.

### Native Shared Fast Path

When the hardware supports `metal_graph_native_session_batch_shared`, DS4 encodes all session tokens in a single graph call. This path, found at lines 61045-61052 in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), eliminates per-session encoding overhead by sharing graph buffers across the batch.

The function `metal_graph_encode_native_session_batch_shared` processes the entire batch as one coherent work unit, significantly reducing CPU-to-GPU communication latency.

### Per-Session Fallback Loop

If native shared batching is unavailable, the engine falls back to a loop that calls `metal_graph_encode_token_raw_swa` for each session individually (lines 61053-61075). While less efficient, this path still benefits from batching because all work is enclosed within a single GPU command batch bracketed by `ds4_gpu_begin_commands` and `ds4_gpu_end_commands`.

After encoding completes, logits are read back in one consolidated operation via `ds4_gpu_tensor_read` (lines 61087-61095), ensuring that memory transfers are coalesced across all sessions.

### Tensor-Parallel (TP) Support

For distributed deployments, `ds4_sessions_eval_batch_metal` checks for tensor-parallel mode and packages batch items into a TP-wire structure (`ds4_sessions_tp_batch_items`) before transmission to worker nodes. The worker processes the batch remotely and returns decoded logits through `ds4_tp_wait_command_ack` and `ds4_sessions_tp_recv_logits`, as implemented in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c).

## Prefill-Aware Batching

DS4 extends basic batching with `ds4_sessions_eval_batch_with_prefill_*` functions, which interleave resumed prefilling chunks with independent decode rows. This advanced pattern, located around lines 63964-63982 in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), first validates strict constraints:

- No distributed/tensor-parallel mode active
- No GLM sessions
- No SSD streaming
- Prefilling region must fit within graph capacity

Once validated, the Metal path processes the prefilling chunk and decode rows within a single epoch, preserving each session's private KV cache while maximizing throughput for heterogeneous workloads.

## CUDA Backend Parity

The CUDA backend mirrors the Metal architecture through `ds4_sessions_eval_batch_cuda`, providing identical semantics for NVIDIA hardware. Test files [`tests/test_cuda_session_batch.c`](https://github.com/antirez/ds4/blob/main/tests/test_cuda_session_batch.c) and [`tests/test_cuda_mixed_batch.c`](https://github.com/antirez/ds4/blob/main/tests/test_cuda_mixed_batch.c) demonstrate that the batching API remains consistent across platforms, allowing applications to switch backends without modifying batching logic.

## Practical Implementation Example

The following minimal example illustrates how to run batched inference on the Metal backend. The same pattern applies to CUDA by substituting `DS4_BACKEND_METAL` with `DS4_BACKEND_CUDA`.

```c
/* 1. Create an engine (already loaded with a model) */
ds4_engine *engine = NULL;
ds4_engine_create(&engine, "model.gguf", DS4_BACKEND_METAL, /* support_kind */ DS4_SUPPORT_NONE);

/* 2. Allocate a few sessions */
#define N_SESSIONS 4
ds4_session *sessions[N_SESSIONS];
for (int i = 0; i < N_SESSIONS; i++) {
    if (ds4_session_create(&sessions[i], engine, /* ctx */ NULL) != 0) {
        fprintf(stderr, "Failed to create session %d\n", i);
        exit(1);
    }
}

/* 3. Prime each session with an initial token (e.g., BOS) */
for (int i = 0; i < N_SESSIONS; i++) {
    ds4_session_feed_token(sessions[i], BOS_TOKEN);
}

/* 4. Build the decode item list */
ds4_decode_item items[N_SESSIONS];
for (int i = 0; i < N_SESSIONS; i++) {
    items[i].session = sessions[i];
    items[i].token   = ds4_session_argmax(sessions[i]);   // token to decode next
}

/* 5. Run the batched evaluation */
char err[256];
int rc = ds4_sessions_eval_batch(items, N_SESSIONS, err, sizeof(err));
if (rc != 0) {
    fprintf(stderr, "Batch failed: %s\n", err);
    exit(1);
}

/* 6. Retrieve the newly generated token from each session */
for (int i = 0; i < N_SESSIONS; i++) {
    uint32_t new_token = ds4_session_argmax(sessions[i]);
    printf("Session %d generated token %u\n", i, new_token);
}

/* 7. Clean up */
for (int i = 0; i < N_SESSIONS; i++) ds4_session_free(sessions[i]);
ds4_engine_free(engine);

```

This workflow creates independent sessions, aggregates them into `ds4_decode_item` structures, and executes `ds4_sessions_eval_batch` to process all tokens simultaneously. The repository's [`tests/test_metal_session_batch.c`](https://github.com/antirez/ds4/blob/main/tests/test_metal_session_batch.c) provides additional examples including error injection and mixed-prefill scenarios.

## Summary

- **`ds4_sessions_eval_batch`** serves as the universal entry point for batched inference, dispatching to Metal or CUDA implementations based on the engine configuration.
- The **native shared fast path** (`metal_graph_encode_native_session_batch_shared`) eliminates per-session overhead when hardware permits, while a fallback loop ensures compatibility across devices.
- **Tensor-parallel support** enables distributed batching through [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c), allowing batches to span multiple workers.
- **Prefill-aware batching** allows mixing prompt continuation with standard decoding, subject to strict validation constraints regarding distributed mode and session types.
- Logits are retrieved via `ds4_gpu_tensor_read` after command buffer finalization, ensuring coalesced memory transfers across the entire batch.

## Frequently Asked Questions

### What is the main entry point for batched inference in DS4?

The primary function is `ds4_sessions_eval_batch`, defined in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) at lines 61387-61441. This function accepts an array of `ds4_decode_item` structures and dispatches to backend-specific implementations (`ds4_sessions_eval_batch_metal` or `ds4_sessions_eval_batch_cuda`) based on the engine's configured backend.

### How does DS4 handle different hardware backends when batching sessions?

DS4 abstracts hardware differences through a unified dispatch layer. The generic `ds4_sessions_eval_batch` function checks `engine->backend` and routes to Metal-specific or CUDA-specific implementations. Both backends follow the same architectural pattern: aggregate work into a single command buffer, encode tokens (either through shared native batching or per-session loops), and read back logits coalesced across all sessions.

### Can prefilling and decoding be mixed in the same batch?

Yes, through the `ds4_sessions_eval_batch_with_prefill_*` functions. However, this requires the prefilling session to meet strict constraints: no tensor-parallel mode, no GLM sessions, no SSD streaming, and the prefilling region must fit within the graph's capacity. When validated, the engine processes the prefilling chunk alongside independent decode rows within a single GPU epoch.

### What happens when native shared batching is not supported?

The engine falls back to a per-session encoding loop within `ds4_sessions_eval_batch_metal`. This loop calls `metal_graph_encode_token_raw_swa` for each session individually (lines 61053-61075 in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)). While this incurs higher encoding overhead than the native shared path, it still benefits from batching because all GPU commands are submitted within a single `ds4_gpu_begin_commands`/`ds4_gpu_end_commands` bracket, preserving the kernel launch reduction benefits.