Running Batched Inference with Multiple Concurrent Sessions in DS4: A Deep Dive
DS4 supports session-level batching through ds4_sessions_eval_batch, enabling multiple independent generation sessions to be evaluated together in a single GPU/Metal command epoch, reducing kernel launch overhead and maximizing throughput.
DS4 (the "deep-seek-4" inference engine) implements a sophisticated session-level batching system that allows many independent generation sessions to share a single GPU command buffer. This architecture significantly reduces per-token latency on both Metal and CUDA backends by amortizing kernel launch costs across concurrent requests.
Core Batching Architecture in ds4.c
The batching implementation lives primarily in ds4.c, where the engine aggregates decode requests from multiple sessions into unified GPU work units.
The Entry Point: ds4_sessions_eval_batch
The generic entry point ds4_sessions_eval_batch serves as the primary interface used by the CLI (ds4_cli.c), server (ds4_server.c), and test harnesses. This function accepts an array of ds4_decode_item structures, each containing a session pointer and the token to process.
In ds4.c (lines 61387-61441), the function validates inputs and dispatches to backend-specific implementations based on the engine's backend field. The signature follows this pattern:
int ds4_sessions_eval_batch(ds4_decode_item *items, int count, char *err, size_t errlen);
Backend Dispatch Logic
The dispatch logic branches immediately based on the active backend:
- Metal: Routes to
ds4_sessions_eval_batch_metal - CUDA: Routes to
ds4_sessions_eval_batch_cuda
This ensures that the high-level batching API remains backend-agnostic while allowing each platform to optimize command encoding for its specific hardware characteristics.
Metal Backend Implementation
The Metal implementation demonstrates DS4's approach to maximizing GPU utilization through two distinct execution paths.
Native Shared Fast Path
When the hardware supports metal_graph_native_session_batch_shared, DS4 encodes all session tokens in a single graph call. This path, found at lines 61045-61052 in ds4.c, eliminates per-session encoding overhead by sharing graph buffers across the batch.
The function metal_graph_encode_native_session_batch_shared processes the entire batch as one coherent work unit, significantly reducing CPU-to-GPU communication latency.
Per-Session Fallback Loop
If native shared batching is unavailable, the engine falls back to a loop that calls metal_graph_encode_token_raw_swa for each session individually (lines 61053-61075). While less efficient, this path still benefits from batching because all work is enclosed within a single GPU command batch bracketed by ds4_gpu_begin_commands and ds4_gpu_end_commands.
After encoding completes, logits are read back in one consolidated operation via ds4_gpu_tensor_read (lines 61087-61095), ensuring that memory transfers are coalesced across all sessions.
Tensor-Parallel (TP) Support
For distributed deployments, ds4_sessions_eval_batch_metal checks for tensor-parallel mode and packages batch items into a TP-wire structure (ds4_sessions_tp_batch_items) before transmission to worker nodes. The worker processes the batch remotely and returns decoded logits through ds4_tp_wait_command_ack and ds4_sessions_tp_recv_logits, as implemented in ds4_tp.c.
Prefill-Aware Batching
DS4 extends basic batching with ds4_sessions_eval_batch_with_prefill_* functions, which interleave resumed prefilling chunks with independent decode rows. This advanced pattern, located around lines 63964-63982 in ds4.c, first validates strict constraints:
- No distributed/tensor-parallel mode active
- No GLM sessions
- No SSD streaming
- Prefilling region must fit within graph capacity
Once validated, the Metal path processes the prefilling chunk and decode rows within a single epoch, preserving each session's private KV cache while maximizing throughput for heterogeneous workloads.
CUDA Backend Parity
The CUDA backend mirrors the Metal architecture through ds4_sessions_eval_batch_cuda, providing identical semantics for NVIDIA hardware. Test files tests/test_cuda_session_batch.c and tests/test_cuda_mixed_batch.c demonstrate that the batching API remains consistent across platforms, allowing applications to switch backends without modifying batching logic.
Practical Implementation Example
The following minimal example illustrates how to run batched inference on the Metal backend. The same pattern applies to CUDA by substituting DS4_BACKEND_METAL with DS4_BACKEND_CUDA.
/* 1. Create an engine (already loaded with a model) */
ds4_engine *engine = NULL;
ds4_engine_create(&engine, "model.gguf", DS4_BACKEND_METAL, /* support_kind */ DS4_SUPPORT_NONE);
/* 2. Allocate a few sessions */
#define N_SESSIONS 4
ds4_session *sessions[N_SESSIONS];
for (int i = 0; i < N_SESSIONS; i++) {
if (ds4_session_create(&sessions[i], engine, /* ctx */ NULL) != 0) {
fprintf(stderr, "Failed to create session %d\n", i);
exit(1);
}
}
/* 3. Prime each session with an initial token (e.g., BOS) */
for (int i = 0; i < N_SESSIONS; i++) {
ds4_session_feed_token(sessions[i], BOS_TOKEN);
}
/* 4. Build the decode item list */
ds4_decode_item items[N_SESSIONS];
for (int i = 0; i < N_SESSIONS; i++) {
items[i].session = sessions[i];
items[i].token = ds4_session_argmax(sessions[i]); // token to decode next
}
/* 5. Run the batched evaluation */
char err[256];
int rc = ds4_sessions_eval_batch(items, N_SESSIONS, err, sizeof(err));
if (rc != 0) {
fprintf(stderr, "Batch failed: %s\n", err);
exit(1);
}
/* 6. Retrieve the newly generated token from each session */
for (int i = 0; i < N_SESSIONS; i++) {
uint32_t new_token = ds4_session_argmax(sessions[i]);
printf("Session %d generated token %u\n", i, new_token);
}
/* 7. Clean up */
for (int i = 0; i < N_SESSIONS; i++) ds4_session_free(sessions[i]);
ds4_engine_free(engine);
This workflow creates independent sessions, aggregates them into ds4_decode_item structures, and executes ds4_sessions_eval_batch to process all tokens simultaneously. The repository's tests/test_metal_session_batch.c provides additional examples including error injection and mixed-prefill scenarios.
Summary
ds4_sessions_eval_batchserves as the universal entry point for batched inference, dispatching to Metal or CUDA implementations based on the engine configuration.- The native shared fast path (
metal_graph_encode_native_session_batch_shared) eliminates per-session overhead when hardware permits, while a fallback loop ensures compatibility across devices. - Tensor-parallel support enables distributed batching through
ds4_tp.c, allowing batches to span multiple workers. - Prefill-aware batching allows mixing prompt continuation with standard decoding, subject to strict validation constraints regarding distributed mode and session types.
- Logits are retrieved via
ds4_gpu_tensor_readafter command buffer finalization, ensuring coalesced memory transfers across the entire batch.
Frequently Asked Questions
What is the main entry point for batched inference in DS4?
The primary function is ds4_sessions_eval_batch, defined in ds4.c at lines 61387-61441. This function accepts an array of ds4_decode_item structures and dispatches to backend-specific implementations (ds4_sessions_eval_batch_metal or ds4_sessions_eval_batch_cuda) based on the engine's configured backend.
How does DS4 handle different hardware backends when batching sessions?
DS4 abstracts hardware differences through a unified dispatch layer. The generic ds4_sessions_eval_batch function checks engine->backend and routes to Metal-specific or CUDA-specific implementations. Both backends follow the same architectural pattern: aggregate work into a single command buffer, encode tokens (either through shared native batching or per-session loops), and read back logits coalesced across all sessions.
Can prefilling and decoding be mixed in the same batch?
Yes, through the ds4_sessions_eval_batch_with_prefill_* functions. However, this requires the prefilling session to meet strict constraints: no tensor-parallel mode, no GLM sessions, no SSD streaming, and the prefilling region must fit within the graph's capacity. When validated, the engine processes the prefilling chunk alongside independent decode rows within a single GPU epoch.
What happens when native shared batching is not supported?
The engine falls back to a per-session encoding loop within ds4_sessions_eval_batch_metal. This loop calls metal_graph_encode_token_raw_swa for each session individually (lines 61053-61075 in ds4.c). While this incurs higher encoding overhead than the native shared path, it still benefits from batching because all GPU commands are submitted within a single ds4_gpu_begin_commands/ds4_gpu_end_commands bracket, preserving the kernel launch reduction benefits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →