How ds4-server Handles Batched Sessions for Multi-User Inference
ds4-server processes concurrent client requests through a three-stage pipeline that aggregates individual token evaluations into GPU batches, using separate mutexes for request queuing and inference execution.
The ds4-server binary in the antirez/ds4 repository implements batched mode to maximize GPU/Metal utilization when serving multiple simultaneous users. Rather than executing each inference request immediately, the server collects pending work into batches that share a single kernel dispatch. This article explains the complete mechanism—from request arrival through batch execution—based on the source code in ds4_server.c and ds4.c.
The Three Stages of Batched Session Handling
Stage 1: Request Enqueue in server_eval_token
When a client thread requests token evaluation, it enters server_eval_token at lines 10801-10809 of ds4_server.c. The path depends on the server configuration:
- Batched mode disabled: The function acquires
inference_mu, callsds4_session_evalsynchronously, and returns immediately. - Batched mode enabled: The request is recorded in its slot (
decode_token,decode_pending = true), andpthread_cond_broadcast(&s->model_cv)wakes the decode worker.
/* Client thread – enqueue a token for batched evaluation */
int rc = server_eval_token(srv, slot, token, err_buf, sizeof(err_buf));
if (rc != 0) { /* handle error */ }
The slot-based design decouples client threads from inference execution. Each slot maintains state flags that the decode worker inspects without blocking incoming connections.
Stage 2: Batch Construction in decode_worker_main
The dedicated decode worker thread runs decode_worker_main (lines 10855-10931 of ds4_server.c). Its operation follows a precise sequence:
- Wait for work: Blocks on
pthread_cond_timedwait(&s->model_cv, &s->model_mu)untils->decode_pending > 0 - Optional coalescing: Delays up to
DS4_SERVER_DECODE_COALESCE_USmicroseconds to accumulate more requests - Build batch array: Populates
ds4_decode_item items[count]from pending slots (lines 10890-10899)
/* Inside decode_worker_main – build the batch */
for (int i = 0; i < s->slot_count; i++) {
server_slot *slot = &s->slots[i];
if (!slot->decode_pending) continue;
slot->decode_pending = false;
slot->decode_in_flight = true;
s->decode_pending--;
members[count] = slot;
items[count].session = slot->session;
items[count].token = slot->decode_token;
count++;
}
The coalescing timeout is critical for throughput: it trades individual request latency for batch efficiency when request arrival is bursty.
Stage 3: Batch Execution in ds4_sessions_eval_batch
The core library function ds4_sessions_eval_batch (lines 61333-61372 of ds4.c) validates and dispatches the batch:
| Validation check | Purpose |
|---|---|
Same ds4_engine for all slots |
Ensures compatible compute context |
| Tokens within vocabulary bounds | Prevents out-of-bounds embedding lookups |
| No duplicate sessions | Avoids race conditions in KV cache |
| Context window not exceeded | Maintains attention mechanism correctness |
/* ds4_sessions_eval_batch – validation & dispatch */
if (count == 1) return ds4_session_eval(items[0].session, items[0].token, err, errlen);
...
if (e->backend == DS4_BACKEND_CUDA)
return ds4_sessions_eval_batch_cuda(items, count, err, errlen);
...
For single-item batches, the function falls back to ds4_session_eval to avoid batch overhead. Multi-item batches route to ds4_sessions_eval_batch_cuda or ds4_sessions_eval_batch_metal depending on the engine's backend type.
Synchronization Architecture
The ds4-server uses two distinct mutexes to minimize contention:
model_mu: Protects slot state flags, pending counters, and condition variable signalinginference_mu: Guards the engine state during actual GPU/Metal execution
This separation allows the decode worker to assemble batches without blocking new request arrivals. Once assembled, the batch executes under inference_mu to ensure exclusive access to the compute context.
| Phase | Lock held | Operation |
|---|---|---|
| Request arrival | model_mu |
Mark slot pending, broadcast model_cv |
| Worker wake-up | model_mu |
Collect slots, optionally coalesce |
| Batch execution | inference_mu |
Call ds4_sessions_eval_batch |
| Result writeback | model_mu |
Store per-slot results, signal clients |
Error Handling and Atomicity
ds4_sessions_eval_batch enforces all-or-nothing semantics: if the backend returns an error, the function invalidates every session in the batch. This prevents partial progress where some slots advance their KV cache while others fail.
After backend completion, decode_worker_main copies results to each slot:
decode_rc– return codedecode_err– error string (if any)decode_done = true– completion flagpthread_cond_broadcast(&s->model_cv)– wake waiting client threads
Clients observe the same interface regardless of batching: they block in server_eval_token until their slot's decode_done flag is set.
Key Source Files
| File | Lines | Responsibility |
|---|---|---|
ds4_server.c |
10801-10931 | server_eval_token, decode_worker_main, slot management |
ds4.c |
61333-61372 | ds4_sessions_eval_batch, validation, backend dispatch |
ds4_session.c |
— | Single-session ds4_session_eval for non-batched fallback |
ds4_tp.c |
— | Thread-pool utilities for slot worker management |
Summary
- ds4-server uses batched mode to aggregate concurrent inference requests into single GPU dispatches
- The three-stage pipeline (enqueue → build → execute) decouples client threads from compute scheduling
- Two-mutex design (
model_mu+inference_mu) maximizes concurrency while protecting engine state - Coalescing timeout (
DS4_SERVER_DECODE_COALESCE_US) trades latency for throughput under variable load - All-or-nothing error handling ensures consistent session state across batch failures
Frequently Asked Questions
How does ds4-server decide between batched and non-batched execution?
The server checks a batched_mode flag at the entry of server_eval_token. When disabled, each request acquires inference_mu immediately and calls ds4_session_eval synchronously. When enabled, requests are deferred to the decode worker thread for batch assembly. This flag is typically set at server startup based on configuration.
What happens if only one request is pending during coalescing?
The batch builder proceeds with a single-item batch. ds4_sessions_eval_batch detects count == 1 and falls back to ds4_session_eval automatically, avoiding unnecessary batch overhead. The coalescing timeout does not delay single requests indefinitely—it triggers as soon as the wait expires or additional requests arrive.
Why does ds4-server use two separate mutexes?
model_mu protects mutable slot state and condition variables accessed frequently by both client threads and the decode worker. inference_mu serializes access to the GPU/Metal context, which has much higher contention cost. Separating them allows request enqueue to proceed in parallel with batch preparation, only serializing during the actual kernel dispatch.
Can batched sessions interfere with each other's results?
No. While sessions share a GPU batch dispatch, each maintains independent KV cache and output buffers. ds4_sessions_eval_batch validates that no session appears twice in the same batch. Per-slot results are written back to separate memory locations before clients are signaled, ensuring isolation despite shared compute.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →