How ds4-server Handles Batched Sessions for Multi-User Inference

ds4-server processes concurrent client requests through a three-stage pipeline that aggregates individual token evaluations into GPU batches, using separate mutexes for request queuing and inference execution.

The ds4-server binary in the antirez/ds4 repository implements batched mode to maximize GPU/Metal utilization when serving multiple simultaneous users. Rather than executing each inference request immediately, the server collects pending work into batches that share a single kernel dispatch. This article explains the complete mechanism—from request arrival through batch execution—based on the source code in ds4_server.c and ds4.c.

The Three Stages of Batched Session Handling

Stage 1: Request Enqueue in server_eval_token

When a client thread requests token evaluation, it enters server_eval_token at lines 10801-10809 of ds4_server.c. The path depends on the server configuration:

  • Batched mode disabled: The function acquires inference_mu, calls ds4_session_eval synchronously, and returns immediately.
  • Batched mode enabled: The request is recorded in its slot (decode_token, decode_pending = true), and pthread_cond_broadcast(&s->model_cv) wakes the decode worker.
/* Client thread – enqueue a token for batched evaluation */
int rc = server_eval_token(srv, slot, token, err_buf, sizeof(err_buf));
if (rc != 0) { /* handle error */ }

The slot-based design decouples client threads from inference execution. Each slot maintains state flags that the decode worker inspects without blocking incoming connections.

Stage 2: Batch Construction in decode_worker_main

The dedicated decode worker thread runs decode_worker_main (lines 10855-10931 of ds4_server.c). Its operation follows a precise sequence:

  1. Wait for work: Blocks on pthread_cond_timedwait(&s->model_cv, &s->model_mu) until s->decode_pending > 0
  2. Optional coalescing: Delays up to DS4_SERVER_DECODE_COALESCE_US microseconds to accumulate more requests
  3. Build batch array: Populates ds4_decode_item items[count] from pending slots (lines 10890-10899)
/* Inside decode_worker_main – build the batch */
for (int i = 0; i < s->slot_count; i++) {
    server_slot *slot = &s->slots[i];
    if (!slot->decode_pending) continue;
    slot->decode_pending = false;
    slot->decode_in_flight = true;
    s->decode_pending--;
    members[count] = slot;
    items[count].session = slot->session;
    items[count].token   = slot->decode_token;
    count++;
}

The coalescing timeout is critical for throughput: it trades individual request latency for batch efficiency when request arrival is bursty.

Stage 3: Batch Execution in ds4_sessions_eval_batch

The core library function ds4_sessions_eval_batch (lines 61333-61372 of ds4.c) validates and dispatches the batch:

Validation check Purpose
Same ds4_engine for all slots Ensures compatible compute context
Tokens within vocabulary bounds Prevents out-of-bounds embedding lookups
No duplicate sessions Avoids race conditions in KV cache
Context window not exceeded Maintains attention mechanism correctness
/* ds4_sessions_eval_batch – validation & dispatch */
if (count == 1) return ds4_session_eval(items[0].session, items[0].token, err, errlen);
...
if (e->backend == DS4_BACKEND_CUDA)
    return ds4_sessions_eval_batch_cuda(items, count, err, errlen);
...

For single-item batches, the function falls back to ds4_session_eval to avoid batch overhead. Multi-item batches route to ds4_sessions_eval_batch_cuda or ds4_sessions_eval_batch_metal depending on the engine's backend type.

Synchronization Architecture

The ds4-server uses two distinct mutexes to minimize contention:

  • model_mu: Protects slot state flags, pending counters, and condition variable signaling
  • inference_mu: Guards the engine state during actual GPU/Metal execution

This separation allows the decode worker to assemble batches without blocking new request arrivals. Once assembled, the batch executes under inference_mu to ensure exclusive access to the compute context.

Phase Lock held Operation
Request arrival model_mu Mark slot pending, broadcast model_cv
Worker wake-up model_mu Collect slots, optionally coalesce
Batch execution inference_mu Call ds4_sessions_eval_batch
Result writeback model_mu Store per-slot results, signal clients

Error Handling and Atomicity

ds4_sessions_eval_batch enforces all-or-nothing semantics: if the backend returns an error, the function invalidates every session in the batch. This prevents partial progress where some slots advance their KV cache while others fail.

After backend completion, decode_worker_main copies results to each slot:

  • decode_rc – return code
  • decode_err – error string (if any)
  • decode_done = true – completion flag
  • pthread_cond_broadcast(&s->model_cv) – wake waiting client threads

Clients observe the same interface regardless of batching: they block in server_eval_token until their slot's decode_done flag is set.

Key Source Files

File Lines Responsibility
ds4_server.c 10801-10931 server_eval_token, decode_worker_main, slot management
ds4.c 61333-61372 ds4_sessions_eval_batch, validation, backend dispatch
ds4_session.c — Single-session ds4_session_eval for non-batched fallback
ds4_tp.c — Thread-pool utilities for slot worker management

Summary

  • ds4-server uses batched mode to aggregate concurrent inference requests into single GPU dispatches
  • The three-stage pipeline (enqueue → build → execute) decouples client threads from compute scheduling
  • Two-mutex design (model_mu + inference_mu) maximizes concurrency while protecting engine state
  • Coalescing timeout (DS4_SERVER_DECODE_COALESCE_US) trades latency for throughput under variable load
  • All-or-nothing error handling ensures consistent session state across batch failures

Frequently Asked Questions

How does ds4-server decide between batched and non-batched execution?

The server checks a batched_mode flag at the entry of server_eval_token. When disabled, each request acquires inference_mu immediately and calls ds4_session_eval synchronously. When enabled, requests are deferred to the decode worker thread for batch assembly. This flag is typically set at server startup based on configuration.

What happens if only one request is pending during coalescing?

The batch builder proceeds with a single-item batch. ds4_sessions_eval_batch detects count == 1 and falls back to ds4_session_eval automatically, avoiding unnecessary batch overhead. The coalescing timeout does not delay single requests indefinitely—it triggers as soon as the wait expires or additional requests arrive.

Why does ds4-server use two separate mutexes?

model_mu protects mutable slot state and condition variables accessed frequently by both client threads and the decode worker. inference_mu serializes access to the GPU/Metal context, which has much higher contention cost. Separating them allows request enqueue to proceed in parallel with batch preparation, only serializing during the actual kernel dispatch.

Can batched sessions interfere with each other's results?

No. While sessions share a GPU batch dispatch, each maintains independent KV cache and output buffers. ds4_sessions_eval_batch validates that no session appears twice in the same batch. Per-slot results are written back to separate memory locations before clients are signaled, ensuring isolation despite shared compute.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →