# How ds4-server Handles Batched Sessions for Multi-User Inference

> Discover how ds4-server efficiently manages batched sessions for multi-user inference. Learn about its three-stage pipeline, GPU batching, and mutexes for seamless request processing.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: internals
- Published: 2026-08-04

---

**ds4-server processes concurrent client requests through a three-stage pipeline that aggregates individual token evaluations into GPU batches, using separate mutexes for request queuing and inference execution.**

The `ds4-server` binary in the antirez/ds4 repository implements **batched mode** to maximize GPU/Metal utilization when serving multiple simultaneous users. Rather than executing each inference request immediately, the server collects pending work into batches that share a single kernel dispatch. This article explains the complete mechanism—from request arrival through batch execution—based on the source code in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) and [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).

## The Three Stages of Batched Session Handling

### Stage 1: Request Enqueue in server_eval_token

When a client thread requests token evaluation, it enters `server_eval_token` at lines 10801-10809 of [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c). The path depends on the server configuration:

- **Batched mode disabled**: The function acquires `inference_mu`, calls `ds4_session_eval` synchronously, and returns immediately.
- **Batched mode enabled**: The request is recorded in its slot (`decode_token`, `decode_pending = true`), and `pthread_cond_broadcast(&s->model_cv)` wakes the decode worker.

```c
/* Client thread – enqueue a token for batched evaluation */
int rc = server_eval_token(srv, slot, token, err_buf, sizeof(err_buf));
if (rc != 0) { /* handle error */ }

```

The slot-based design decouples client threads from inference execution. Each slot maintains state flags that the decode worker inspects without blocking incoming connections.

### Stage 2: Batch Construction in decode_worker_main

The dedicated decode worker thread runs `decode_worker_main` (lines 10855-10931 of [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)). Its operation follows a precise sequence:

1. **Wait for work**: Blocks on `pthread_cond_timedwait(&s->model_cv, &s->model_mu)` until `s->decode_pending > 0`
2. **Optional coalescing**: Delays up to `DS4_SERVER_DECODE_COALESCE_US` microseconds to accumulate more requests
3. **Build batch array**: Populates `ds4_decode_item items[count]` from pending slots (lines 10890-10899)

```c
/* Inside decode_worker_main – build the batch */
for (int i = 0; i < s->slot_count; i++) {
    server_slot *slot = &s->slots[i];
    if (!slot->decode_pending) continue;
    slot->decode_pending = false;
    slot->decode_in_flight = true;
    s->decode_pending--;
    members[count] = slot;
    items[count].session = slot->session;
    items[count].token   = slot->decode_token;
    count++;
}

```

The coalescing timeout is critical for throughput: it trades individual request latency for batch efficiency when request arrival is bursty.

### Stage 3: Batch Execution in ds4_sessions_eval_batch

The core library function `ds4_sessions_eval_batch` (lines 61333-61372 of [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)) validates and dispatches the batch:

| Validation check | Purpose |
|------------------|---------|
| Same `ds4_engine` for all slots | Ensures compatible compute context |
| Tokens within vocabulary bounds | Prevents out-of-bounds embedding lookups |
| No duplicate sessions | Avoids race conditions in KV cache |
| Context window not exceeded | Maintains attention mechanism correctness |

```c
/* ds4_sessions_eval_batch – validation & dispatch */
if (count == 1) return ds4_session_eval(items[0].session, items[0].token, err, errlen);
...
if (e->backend == DS4_BACKEND_CUDA)
    return ds4_sessions_eval_batch_cuda(items, count, err, errlen);
...

```

For single-item batches, the function falls back to `ds4_session_eval` to avoid batch overhead. Multi-item batches route to `ds4_sessions_eval_batch_cuda` or `ds4_sessions_eval_batch_metal` depending on the engine's backend type.

## Synchronization Architecture

The ds4-server uses **two distinct mutexes** to minimize contention:

- **`model_mu`**: Protects slot state flags, pending counters, and condition variable signaling
- **`inference_mu`**: Guards the engine state during actual GPU/Metal execution

This separation allows the decode worker to assemble batches without blocking new request arrivals. Once assembled, the batch executes under `inference_mu` to ensure exclusive access to the compute context.

| Phase | Lock held | Operation |
|-------|-----------|-----------|
| Request arrival | `model_mu` | Mark slot pending, broadcast `model_cv` |
| Worker wake-up | `model_mu` | Collect slots, optionally coalesce |
| Batch execution | `inference_mu` | Call `ds4_sessions_eval_batch` |
| Result writeback | `model_mu` | Store per-slot results, signal clients |

## Error Handling and Atomicity

`ds4_sessions_eval_batch` enforces **all-or-nothing semantics**: if the backend returns an error, the function invalidates every session in the batch. This prevents partial progress where some slots advance their KV cache while others fail.

After backend completion, `decode_worker_main` copies results to each slot:
- `decode_rc` – return code
- `decode_err` – error string (if any)
- `decode_done = true` – completion flag
- `pthread_cond_broadcast(&s->model_cv)` – wake waiting client threads

Clients observe the same interface regardless of batching: they block in `server_eval_token` until their slot's `decode_done` flag is set.

## Key Source Files

| File | Lines | Responsibility |
|------|-------|----------------|
| [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) | 10801-10931 | `server_eval_token`, `decode_worker_main`, slot management |
| [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) | 61333-61372 | `ds4_sessions_eval_batch`, validation, backend dispatch |
| [`ds4_session.c`](https://github.com/antirez/ds4/blob/main/ds4_session.c) | — | Single-session `ds4_session_eval` for non-batched fallback |
| [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) | — | Thread-pool utilities for slot worker management |

## Summary

- ds4-server uses **batched mode** to aggregate concurrent inference requests into single GPU dispatches
- The **three-stage pipeline** (enqueue → build → execute) decouples client threads from compute scheduling
- **Two-mutex design** (`model_mu` + `inference_mu`) maximizes concurrency while protecting engine state
- **Coalescing timeout** (`DS4_SERVER_DECODE_COALESCE_US`) trades latency for throughput under variable load
- **All-or-nothing error handling** ensures consistent session state across batch failures

## Frequently Asked Questions

### How does ds4-server decide between batched and non-batched execution?

The server checks a `batched_mode` flag at the entry of `server_eval_token`. When disabled, each request acquires `inference_mu` immediately and calls `ds4_session_eval` synchronously. When enabled, requests are deferred to the decode worker thread for batch assembly. This flag is typically set at server startup based on configuration.

### What happens if only one request is pending during coalescing?

The batch builder proceeds with a single-item batch. `ds4_sessions_eval_batch` detects `count == 1` and falls back to `ds4_session_eval` automatically, avoiding unnecessary batch overhead. The coalescing timeout does not delay single requests indefinitely—it triggers as soon as the wait expires or additional requests arrive.

### Why does ds4-server use two separate mutexes?

`model_mu` protects mutable slot state and condition variables accessed frequently by both client threads and the decode worker. `inference_mu` serializes access to the GPU/Metal context, which has much higher contention cost. Separating them allows request enqueue to proceed in parallel with batch preparation, only serializing during the actual kernel dispatch.

### Can batched sessions interfere with each other's results?

No. While sessions share a GPU batch dispatch, each maintains independent KV cache and output buffers. `ds4_sessions_eval_batch` validates that no session appears twice in the same batch. Per-slot results are written back to separate memory locations before clients are signaled, ensuring isolation despite shared compute.