# How to Configure Batch Sessions for Multi-User Serving with ds4-server

> Configure batched sessions for ds4-server to boost multi-user serving throughput. Group client decode requests into GPU batches for dramatic performance gains.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Enable batched mode by setting `--batched-sessions N` when launching ds4-server to group multiple client decode requests into a single GPU batch, dramatically improving throughput for concurrent users.**

The `ds4-server` binary in the antirez/ds4 repository supports high-throughput multi-user inference through an intelligent batching mechanism. By configuring batch sessions, you consolidate decode tokens from multiple active clients into unified GPU or Metal operations, eliminating per-token overhead while maintaining strict session isolation. This configuration is essential for production deployments serving many simultaneous users.

## How Batched Mode Works

### Configuration Flag Activation

In [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) at line 13379, the server initializes batched mode by evaluating the `batched_sessions` configuration value:

```c
s.batched_mode = cfg.batched_sessions > 0;

```

When this condition evaluates to true, the server transitions from single-session processing to batch evaluation mode, fundamentally changing how the decode worker handles incoming token requests.

### Slot Management and Request Accumulation

Each client connection receives a dedicated `server_slot` structure. When a client submits a token request, the server marks the slot with `decode_pending = true` and increments the global `s.decode_pending` counter. This lock-free coordination allows the decode worker to identify active sessions requiring processing without blocking client I/O threads.

### The Decode Worker Thread

The `decode_worker_main` thread continuously monitors `s.decode_pending`. If the pending count remains below the total slot capacity, the worker optionally waits for `DS4_SERVER_DECODE_COALESCE_US` microseconds (default 2000 µs) to accumulate additional requests. Once the window expires or the batch fills, the worker collects all pending slots into an `items[]` array and invokes `ds4_sessions_eval_batch(items, count, …)` at line 11038 of [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c).

## Configuration Options

You can configure batch sessions through command-line arguments or environment variables parsed in [`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c).

**Command-Line Interface:**

```bash
ds4-server --batched-sessions 32 --model /path/to/model.gguf

```

**Environment Variables:**

- `DS4_SERVER_BATCHED_SESSIONS`: Sets the maximum number of concurrent sessions to batch (equivalent to `--batched-sessions`)
- `DS4_SERVER_DECODE_COALESCE_US`: Tunes the microsecond delay for request coalescing (default: 2000)
- `DS4_SERVER_BATCH_LOG`: Enables per-batch logging when set to any non-empty value

## Step-by-Step Configuration Guide

1. **Determine your concurrency requirements.** Calculate the expected number of simultaneous users. Set `--batched-sessions` to this value or slightly higher to accommodate peak loads.

2. **Launch with batching enabled.**

   ```bash
   ds4-server --batched-sessions 64 --host 0.0.0.0 --port 8080
   ```

3. **Tune the coalescing window.** For high-latency networks or bursty traffic, increase the wait time:

   ```bash
   export DS4_SERVER_DECODE_COALESCE_US=5000
   ```

4. **Enable batch logging for monitoring:**

   ```bash
   export DS4_SERVER_BATCH_LOG=1
   ```

5. **Deploy with Docker.** Pass environment variables through your container orchestration:

   ```bash
   docker run -d \
     -p 8080:8080 \
     -e DS4_SERVER_BATCHED_SESSIONS=64 \
     -e DS4_SERVER_DECODE_COALESCE_US=3000 \
     -e DS4_SERVER_BATCH_LOG=1 \
     antirez/ds4:latest \
     ds4-server --model /models/gguf/model.gguf
   ```

## Monitoring Batch Performance

When `DS4_SERVER_BATCH_LOG` is enabled, the server emits structured log entries for each batch operation:

```

ds4-server: decode batch count=28 elapsed=12.3 ms status=ok

```

Monitor these metrics to optimize your configuration:

- **Batch count**: Should consistently approach your `--batched-sessions` value during peak load
- **Elapsed time**: Indicates GPU processing latency for the combined batch
- **Status**: Reports execution errors affecting the entire batch

## Summary

- Enable multi-user batching by setting `--batched-sessions N` or `DS4_SERVER_BATCHED_SESSIONS` to a value greater than zero
- The server aggregates pending decode requests in `server_slot` structures and processes them via `ds4_sessions_eval_batch()` in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)
- Tune `DS4_SERVER_DECODE_COALESCE_US` to balance latency against throughput (default 2000 µs)
- Enable `DS4_SERVER_BATCH_LOG` to monitor batch sizes and execution times
- Changes require a server restart as the `s.batched_mode` flag initializes at startup in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)

## Frequently Asked Questions

### What is the optimal number for batched-sessions?

Set this value to your expected peak concurrent user count plus a 20% buffer. Each slot consumes memory for KV-cache storage, so balance throughput against available GPU VRAM. Monitor batch utilization logs to verify the server consistently fills batches during peak periods.

### How does request coalescing affect latency?

The `DS4_SERVER_DECODE_COALESCE_US` parameter introduces a deliberate microsecond delay to accumulate requests. A longer window increases throughput by creating larger batches but adds latency for individual users. For real-time applications, reduce this value to 500-1000 µs; for throughput-optimized workloads, increase to 5000-10000 µs.

### Can I enable batching on a running server?

No. The `s.batched_mode` boolean initializes at server startup in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) based on the `cfg.batched_sessions` value. Changing this configuration requires restarting the `ds4-server` process to reallocate slot arrays and spawn the batch decode worker thread.

### Where are batch errors logged?

Batch evaluation errors surface through the `ds4_sessions_eval_batch()` return code in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c). When the batch completes, each slot receives identical `rc` values and error messages. Enable `DS4_SERVER_BATCH_LOG` to capture batch-level status, or check individual client responses for per-session error propagation.