# Best Practices for Multi-Session Batching in ds4-server

> Discover ds4 server best practices for multi-session batching. Maximize throughput and control latency with GPU kernel launch coalescing and pre-allocated server slots.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: best-practices
- Published: 2026-08-08

---

**ds4-server groups independent decode sessions into a single GPU kernel launch using a coalescing window and pre-allocated server slots to maximize throughput while controlling latency.**

The `ds4-server` component of the antirez/ds4 repository implements a high-performance inference server capable of processing multiple independent sessions in parallel. By leveraging multi-session batching, the server aggregates decode requests from different clients into a single GPU kernel execution, significantly improving hardware utilization. This article examines the implementation details in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) and [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) to provide concrete configuration strategies for production deployments.

## How Multi-Session Batching Works

The core batching mechanism resides in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c), where a dedicated worker thread coordinates between client connection slots and the GPU compute backend.

### The Decode Worker Thread

The `decode_worker_main` function (lines 97-130) implements the central batching loop. It maintains an array of pending decode items, monitors slot states, and invokes the generic batch API `ds4_sessions_eval_batch` to execute the computation. This design allows the server to process many independent sessions within a single kernel launch.

### Per-Slot State Management

Each client connection owns a `server_slot` structure. When a token is ready for decoding, the slot sets the `decode_pending` flag and stores the token in `decode_token` (lines 54-64). This state machine enables the worker thread to identify ready work without scanning the entire connection pool, eliminating lock contention during request handling.

### Coalescing and Batch Formation

To maximize batch size without introducing excessive latency, the worker implements a coalescing window controlled by `DS4_SERVER_DECODE_COALESCE_US`. The thread uses `pthread_cond_timedwait` to pause execution, allowing additional slots to set `decode_pending` before proceeding. The worker collects all ready slots into `items[]` and `members[]` arrays (lines 31-43), with the resulting `count` variable determining the final batch size.

### Kernel Execution and Result Distribution

The actual GPU computation occurs in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) within `ds4_sessions_eval_batch` (lines 61387-61445), which dispatches to Metal or CUDA specializations. After execution completes, errors propagate back to individual slots via the `decode_rc` field, and the main thread wakes waiting clients to return results (lines 60-70).

## Recommended Configuration Parameters

| Setting | Description | Recommendation |
|---------|-------------|----------------|
| `--batched-session N` | Pre-allocates **N** independent KV sessions | Set to expected concurrent users (16-64) |
| `DS4_SERVER_DECODE_COALESCE_US` | Microseconds to wait for additional requests before forming a batch | ≤2000 µs for latency-sensitive workloads; 5000-10000 µs for throughput-oriented servers |
| `DS4_SERVER_BATCH_LOG` | Enables per-batch logging of count and elapsed time | Enable during debugging; disable in production to avoid I/O overhead |
| GPU selection (`--gpu` / `CUDA_VISIBLE_DEVICES`) | Determines target GPU device | On multi-GPU nodes, allocate one GPU per server process |

## Tuning the Coalescing Window

The coalescing window represents the primary trade-off between throughput and latency in multi-session batching. A window set below 500 µs frequently produces single-token batches, leaving GPU tensor cores underutilized. Conversely, exceeding 10 ms introduces perceptible delays that violate real-time chat expectations. For interactive applications, maintain the window under 2 ms to preserve sub-10 ms per-token latency. For offline batch processing, extend the window to 5-10 ms to maximize aggregate throughput.

## Common Pitfalls and Solutions

**Insufficient Session Pre-Allocation**
Setting `--batched-session` too low causes "session already has a decode in flight" errors when the `decode_pending` flag remains set on all slots. Increase this parameter to exceed your peak concurrency level.

**Zero Coalescing Timeout**
Disabling the coalescing window (setting `DS4_SERVER_DECODE_COALESCE_US` to 0) forces every decode to run as a separate batch, drastically reducing throughput. Always specify a positive value, such as 2000 µs.

**GPU Memory Exhaustion**
If `ds4_sessions_eval_batch` returns "batched decode failed," the combined batch exceeds available VRAM. Reduce `--batched-session` or decrease the coalescing timeout to limit the number of concurrent contexts per kernel.

**Backend Mismatch**
Running `make test-cuda-session-batch` against a Metal build produces "model-backend mismatch" errors. Ensure the server and test suites use consistent backends (CUDA or Metal) as verified in [`tests/test_metal_session_batch.c`](https://github.com/antirez/ds4/blob/main/tests/test_metal_session_batch.c) and [`tests/test_cuda_session_batch.c`](https://github.com/antirez/ds4/blob/main/tests/test_cuda_session_batch.c).

## Practical Configuration Examples

### High-Throughput Server Setup

```bash
export DS4_SERVER_DECODE_COALESCE_US=5000
export DS4_SERVER_BATCH_LOG=1

./ds4_server \
    --listen 0.0.0.0:8080 \
    --gpu 0 \
    --batched-session 32 \
    --model deepseek-v4-flash \
    --model-path /models/deepseek-v4-flash.gguf

```

This configuration allows up to 5 milliseconds for batch formation across 32 pre-allocated sessions, maximizing GPU utilization during peak load.

### Low-Latency Chat Configuration

```bash
export DS4_SERVER_DECODE_COALESCE_US=1000
export DS4_SERVER_BATCH_LOG=0

./ds4_server --batched-session 16 --gpu 0 ...

```

The 1 millisecond window maintains responsive sub-10 ms per-token latency while still grouping 2-4 concurrent requests when users type simultaneously.

### Debugging Batch Performance

Enable `DS4_SERVER_BATCH_LOG` to output diagnostic lines such as `decode batch count=12 elapsed=3.214 ms status=ok` (as implemented in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) lines 54-58). Use this output to verify that batches grow beyond single tokens during peak traffic and to identify when coalescing timeouts are too aggressive.

## Summary

- Pre-allocate sessions using `--batched-session` to match expected concurrency and eliminate dynamic allocation overhead during request handling.
- Tune `DS4_SERVER_DECODE_COALESCE_US` between 1000-10000 µs to balance latency requirements against GPU utilization.
- Monitor actual batch sizes via `DS4_SERVER_BATCH_LOG` during performance validation to ensure the coalescing window functions as expected.
- Ensure batch sizes remain within GPU memory limits to avoid `ds4_sessions_eval_batch` failures; reduce slots or coalescing time if "batched decode failed" errors occur.
- Match backend types (CUDA/Metal) between the server binary and test suites to prevent runtime compatibility errors.

## Frequently Asked Questions

### How does ds4-server handle multiple client sessions simultaneously?

ds4-server assigns each client a `server_slot` structure. The `decode_worker_main` thread continuously scans these slots, groups those with the `decode_pending` flag set into a batch array, and executes them via `ds4_sessions_eval_batch` in a single GPU kernel launch, with results distributed back to individual slots via `decode_rc`.

### What is the optimal value for DS4_SERVER_DECODE_COALESCE_US?

For latency-sensitive applications like real-time chat, keep the value at or below 2000 microseconds to maintain responsiveness. For throughput-oriented workloads where individual request latency matters less, use 5000-10000 microseconds to allow larger batches to form and maximize GPU utilization.

### Why am I seeing "session already has a decode in flight" errors?

This error occurs when all pre-allocated slots are occupied with pending decodes and new requests arrive. Increase the `--batched-session` parameter to provide more slots than your peak number of concurrent users, ensuring the `decode_pending` flag can clear before new requests arrive.

### Can I run multiple ds4-server processes on different GPUs?

Yes. On multi-GPU nodes, launch one process per GPU using the `--gpu` flag or `CUDA_VISIBLE_DEVICES` environment variable, ensuring each instance uses the same `--batched-session` value. This scales total throughput linearly while maintaining independent batching per device.