Best Practices for Multi-Session Batching in ds4-server
ds4-server groups independent decode sessions into a single GPU kernel launch using a coalescing window and pre-allocated server slots to maximize throughput while controlling latency.
The ds4-server component of the antirez/ds4 repository implements a high-performance inference server capable of processing multiple independent sessions in parallel. By leveraging multi-session batching, the server aggregates decode requests from different clients into a single GPU kernel execution, significantly improving hardware utilization. This article examines the implementation details in ds4_server.c and ds4.c to provide concrete configuration strategies for production deployments.
How Multi-Session Batching Works
The core batching mechanism resides in ds4_server.c, where a dedicated worker thread coordinates between client connection slots and the GPU compute backend.
The Decode Worker Thread
The decode_worker_main function (lines 97-130) implements the central batching loop. It maintains an array of pending decode items, monitors slot states, and invokes the generic batch API ds4_sessions_eval_batch to execute the computation. This design allows the server to process many independent sessions within a single kernel launch.
Per-Slot State Management
Each client connection owns a server_slot structure. When a token is ready for decoding, the slot sets the decode_pending flag and stores the token in decode_token (lines 54-64). This state machine enables the worker thread to identify ready work without scanning the entire connection pool, eliminating lock contention during request handling.
Coalescing and Batch Formation
To maximize batch size without introducing excessive latency, the worker implements a coalescing window controlled by DS4_SERVER_DECODE_COALESCE_US. The thread uses pthread_cond_timedwait to pause execution, allowing additional slots to set decode_pending before proceeding. The worker collects all ready slots into items[] and members[] arrays (lines 31-43), with the resulting count variable determining the final batch size.
Kernel Execution and Result Distribution
The actual GPU computation occurs in ds4.c within ds4_sessions_eval_batch (lines 61387-61445), which dispatches to Metal or CUDA specializations. After execution completes, errors propagate back to individual slots via the decode_rc field, and the main thread wakes waiting clients to return results (lines 60-70).
Recommended Configuration Parameters
| Setting | Description | Recommendation |
|---|---|---|
--batched-session N |
Pre-allocates N independent KV sessions | Set to expected concurrent users (16-64) |
DS4_SERVER_DECODE_COALESCE_US |
Microseconds to wait for additional requests before forming a batch | ≤2000 µs for latency-sensitive workloads; 5000-10000 µs for throughput-oriented servers |
DS4_SERVER_BATCH_LOG |
Enables per-batch logging of count and elapsed time | Enable during debugging; disable in production to avoid I/O overhead |
GPU selection (--gpu / CUDA_VISIBLE_DEVICES) |
Determines target GPU device | On multi-GPU nodes, allocate one GPU per server process |
Tuning the Coalescing Window
The coalescing window represents the primary trade-off between throughput and latency in multi-session batching. A window set below 500 µs frequently produces single-token batches, leaving GPU tensor cores underutilized. Conversely, exceeding 10 ms introduces perceptible delays that violate real-time chat expectations. For interactive applications, maintain the window under 2 ms to preserve sub-10 ms per-token latency. For offline batch processing, extend the window to 5-10 ms to maximize aggregate throughput.
Common Pitfalls and Solutions
Insufficient Session Pre-Allocation
Setting --batched-session too low causes "session already has a decode in flight" errors when the decode_pending flag remains set on all slots. Increase this parameter to exceed your peak concurrency level.
Zero Coalescing Timeout
Disabling the coalescing window (setting DS4_SERVER_DECODE_COALESCE_US to 0) forces every decode to run as a separate batch, drastically reducing throughput. Always specify a positive value, such as 2000 µs.
GPU Memory Exhaustion
If ds4_sessions_eval_batch returns "batched decode failed," the combined batch exceeds available VRAM. Reduce --batched-session or decrease the coalescing timeout to limit the number of concurrent contexts per kernel.
Backend Mismatch
Running make test-cuda-session-batch against a Metal build produces "model-backend mismatch" errors. Ensure the server and test suites use consistent backends (CUDA or Metal) as verified in tests/test_metal_session_batch.c and tests/test_cuda_session_batch.c.
Practical Configuration Examples
High-Throughput Server Setup
export DS4_SERVER_DECODE_COALESCE_US=5000
export DS4_SERVER_BATCH_LOG=1
./ds4_server \
--listen 0.0.0.0:8080 \
--gpu 0 \
--batched-session 32 \
--model deepseek-v4-flash \
--model-path /models/deepseek-v4-flash.gguf
This configuration allows up to 5 milliseconds for batch formation across 32 pre-allocated sessions, maximizing GPU utilization during peak load.
Low-Latency Chat Configuration
export DS4_SERVER_DECODE_COALESCE_US=1000
export DS4_SERVER_BATCH_LOG=0
./ds4_server --batched-session 16 --gpu 0 ...
The 1 millisecond window maintains responsive sub-10 ms per-token latency while still grouping 2-4 concurrent requests when users type simultaneously.
Debugging Batch Performance
Enable DS4_SERVER_BATCH_LOG to output diagnostic lines such as decode batch count=12 elapsed=3.214 ms status=ok (as implemented in ds4_server.c lines 54-58). Use this output to verify that batches grow beyond single tokens during peak traffic and to identify when coalescing timeouts are too aggressive.
Summary
- Pre-allocate sessions using
--batched-sessionto match expected concurrency and eliminate dynamic allocation overhead during request handling. - Tune
DS4_SERVER_DECODE_COALESCE_USbetween 1000-10000 µs to balance latency requirements against GPU utilization. - Monitor actual batch sizes via
DS4_SERVER_BATCH_LOGduring performance validation to ensure the coalescing window functions as expected. - Ensure batch sizes remain within GPU memory limits to avoid
ds4_sessions_eval_batchfailures; reduce slots or coalescing time if "batched decode failed" errors occur. - Match backend types (CUDA/Metal) between the server binary and test suites to prevent runtime compatibility errors.
Frequently Asked Questions
How does ds4-server handle multiple client sessions simultaneously?
ds4-server assigns each client a server_slot structure. The decode_worker_main thread continuously scans these slots, groups those with the decode_pending flag set into a batch array, and executes them via ds4_sessions_eval_batch in a single GPU kernel launch, with results distributed back to individual slots via decode_rc.
What is the optimal value for DS4_SERVER_DECODE_COALESCE_US?
For latency-sensitive applications like real-time chat, keep the value at or below 2000 microseconds to maintain responsiveness. For throughput-oriented workloads where individual request latency matters less, use 5000-10000 microseconds to allow larger batches to form and maximize GPU utilization.
Why am I seeing "session already has a decode in flight" errors?
This error occurs when all pre-allocated slots are occupied with pending decodes and new requests arrive. Increase the --batched-session parameter to provide more slots than your peak number of concurrent users, ensuring the decode_pending flag can clear before new requests arrive.
Can I run multiple ds4-server processes on different GPUs?
Yes. On multi-GPU nodes, launch one process per GPU using the --gpu flag or CUDA_VISIBLE_DEVICES environment variable, ensuring each instance uses the same --batched-session value. This scales total throughput linearly while maintaining independent batching per device.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →