How to Configure Batch Sessions for Multi-User Serving with ds4-server
Enable batched mode by setting --batched-sessions N when launching ds4-server to group multiple client decode requests into a single GPU batch, dramatically improving throughput for concurrent users.
The ds4-server binary in the antirez/ds4 repository supports high-throughput multi-user inference through an intelligent batching mechanism. By configuring batch sessions, you consolidate decode tokens from multiple active clients into unified GPU or Metal operations, eliminating per-token overhead while maintaining strict session isolation. This configuration is essential for production deployments serving many simultaneous users.
How Batched Mode Works
Configuration Flag Activation
In ds4_server.c at line 13379, the server initializes batched mode by evaluating the batched_sessions configuration value:
s.batched_mode = cfg.batched_sessions > 0;
When this condition evaluates to true, the server transitions from single-session processing to batch evaluation mode, fundamentally changing how the decode worker handles incoming token requests.
Slot Management and Request Accumulation
Each client connection receives a dedicated server_slot structure. When a client submits a token request, the server marks the slot with decode_pending = true and increments the global s.decode_pending counter. This lock-free coordination allows the decode worker to identify active sessions requiring processing without blocking client I/O threads.
The Decode Worker Thread
The decode_worker_main thread continuously monitors s.decode_pending. If the pending count remains below the total slot capacity, the worker optionally waits for DS4_SERVER_DECODE_COALESCE_US microseconds (default 2000 µs) to accumulate additional requests. Once the window expires or the batch fills, the worker collects all pending slots into an items[] array and invokes ds4_sessions_eval_batch(items, count, …) at line 11038 of ds4_server.c.
Configuration Options
You can configure batch sessions through command-line arguments or environment variables parsed in ds4_cli.c.
Command-Line Interface:
ds4-server --batched-sessions 32 --model /path/to/model.gguf
Environment Variables:
DS4_SERVER_BATCHED_SESSIONS: Sets the maximum number of concurrent sessions to batch (equivalent to--batched-sessions)DS4_SERVER_DECODE_COALESCE_US: Tunes the microsecond delay for request coalescing (default: 2000)DS4_SERVER_BATCH_LOG: Enables per-batch logging when set to any non-empty value
Step-by-Step Configuration Guide
-
Determine your concurrency requirements. Calculate the expected number of simultaneous users. Set
--batched-sessionsto this value or slightly higher to accommodate peak loads. -
Launch with batching enabled.
ds4-server --batched-sessions 64 --host 0.0.0.0 --port 8080 -
Tune the coalescing window. For high-latency networks or bursty traffic, increase the wait time:
export DS4_SERVER_DECODE_COALESCE_US=5000 -
Enable batch logging for monitoring:
export DS4_SERVER_BATCH_LOG=1 -
Deploy with Docker. Pass environment variables through your container orchestration:
docker run -d \ -p 8080:8080 \ -e DS4_SERVER_BATCHED_SESSIONS=64 \ -e DS4_SERVER_DECODE_COALESCE_US=3000 \ -e DS4_SERVER_BATCH_LOG=1 \ antirez/ds4:latest \ ds4-server --model /models/gguf/model.gguf
Monitoring Batch Performance
When DS4_SERVER_BATCH_LOG is enabled, the server emits structured log entries for each batch operation:
ds4-server: decode batch count=28 elapsed=12.3 ms status=ok
Monitor these metrics to optimize your configuration:
- Batch count: Should consistently approach your
--batched-sessionsvalue during peak load - Elapsed time: Indicates GPU processing latency for the combined batch
- Status: Reports execution errors affecting the entire batch
Summary
- Enable multi-user batching by setting
--batched-sessions NorDS4_SERVER_BATCHED_SESSIONSto a value greater than zero - The server aggregates pending decode requests in
server_slotstructures and processes them viads4_sessions_eval_batch()inds4_server.c - Tune
DS4_SERVER_DECODE_COALESCE_USto balance latency against throughput (default 2000 µs) - Enable
DS4_SERVER_BATCH_LOGto monitor batch sizes and execution times - Changes require a server restart as the
s.batched_modeflag initializes at startup inds4_server.c
Frequently Asked Questions
What is the optimal number for batched-sessions?
Set this value to your expected peak concurrent user count plus a 20% buffer. Each slot consumes memory for KV-cache storage, so balance throughput against available GPU VRAM. Monitor batch utilization logs to verify the server consistently fills batches during peak periods.
How does request coalescing affect latency?
The DS4_SERVER_DECODE_COALESCE_US parameter introduces a deliberate microsecond delay to accumulate requests. A longer window increases throughput by creating larger batches but adds latency for individual users. For real-time applications, reduce this value to 500-1000 µs; for throughput-optimized workloads, increase to 5000-10000 µs.
Can I enable batching on a running server?
No. The s.batched_mode boolean initializes at server startup in ds4_server.c based on the cfg.batched_sessions value. Changing this configuration requires restarting the ds4-server process to reallocate slot arrays and spawn the batch decode worker thread.
Where are batch errors logged?
Batch evaluation errors surface through the ds4_sessions_eval_batch() return code in ds4_server.c. When the batch completes, each slot receives identical rc values and error messages. Enable DS4_SERVER_BATCH_LOG to capture batch-level status, or check individual client responses for per-session error propagation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →