How to Configure Pipeline Parallelism Across Multiple Machines for ds4
To configure pipeline parallelism across multiple machines for ds4, deploy worker nodes with specific layer ranges and a coordinator node that drives the sampling loop, then tune network latency hiding via --dist-prefill-chunk and --dist-prefill-window flags.
ds4 is an open-source LLM inference engine developed by antirez that enables distributed execution of large language models across multiple physical hosts. When you configure pipeline parallelism across multiple machines for ds4, you partition the model weights and KV-cache across workers while the coordinator manages prompt ingestion and token generation, allowing you to run models that exceed single-machine memory limits.
Distributed Inference Architecture
ds4 implements two complementary parallelism models that determine how computation and state are partitioned across the network.
Layer-Range Partitioning (Coordinator and Workers)
The standard distributed mode splits the model by layer ranges. Each worker owns a contiguous slice of the model—including its weights and associated KV-cache—while the coordinator maintains the prompt buffer, sampling state, and client API endpoint.
In ds4_distributed.c, the function dist_coordinator_prefill_prompt_pipelined orchestrates the flow of hidden-state tensors through the worker chain. The coordinator streams activations sequentially, enabling the full model to be evaluated as if it resided on a single machine despite being physically partitioned.
Tensor-Parallelism Mode (Two-Machine Setup)
For deployments with exactly two machines, ds4 provides a specialized tensor-parallelism path activated by the --tensor-parallel flag. Implemented in ds4_tp.c, this mode automatically configures a 50/50 layer split and establishes a single-step pipeline that exchanges partial attention and feed-forward results between the two peers.
Configuring the Pre-Fill Pipeline
When processing long prompts, ds4 can pipeline the pre-fill phase so that workers begin computing on early token chunks while the coordinator continues transmitting subsequent chunks. This overlaps computation with communication and reduces worker idle time.
Two runtime options control this behavior, defined in ds4_help.c:
--dist-prefill-chunk N: Specifies the size of each pre-fill chunk (in tokens) sent to workers before waiting for results. The default is the session token-cap.--dist-prefill-window N: Sets the maximum number of chunks that may be "in-flight" (sent but not yet acknowledged). The default isworkers + 2, capped at 8.
The coordinator breaks the prompt into chunks of the specified size, pushes the first window of chunks to workers, and slides the window forward as results return. The logic in ds4_distributed.c function dist_coordinator_can_pipeline_prefill determines when additional chunks can be safely dispatched without overwhelming the worker buffer.
Activation Compression for Bandwidth Optimization
Hidden-state transport between machines can be quantized to reduce network bandwidth. Use the --dist-activation-bits N flag to configure the transmission precision:
32(default): Full single-precision floating point16: Half-precision (FP16/BF16)8: 8-bit quantized transmission
Lowering the bit width reduces bandwidth requirements at the cost of small numerical precision loss, which is often negligible for inference workloads.
Multi-Machine Configuration Examples
Standard Distributed Setup with Layer Partitioning
First, start the workers with their respective layer slices and the coordinator address:
# Worker 1 (first half of model)
ds4 --role worker \
--layers 0:19 \
--listen 0.0.0.0 5001 \
--coordinator 192.168.1.100 5000
# Worker 2 (second half of model)
ds4 --role worker \
--layers 20:output \
--listen 0.0.0.0 5002 \
--coordinator 192.168.1.100 5000
Then launch the coordinator with pipeline tuning parameters:
ds4 --role coordinator \
--listen 0.0.0.0 5000 \
--dist-prefill-chunk 128 \
--dist-prefill-window 4 \
--dist-activation-bits 16
Two-Machine Tensor-Parallel Mode
For a simplified two-machine deployment, use the tensor-parallel shortcut:
# Machine 1 (worker role)
ds4 --role worker \
--tensor-parallel \
--listen 0.0.0.0 6000 \
--coordinator 192.168.1.200 6001
# Machine 2 (coordinator role)
ds4 --role coordinator \
--tensor-parallel \
--listen 0.0.0.0 6001 \
--dist-prefill-chunk 256 \
--dist-prefill-window 2
The --tensor-parallel flag automatically configures the appropriate layer slices and activates the specialized transport protocol in ds4_tp.c that exchanges partial results between the two machines.
Key Source Files and Implementation Details
The following files in the antirez/ds4 repository contain the core logic for distributed pipeline execution:
| File | Purpose | Key Components |
|---|---|---|
ds4_help.c |
CLI documentation | Defines --dist-prefill-chunk, --dist-prefill-window, --tensor-parallel, and --dist-activation-bits flags |
ds4_distributed.c |
Core distributed runtime | Implements dist_coordinator_can_pipeline_prefill and dist_coordinator_prefill_prompt_pipelined for chunk management |
ds4_tp.c |
Tensor-parallel implementation | Handles two-machine mode transport, address switching, and partial result aggregation |
ds4.c |
Session orchestration | Manages pipeline capture flags and generic pre-fill/decode pipeline plumbing |
ds4_server.c |
HTTP/CLI server facade | Entry point that parses command-line arguments and launches the appropriate runtime mode |
Summary
- ds4 supports two distributed modes: layer-range partitioning for arbitrary worker counts and tensor-parallelism optimized specifically for two machines.
- Configure pipeline parallelism using
--dist-prefill-chunkto set token chunk sizes and--dist-prefill-windowto control concurrent in-flight chunks. - Optimize network bandwidth with
--dist-activation-bitsby compressing hidden-state tensors to 16 or 8 bits. - Workers specify their model slice via
--layers A:Band connect to the coordinator at--coordinator HOST PORT. - The coordinator drives the sampling loop and automatically pipelines pre-fill and decode stages across the worker network according to the constraints defined in
ds4_distributed.c.
Frequently Asked Questions
What is the difference between the coordinator and worker roles in ds4?
The coordinator holds the prompt, manages the autoregressive sampling loop, and exposes the API endpoint, while workers own specific layer ranges of the model weights and KV-cache and process hidden-state tensors sent by the coordinator. The coordinator initiates all pipeline stages and aggregates results from workers to produce final token outputs.
How does the pre-fill pipeline improve distributed performance?
The pre-fill pipeline allows workers to begin computing on early token chunks while the coordinator continues transmitting subsequent portions of the prompt, effectively hiding network latency. By adjusting --dist-prefill-window, you control how many chunks remain "in-flight," ensuring workers stay busy rather than waiting for the entire prompt to arrive.
Can tensor-parallelism mode be used with more than two machines?
No. The --tensor-parallel flag in ds4_tp.c is specifically designed as a two-machine optimization that automatically splits the model 50/50 between peers. For deployments with three or more machines, you must use the standard distributed mode with explicit --layers ranges assigned to each worker node.
What is the default activation bit width for hidden-state transfers?
If --dist-activation-bits is not specified, ds4 defaults to 32 bits for all hidden-state transfers between machines. While this preserves full numerical accuracy, it maximizes network bandwidth usage; most deployments benefit from switching to 16-bit transmission to reduce latency without significant quality degradation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →