How to Configure Pipeline Parallelism Across Multiple Machines for Distributed Inference in DwarfStar (ds4)
Pipeline parallelism in DwarfStar (ds4) lets you split transformer layers across multiple hosts using a coordinator-worker architecture, enabling inference on models too large for single-machine RAM.
Distributed inference with pipeline parallelism is essential when working with massive language models that exceed the memory capacity of a single GPU or host. The antirez/ds4 repository implements this through a lightweight binary protocol defined in ds4_distributed.c, allowing you to partition model layers between a coordinator (handling input tokenization and output generation) and one or more workers (executing downstream layer computations).
Understanding the Distributed Architecture
DwarfStar operates in two distinct roles that communicate over TCP or RDMA (Thunderbolt).
| Role | Responsibility |
|---|---|
| Coordinator | Loads initial layers, tokenizes input, manages the output head, and drives the pre-fill pipeline by sending activation chunks to workers. |
| Worker | Holds a contiguous slice of the model (e.g., layers 31 through output) and computes its portion of the forward pass, returning logits to the coordinator. |
The protocol implementation in ds4_distributed.c handles activation transport, while CLI flags controlling this behavior are defined in ds4.h.
Prerequisites and Model Preparation
Before starting distributed inference, you must obtain split GGUF model files where each host loads only the tensors for its assigned layer slice.
Download Split GGUFs
Use the provided script to download the appropriate shards for your role:
# On the coordinator machine
./download_model.sh pro-q4-layers00-30
# On the worker machine
./download_model.sh pro-q4-layers31-output
This creates:
gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf(coordinator)gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf(worker)
Define Layer Ranges
Specify slices using START:END (inclusive) or START:output syntax:
- Coordinator:
--layers 0:30 - Worker:
--layers 31:output
Each host must have a GGUF containing only the tensors for its specified range.
Step-by-Step Configuration
Start the Coordinator
The coordinator acts as the control plane and inference entry point. Launch it with --role coordinator and bind to a reachable address:
./ds4 \
-m gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf \
--role coordinator \
--layers 0:30 \
--listen 169.254.43.68 1234
Key flags explained:
--role coordinatoractivates distributed mode in the runtime.--layersdefines which transformer slice to load into local memory.--listensets the IP and port where workers will connect.
Connect the Worker
The worker connects upstream to the coordinator and executes its assigned layer range:
./ds4 \
-m gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf \
--role worker \
--layers 31:output \
--coordinator 169.254.43.68 1234
Critical parameters:
--role workerconfigures the instance as a downstream executor.--coordinatorspecifies the target IP and port for the upstream connection.
Run Distributed Inference
Once connected, interact with the coordinator exactly like a standalone ds4 instance:
./ds4 -p "Explain quantum computing in simple terms."
The coordinator tokenizes the prompt, pipelines the pre-fill phase across machines, and manages autoregressive generation using the worker's returned logits.
Optimizing Network Performance
Reduce Activation Bandwidth
By default, activations transmit as 32-bit floats. Reduce traffic with quantization:
# On both coordinator and worker
./ds4 ... --dist-activation-bits 16
Valid options are 16 (recommended) or 8 (experimental). According to the source in ds4_distributed.c, this halves or quarters inter-node traffic with minimal accuracy impact.
Enable RDMA Over Thunderbolt
For ultra-low latency between Macs, use RDMA instead of TCP:
# Coordinator
./ds4 \
-m gguf/DeepSeek-V4-Flash-Q4K.gguf \
--role coordinator \
--layers 0:19 \
--listen 10.99.0.2 9911 \
--transport rdma
# Worker
./ds4 \
-m gguf/DeepSeek-V4-Flash-Q4K.gguf \
--role worker \
--layers 20:output \
--coordinator 10.99.0.2 9911 \
--transport rdma
Add --rdma-device or --rdma-gid-index if multiple RDMA interfaces are present.
Tune the Prefill Pipeline
During long pre-fills, the coordinator processes chunk N+1 while the worker handles chunk N. Increase the pipeline depth on high-latency links:
./ds4 ... --dist-prefill-window 4
The default window is 1; increasing this value improves throughput for distributed pre-fill at the cost of higher memory usage.
Debugging and Telemetry
Enable detailed per-hop timing and bandwidth statistics:
./ds4 ... --debug
As implemented in ds4_distributed.c around line 442, this prints:
- Bytes sent and received per activation transfer
- Layer-wise computation latency
- Network protocol overhead
Use this data to verify that activation bits and prefill windows are optimally configured for your specific network topology.
Summary
- Pipeline parallelism requires splitting GGUF files by layer ranges and assigning
coordinatororworkerroles via CLI flags defined inds4.h. - The coordinator handles tokenization, pre-fill pipelining, and generation, while workers compute specific layer slices as defined in
ds4_distributed.c. - Communication uses TCP by default, but RDMA over Thunderbolt provides lower latency for Mac-to-Mac connections.
- Optimize bandwidth with
--dist-activation-bits 16and tune pipeline depth with--dist-prefill-windowto match your network latency. - Generation remains strictly autoregressive and runs locally on the coordinator, making distributed generation slower than single-node inference despite faster pre-fill.
Frequently Asked Questions
What network protocol does ds4 use for distributed inference?
DwarfStar uses a lightweight binary protocol implemented in ds4_distributed.c that runs over plain TCP sockets by default. For Apple Silicon machines connected via Thunderbolt, you can specify --transport rdma to bypass the kernel networking stack and achieve lower latency activation transfers.
Can I use pipeline parallelism with more than two machines?
Yes. While the examples show one coordinator and one worker, the architecture supports chaining multiple workers. Each worker connects to an upstream node (either the coordinator or another worker) and specifies its layer slice with --layers START:END. The final worker in the chain uses START:output to indicate it hosts the output head.
Why is distributed generation slower than single-node inference?
Generation is strictly autoregressive, meaning each token depends on the previous one. Unlike the pre-fill phase where chunks can pipeline in parallel, generation requires round-trip communication for every single token across the network. According to the source code analysis, this serial dependency eliminates the parallelism benefits during the generation phase.
How do I verify my layer splits are correct?
Run the coordinator with --debug to enable telemetry output. The logs will show successful connection handshakes in ds4_distributed.c and report the exact layer ranges loaded by each node. If layer ranges overlap or leave gaps, the initialization will fail with a protocol error before inference begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →