How to Configure Pipeline Parallelism Across Multiple Machines for Distributed Inference in DwarfStar (ds4)

Pipeline parallelism in DwarfStar (ds4) lets you split transformer layers across multiple hosts using a coordinator-worker architecture, enabling inference on models too large for single-machine RAM.

Distributed inference with pipeline parallelism is essential when working with massive language models that exceed the memory capacity of a single GPU or host. The antirez/ds4 repository implements this through a lightweight binary protocol defined in ds4_distributed.c, allowing you to partition model layers between a coordinator (handling input tokenization and output generation) and one or more workers (executing downstream layer computations).

Understanding the Distributed Architecture

DwarfStar operates in two distinct roles that communicate over TCP or RDMA (Thunderbolt).

Role Responsibility
Coordinator Loads initial layers, tokenizes input, manages the output head, and drives the pre-fill pipeline by sending activation chunks to workers.
Worker Holds a contiguous slice of the model (e.g., layers 31 through output) and computes its portion of the forward pass, returning logits to the coordinator.

The protocol implementation in ds4_distributed.c handles activation transport, while CLI flags controlling this behavior are defined in ds4.h.

Prerequisites and Model Preparation

Before starting distributed inference, you must obtain split GGUF model files where each host loads only the tensors for its assigned layer slice.

Download Split GGUFs

Use the provided script to download the appropriate shards for your role:


# On the coordinator machine

./download_model.sh pro-q4-layers00-30

# On the worker machine  

./download_model.sh pro-q4-layers31-output

This creates:

  • gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf (coordinator)
  • gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf (worker)

Define Layer Ranges

Specify slices using START:END (inclusive) or START:output syntax:

  • Coordinator: --layers 0:30
  • Worker: --layers 31:output

Each host must have a GGUF containing only the tensors for its specified range.

Step-by-Step Configuration

Start the Coordinator

The coordinator acts as the control plane and inference entry point. Launch it with --role coordinator and bind to a reachable address:

./ds4 \
  -m gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf \
  --role coordinator \
  --layers 0:30 \
  --listen 169.254.43.68 1234

Key flags explained:

  • --role coordinator activates distributed mode in the runtime.
  • --layers defines which transformer slice to load into local memory.
  • --listen sets the IP and port where workers will connect.

Connect the Worker

The worker connects upstream to the coordinator and executes its assigned layer range:

./ds4 \
  -m gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf \
  --role worker \
  --layers 31:output \
  --coordinator 169.254.43.68 1234

Critical parameters:

  • --role worker configures the instance as a downstream executor.
  • --coordinator specifies the target IP and port for the upstream connection.

Run Distributed Inference

Once connected, interact with the coordinator exactly like a standalone ds4 instance:

./ds4 -p "Explain quantum computing in simple terms."

The coordinator tokenizes the prompt, pipelines the pre-fill phase across machines, and manages autoregressive generation using the worker's returned logits.

Optimizing Network Performance

Reduce Activation Bandwidth

By default, activations transmit as 32-bit floats. Reduce traffic with quantization:


# On both coordinator and worker

./ds4 ... --dist-activation-bits 16

Valid options are 16 (recommended) or 8 (experimental). According to the source in ds4_distributed.c, this halves or quarters inter-node traffic with minimal accuracy impact.

Enable RDMA Over Thunderbolt

For ultra-low latency between Macs, use RDMA instead of TCP:


# Coordinator

./ds4 \
  -m gguf/DeepSeek-V4-Flash-Q4K.gguf \
  --role coordinator \
  --layers 0:19 \
  --listen 10.99.0.2 9911 \
  --transport rdma

# Worker

./ds4 \
  -m gguf/DeepSeek-V4-Flash-Q4K.gguf \
  --role worker \
  --layers 20:output \
  --coordinator 10.99.0.2 9911 \
  --transport rdma

Add --rdma-device or --rdma-gid-index if multiple RDMA interfaces are present.

Tune the Prefill Pipeline

During long pre-fills, the coordinator processes chunk N+1 while the worker handles chunk N. Increase the pipeline depth on high-latency links:

./ds4 ... --dist-prefill-window 4

The default window is 1; increasing this value improves throughput for distributed pre-fill at the cost of higher memory usage.

Debugging and Telemetry

Enable detailed per-hop timing and bandwidth statistics:

./ds4 ... --debug

As implemented in ds4_distributed.c around line 442, this prints:

  • Bytes sent and received per activation transfer
  • Layer-wise computation latency
  • Network protocol overhead

Use this data to verify that activation bits and prefill windows are optimally configured for your specific network topology.

Summary

  • Pipeline parallelism requires splitting GGUF files by layer ranges and assigning coordinator or worker roles via CLI flags defined in ds4.h.
  • The coordinator handles tokenization, pre-fill pipelining, and generation, while workers compute specific layer slices as defined in ds4_distributed.c.
  • Communication uses TCP by default, but RDMA over Thunderbolt provides lower latency for Mac-to-Mac connections.
  • Optimize bandwidth with --dist-activation-bits 16 and tune pipeline depth with --dist-prefill-window to match your network latency.
  • Generation remains strictly autoregressive and runs locally on the coordinator, making distributed generation slower than single-node inference despite faster pre-fill.

Frequently Asked Questions

What network protocol does ds4 use for distributed inference?

DwarfStar uses a lightweight binary protocol implemented in ds4_distributed.c that runs over plain TCP sockets by default. For Apple Silicon machines connected via Thunderbolt, you can specify --transport rdma to bypass the kernel networking stack and achieve lower latency activation transfers.

Can I use pipeline parallelism with more than two machines?

Yes. While the examples show one coordinator and one worker, the architecture supports chaining multiple workers. Each worker connects to an upstream node (either the coordinator or another worker) and specifies its layer slice with --layers START:END. The final worker in the chain uses START:output to indicate it hosts the output head.

Why is distributed generation slower than single-node inference?

Generation is strictly autoregressive, meaning each token depends on the previous one. Unlike the pre-fill phase where chunks can pipeline in parallel, generation requires round-trip communication for every single token across the network. According to the source code analysis, this serial dependency eliminates the parallelism benefits during the generation phase.

How do I verify my layer splits are correct?

Run the coordinator with --debug to enable telemetry output. The logs will show successful connection handshakes in ds4_distributed.c and report the exact layer ranges loaded by each node. If layer ranges overlap or leave gaps, the initialization will fail with a protocol error before inference begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →