# How to Configure Pipeline Parallelism Across Multiple Machines for Distributed Inference in DwarfStar (ds4)

> Configure pipeline parallelism across multiple machines for distributed inference in DwarfStar ds4. Split transformer layers across hosts for models too large for single-machine RAM.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-05

---

**Pipeline parallelism in DwarfStar (ds4) lets you split transformer layers across multiple hosts using a coordinator-worker architecture, enabling inference on models too large for single-machine RAM.**

Distributed inference with pipeline parallelism is essential when working with massive language models that exceed the memory capacity of a single GPU or host. The `antirez/ds4` repository implements this through a lightweight binary protocol defined in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c), allowing you to partition model layers between a **coordinator** (handling input tokenization and output generation) and one or more **workers** (executing downstream layer computations).

## Understanding the Distributed Architecture

DwarfStar operates in two distinct roles that communicate over TCP or RDMA (Thunderbolt).

| Role | Responsibility |
|------|----------------|
| **Coordinator** | Loads initial layers, tokenizes input, manages the output head, and drives the pre-fill pipeline by sending activation chunks to workers. |
| **Worker** | Holds a contiguous slice of the model (e.g., layers 31 through output) and computes its portion of the forward pass, returning logits to the coordinator. |

The protocol implementation in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c) handles activation transport, while CLI flags controlling this behavior are defined in [`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h).

## Prerequisites and Model Preparation

Before starting distributed inference, you must obtain split GGUF model files where each host loads only the tensors for its assigned layer slice.

### Download Split GGUFs

Use the provided script to download the appropriate shards for your role:

```bash

# On the coordinator machine

./download_model.sh pro-q4-layers00-30

# On the worker machine  

./download_model.sh pro-q4-layers31-output

```

This creates:
- `gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf` (coordinator)
- `gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf` (worker)

### Define Layer Ranges

Specify slices using `START:END` (inclusive) or `START:output` syntax:
- Coordinator: `--layers 0:30`
- Worker: `--layers 31:output`

Each host must have a GGUF containing only the tensors for its specified range.

## Step-by-Step Configuration

### Start the Coordinator

The coordinator acts as the control plane and inference entry point. Launch it with `--role coordinator` and bind to a reachable address:

```bash
./ds4 \
  -m gguf/DeepSeek-V4-Pro-Q4K-Layers00-30.gguf \
  --role coordinator \
  --layers 0:30 \
  --listen 169.254.43.68 1234

```

**Key flags explained:**
- `--role coordinator` activates distributed mode in the runtime.
- `--layers` defines which transformer slice to load into local memory.
- `--listen` sets the IP and port where workers will connect.

### Connect the Worker

The worker connects upstream to the coordinator and executes its assigned layer range:

```bash
./ds4 \
  -m gguf/DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf \
  --role worker \
  --layers 31:output \
  --coordinator 169.254.43.68 1234

```

**Critical parameters:**
- `--role worker` configures the instance as a downstream executor.
- `--coordinator` specifies the target IP and port for the upstream connection.

### Run Distributed Inference

Once connected, interact with the coordinator exactly like a standalone `ds4` instance:

```bash
./ds4 -p "Explain quantum computing in simple terms."

```

The coordinator tokenizes the prompt, pipelines the pre-fill phase across machines, and manages autoregressive generation using the worker's returned logits.

## Optimizing Network Performance

### Reduce Activation Bandwidth

By default, activations transmit as 32-bit floats. Reduce traffic with quantization:

```bash

# On both coordinator and worker

./ds4 ... --dist-activation-bits 16

```

Valid options are `16` (recommended) or `8` (experimental). According to the source in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c), this halves or quarters inter-node traffic with minimal accuracy impact.

### Enable RDMA Over Thunderbolt

For ultra-low latency between Macs, use RDMA instead of TCP:

```bash

# Coordinator

./ds4 \
  -m gguf/DeepSeek-V4-Flash-Q4K.gguf \
  --role coordinator \
  --layers 0:19 \
  --listen 10.99.0.2 9911 \
  --transport rdma

# Worker

./ds4 \
  -m gguf/DeepSeek-V4-Flash-Q4K.gguf \
  --role worker \
  --layers 20:output \
  --coordinator 10.99.0.2 9911 \
  --transport rdma

```

Add `--rdma-device` or `--rdma-gid-index` if multiple RDMA interfaces are present.

### Tune the Prefill Pipeline

During long pre-fills, the coordinator processes chunk N+1 while the worker handles chunk N. Increase the pipeline depth on high-latency links:

```bash
./ds4 ... --dist-prefill-window 4

```

The default window is `1`; increasing this value improves throughput for distributed pre-fill at the cost of higher memory usage.

## Debugging and Telemetry

Enable detailed per-hop timing and bandwidth statistics:

```bash
./ds4 ... --debug

```

As implemented in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c) around line 442, this prints:
- Bytes sent and received per activation transfer
- Layer-wise computation latency
- Network protocol overhead

Use this data to verify that activation bits and prefill windows are optimally configured for your specific network topology.

## Summary

- **Pipeline parallelism** requires splitting GGUF files by layer ranges and assigning `coordinator` or `worker` roles via CLI flags defined in [`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h).
- The coordinator handles tokenization, pre-fill pipelining, and generation, while workers compute specific layer slices as defined in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c).
- Communication uses TCP by default, but **RDMA over Thunderbolt** provides lower latency for Mac-to-Mac connections.
- Optimize bandwidth with `--dist-activation-bits 16` and tune pipeline depth with `--dist-prefill-window` to match your network latency.
- Generation remains strictly autoregressive and runs locally on the coordinator, making distributed generation slower than single-node inference despite faster pre-fill.

## Frequently Asked Questions

### What network protocol does ds4 use for distributed inference?

DwarfStar uses a lightweight binary protocol implemented in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c) that runs over plain TCP sockets by default. For Apple Silicon machines connected via Thunderbolt, you can specify `--transport rdma` to bypass the kernel networking stack and achieve lower latency activation transfers.

### Can I use pipeline parallelism with more than two machines?

Yes. While the examples show one coordinator and one worker, the architecture supports chaining multiple workers. Each worker connects to an upstream node (either the coordinator or another worker) and specifies its layer slice with `--layers START:END`. The final worker in the chain uses `START:output` to indicate it hosts the output head.

### Why is distributed generation slower than single-node inference?

Generation is strictly autoregressive, meaning each token depends on the previous one. Unlike the pre-fill phase where chunks can pipeline in parallel, generation requires round-trip communication for every single token across the network. According to the source code analysis, this serial dependency eliminates the parallelism benefits during the generation phase.

### How do I verify my layer splits are correct?

Run the coordinator with `--debug` to enable telemetry output. The logs will show successful connection handshakes in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c) and report the exact layer ranges loaded by each node. If layer ranges overlap or leave gaps, the initialization will fail with a protocol error before inference begins.