# Pipeline Parallelism vs Tensor Parallelism in Distributed Inference: Key Differences Explained

> Understand pipeline parallelism vs tensor parallelism for distributed inference. Learn how they split models and reduce latency for faster AI.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-04

---

**Pipeline parallelism splits transformer layers across hosts to fit large models, while tensor parallelism splits compute within a layer across GPUs to reduce latency.**

Both strategies in **DwarfStar (ds4)** let you run models across multiple machines or GPUs, but they target different bottlenecks and trade-offs. Understanding which to use depends on whether your priority is model capacity, prefill throughput, or generation latency.

## What Gets Split: Layers vs. Layer Internals

The fundamental architectural difference lies in what work is divided.

**Pipeline parallelism** divides **entire transformer layers** into disjoint layer slices assigned to different hosts. In [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c), the engine handles this by registering layer ranges—each host loads only its assigned slice. When a token enters the system, it traverses the pipeline sequentially: Host A processes layers 0–30, then passes hidden-state activations to Host B for layers 31–output.

**Tensor parallelism** divides **the work inside a single layer**—specifically the heavy routed-expert GEMMs—between two GPUs. Both GPUs work on the *same token simultaneously*, exchanging partial sums after each synchronization point. The [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) file implements this 50/50 split automatically when `--tensor-parallel` is enabled.

## When to Use Each Strategy

| Goal | Best Strategy | Why |
|------|-------------|-----|
| Fit a model that exceeds single-machine memory | Pipeline parallelism | Layers can be distributed across multiple hosts (e.g., 4-bit Flash quant split across two 128 GB MacBooks). |
| Minimize per-token latency | Tensor parallelism | Both GPUs work on identical tokens concurrently, cutting decode time. |
| Maximize prefill throughput | Either; depends on hardware | Pipeline parallelism enables chunked pipelining; tensor parallelism provides near-linear speedup. |
| Simplest setup | Tensor parallelism | No `--layers` configuration needed—automatic 50/50 split. |

## Performance Characteristics

### Prefill Phase

Pipeline parallelism can pipeline prefill chunks: while the coordinator processes chunk *N + 1*, a worker processes chunk *N*. According to the README documentation around line 79, this yields **1.3–1.9× speedup** for large prompts.

Tensor parallelism parallelizes the entire layer work jointly. Benchmarks show **~94 t/s on two Macs** versus **~4.8 t/s on a single-GPU SSD-streaming setup** (README line ~1000)—roughly linear scaling with GPU count.

### Generation Phase

Here the strategies diverge dramatically.

Pipeline parallelism is **strictly autoregressive**: a token cannot be generated until the previous token traverses the entire pipeline. This adds at least one cross-machine hop per token, making generation **~19% slower** than single-process runs (README line ~24).

Tensor parallelism keeps generation fast because both machines work on the *same* token. The two-Mac tensor-parallel decode reaches **~16.8 t/s**—approximately **3× faster** than the SSD-streaming baseline (README line ~71).

## Configuration and Transport

### Pipeline Parallelism Setup

Use `--layers` to define slices and `--role`/`--listen`/`--coordinator` for topology:

```sh

# Worker machine: loads latter half

./ds4 -m gguf/DeepSeek-V4-Flash-Q4K-Layers-31-output.gguf \
      --role worker \
      --layers 31:output \
      --coordinator 169.254.43.68 1234

# Coordinator machine: loads first half

./ds4 -m gguf/DeepSeek-V4-Flash-Q4K-Layers00-30.gguf \
      --role coordinator \
      --layers 0:30 \
      --listen 169.254.43.68 1234

```

Activations flow **host-to-host over plain TCP**. The coordinator does not relay traffic—workers forward directly to the next worker (README line ~100).

### Tensor Parallelism Setup

Enable with `--tensor-parallel` (or `--cuda-tensor-parallel` for CUDA). No `--layers` required:

```sh
MODEL=gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf

# Worker starts first

./ds4 -m "$MODEL" --tensor-parallel --role worker \
      --coordinator 10.99.0.2 9911 --transport rdma

# Coordinator follows

./ds4 -m "$MODEL" --tensor-parallel --role coordinator \
      --listen 10.99.0.2 9911 --transport rdma -c 8192 \
      -p "Tell me something about the sea."

```

By default, tensor parallelism uses **RDMA over Thunderbolt** when available, falling back to TCP. Force transport with `--transport rdma` or `--transport tcp`.

## Implementation in ds4 Source Code

The codebase cleanly separates these paths:

- **[`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c)** (lines ~357–463): Parses `--tensor-parallel`, validates role/transport combinations, and drives the 50/50 TP setup.
- **[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)** (engine-side starting line ~35791): Implements both pipeline and tensor-parallel state management.
- **[`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)** (line ~12831): Handles CUDA-tensor-parallel server logic.
- **[`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c)** (lines ~244–250): Provides user-facing help text for `--tensor-parallel`.

## Summary

- **Pipeline parallelism** splits **model layers** across hosts to fit oversized models and accelerate prefill through chunk pipelining—at the cost of slower generation due to per-token latency accumulation.
- **Tensor parallelism** splits **intra-layer compute** across GPUs to minimize per-token latency and maximize generation throughput—ideal when both machines can hold the full model.
- Choose **pipeline** when memory is the constraint; choose **tensor** when latency is the constraint.
- Configure pipeline with explicit `--layers` ranges; tensor parallelism auto-configures a 50/50 split via `--tensor-parallel`.

## Frequently Asked Questions

### Can I combine pipeline and tensor parallelism in ds4?

No—the current implementation in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) enforces mutual exclusion. The option parser validates that `--tensor-parallel` and `--layers` are not used together. You must choose one strategy per inference run.

### Why is generation slower with pipeline parallelism despite using two machines?

Each generated token must traverse the full pipeline sequentially. The cross-machine hop adds latency that cannot be overlapped because generation is autoregressive—token *N+1* depends on token *N* being complete. This structural constraint makes pipeline parallelism ~19% slower for decode than single-process runs.

### Does tensor parallelism require identical GPUs?

The ds4 implementation assumes symmetric capacity for the 50/50 split, but does not strictly require identical hardware. Performance depends on the slower GPU. For best results with `--tensor-parallel`, use matched machines connected by low-latency links like Thunderbolt 5 or RDMA-capable NICs.

### What transport should I use for tensor parallelism?

Prefer **`--transport rdma`** when available (Thunderbolt 5 on modern Macs). This bypasses TCP overhead and reduces synchronization latency between GPUs. The fallback TCP transport works but adds measurable latency to partial-sum exchanges.