Pipeline Parallelism vs Tensor Parallelism in Distributed Inference: Key Differences Explained
Pipeline parallelism splits transformer layers across hosts to fit large models, while tensor parallelism splits compute within a layer across GPUs to reduce latency.
Both strategies in DwarfStar (ds4) let you run models across multiple machines or GPUs, but they target different bottlenecks and trade-offs. Understanding which to use depends on whether your priority is model capacity, prefill throughput, or generation latency.
What Gets Split: Layers vs. Layer Internals
The fundamental architectural difference lies in what work is divided.
Pipeline parallelism divides entire transformer layers into disjoint layer slices assigned to different hosts. In ds4.c, the engine handles this by registering layer ranges—each host loads only its assigned slice. When a token enters the system, it traverses the pipeline sequentially: Host A processes layers 0–30, then passes hidden-state activations to Host B for layers 31–output.
Tensor parallelism divides the work inside a single layer—specifically the heavy routed-expert GEMMs—between two GPUs. Both GPUs work on the same token simultaneously, exchanging partial sums after each synchronization point. The ds4_tp.c file implements this 50/50 split automatically when --tensor-parallel is enabled.
When to Use Each Strategy
| Goal | Best Strategy | Why |
|---|---|---|
| Fit a model that exceeds single-machine memory | Pipeline parallelism | Layers can be distributed across multiple hosts (e.g., 4-bit Flash quant split across two 128 GB MacBooks). |
| Minimize per-token latency | Tensor parallelism | Both GPUs work on identical tokens concurrently, cutting decode time. |
| Maximize prefill throughput | Either; depends on hardware | Pipeline parallelism enables chunked pipelining; tensor parallelism provides near-linear speedup. |
| Simplest setup | Tensor parallelism | No --layers configuration needed—automatic 50/50 split. |
Performance Characteristics
Prefill Phase
Pipeline parallelism can pipeline prefill chunks: while the coordinator processes chunk N + 1, a worker processes chunk N. According to the README documentation around line 79, this yields 1.3–1.9× speedup for large prompts.
Tensor parallelism parallelizes the entire layer work jointly. Benchmarks show ~94 t/s on two Macs versus ~4.8 t/s on a single-GPU SSD-streaming setup (README line ~1000)—roughly linear scaling with GPU count.
Generation Phase
Here the strategies diverge dramatically.
Pipeline parallelism is strictly autoregressive: a token cannot be generated until the previous token traverses the entire pipeline. This adds at least one cross-machine hop per token, making generation ~19% slower than single-process runs (README line ~24).
Tensor parallelism keeps generation fast because both machines work on the same token. The two-Mac tensor-parallel decode reaches ~16.8 t/s—approximately 3× faster than the SSD-streaming baseline (README line ~71).
Configuration and Transport
Pipeline Parallelism Setup
Use --layers to define slices and --role/--listen/--coordinator for topology:
# Worker machine: loads latter half
./ds4 -m gguf/DeepSeek-V4-Flash-Q4K-Layers-31-output.gguf \
--role worker \
--layers 31:output \
--coordinator 169.254.43.68 1234
# Coordinator machine: loads first half
./ds4 -m gguf/DeepSeek-V4-Flash-Q4K-Layers00-30.gguf \
--role coordinator \
--layers 0:30 \
--listen 169.254.43.68 1234
Activations flow host-to-host over plain TCP. The coordinator does not relay traffic—workers forward directly to the next worker (README line ~100).
Tensor Parallelism Setup
Enable with --tensor-parallel (or --cuda-tensor-parallel for CUDA). No --layers required:
MODEL=gguf/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf
# Worker starts first
./ds4 -m "$MODEL" --tensor-parallel --role worker \
--coordinator 10.99.0.2 9911 --transport rdma
# Coordinator follows
./ds4 -m "$MODEL" --tensor-parallel --role coordinator \
--listen 10.99.0.2 9911 --transport rdma -c 8192 \
-p "Tell me something about the sea."
By default, tensor parallelism uses RDMA over Thunderbolt when available, falling back to TCP. Force transport with --transport rdma or --transport tcp.
Implementation in ds4 Source Code
The codebase cleanly separates these paths:
ds4_tp.c(lines ~357–463): Parses--tensor-parallel, validates role/transport combinations, and drives the 50/50 TP setup.ds4.c(engine-side starting line ~35791): Implements both pipeline and tensor-parallel state management.ds4_server.c(line ~12831): Handles CUDA-tensor-parallel server logic.ds4_help.c(lines ~244–250): Provides user-facing help text for--tensor-parallel.
Summary
- Pipeline parallelism splits model layers across hosts to fit oversized models and accelerate prefill through chunk pipelining—at the cost of slower generation due to per-token latency accumulation.
- Tensor parallelism splits intra-layer compute across GPUs to minimize per-token latency and maximize generation throughput—ideal when both machines can hold the full model.
- Choose pipeline when memory is the constraint; choose tensor when latency is the constraint.
- Configure pipeline with explicit
--layersranges; tensor parallelism auto-configures a 50/50 split via--tensor-parallel.
Frequently Asked Questions
Can I combine pipeline and tensor parallelism in ds4?
No—the current implementation in ds4_tp.c enforces mutual exclusion. The option parser validates that --tensor-parallel and --layers are not used together. You must choose one strategy per inference run.
Why is generation slower with pipeline parallelism despite using two machines?
Each generated token must traverse the full pipeline sequentially. The cross-machine hop adds latency that cannot be overlapped because generation is autoregressive—token N+1 depends on token N being complete. This structural constraint makes pipeline parallelism ~19% slower for decode than single-process runs.
Does tensor parallelism require identical GPUs?
The ds4 implementation assumes symmetric capacity for the 50/50 split, but does not strictly require identical hardware. Performance depends on the slower GPU. For best results with --tensor-parallel, use matched machines connected by low-latency links like Thunderbolt 5 or RDMA-capable NICs.
What transport should I use for tensor parallelism?
Prefer --transport rdma when available (Thunderbolt 5 on modern Macs). This bypasses TCP overhead and reduces synchronization latency between GPUs. The fallback TCP transport works but adds measurable latency to partial-sum exchanges.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →