Difference Between Pipeline Parallelism and Tensor Parallelism in ds4

Pipeline parallelism distributes model layers across multiple workers to enable multi-device inference, while tensor parallelism splits individual layer computations across two identical machines to double per-layer throughput.

The antirez/ds4 inference engine implements two distinct distributed strategies to scale large language models beyond single-device memory limits. Understanding the difference between pipeline parallelism and tensor parallelism in ds4 is essential for optimizing throughput based on your hardware topology and model architecture.

Architectural Goals and Layer Distribution

The fundamental distinction lies in how each approach partitions the model graph across compute nodes.

Pipeline Parallelism: Stage-Based Execution

Pipeline parallelism splits the model by layers across multiple workers, assigning each worker a specific stage of the network. As implemented in ds4_distributed.c (lines 3378–3830), this approach enables pipelined prefill, where input tokens stream through sequential stages. Each worker executes its assigned layers and passes intermediate activations to the next stage, overlapping computation with KV-cache updates. This method reduces per-device memory pressure by sharding the model vertically, making it ideal for very deep networks that exceed single-GPU memory capacity.

Tensor Parallelism: Intra-Layer Sharding

Tensor parallelism takes the opposite approach, splitting computation inside each layer between two identical machines using a 50/50 expert split. According to ds4_tp.c (lines 55–63), the entire model remains resident on both sides of the connection, but each node computes half of the matrix-multiplication or MoE expert work. This horizontal sharding increases per-layer compute bandwidth by engaging two GPUs simultaneously, benefiting models that fit in memory but are bottlenecked by raw computation speed.

Configuration and Command-Line Interface

Each parallelism mode requires distinct CLI flags and role assignments.

Pipeline parallelism activates through the distributed prefill interface:

  • --dist-prefill-chunk: Defines the token chunk size for pipeline stages
  • --dist-prefill-window: Controls the pipeline depth for overlapping operations
  • Roles: coordinator (stage 0) and worker (subsequent stages)

When initialized, the coordinator logs pipeline activity such as pipelined prefill 2 chunks of up to 1024 tokens, confirming the staged execution path in ds4_distributed.c.

Tensor parallelism requires the --tensor-parallel flag alongside explicit role definition:

ds4 --tensor-parallel --role coordinator --listen 0.0.0.0 6000
ds4 --tensor-parallel --role worker --coordinator <ip> 6000

As documented in ds4_tp.c (lines 98–104), this mode strictly requires the Metal backend and cannot be combined with --dist-prefill-chunk or other distributed layer-slicing options.

Data Movement and Communication Patterns

The transport layer behavior differs significantly between the two approaches.

In pipeline parallelism, workers exchange inter-stage activations and KV-cache slices over RDMA or TCP connections. The logging code in ds4_distributed.c (line 6685) references these "worker-to-worker connections" that move partial results between pipeline stages, with each transfer representing a complete layer boundary crossing.

Tensor parallelism exchanges full-tensor shards (such as split-matrix rows) at every layer boundary using the TP transport layer. Since both machines maintain identical weight replicas, the communication focuses on synchronizing partial computation results rather than passing sequential activations through a deep pipeline.

Hardware Constraints and Backend Requirements

The choice of parallelism is often dictated by hardware capabilities rather than preference.

Pipeline parallelism offers maximum flexibility, functioning with any backend that supports distributed communication. Stages can be unevenly sized to accommodate heterogeneous hardware, and the approach scales beyond two machines.

Tensor parallelism imposes strict constraints:

  • Backend restriction: Metal backend only (ds4_tp.c, lines 98–104)
  • Topology limitation: Exactly two identical machines (50/50 split)
  • Mutual exclusivity: Cannot combine with --dist-* flags or other distribution strategies

These restrictions exist because tensor parallelism relies on symmetric memory residency and specialized Metal kernel implementations for split-matrix operations.

Practical Implementation Examples

Pipeline Parallelism Setup

Deploy a two-stage pipeline across separate nodes:


# Stage 0: Coordinator node

ds4 --role coordinator --listen 0.0.0.0 5000 \
    --dist-prefill-chunk 1024 --dist-prefill-window 4 \
    --model /path/to/model.gguf

# Stage 1: Worker node connecting to coordinator

ds4 --role worker --coordinator 192.168.1.100 5000 \
    --dist-prefill-chunk 1024 --dist-prefill-window 4 \
    --model /path/to/model.gguf

The coordinator output will confirm pipeline activation: distributed coordinator: pipelined prefill 2 chunks of up to 1024 tokens.

Tensor Parallelism Setup

Configure a 50/50 split across two identical machines:


# Machine A (Leader)

ds4 --tensor-parallel --role coordinator --listen 0.0.0.0 6000 \
    --transport rdma --model /path/to/model.gguf

# Machine B (Follower)

ds4 --tensor-parallel --role worker --coordinator 10.0.0.1 6000 \
    --transport rdma --model /path/to/model.gguf

The CLI validates Metal backend availability before initializing the tensor-parallel transport layer.

Summary

  • Pipeline parallelism splits models vertically by layers, enabling multi-stage inference across heterogeneous devices with flexible stage sizing.
  • Tensor parallelism splits computations horizontally within layers, requiring exactly two identical Metal-backend machines with full model replication.
  • Pipeline mode uses --dist-prefill-chunk and supports arbitrary backend configurations, while tensor mode requires --tensor-parallel and excludes other distribution flags.
  • Data movement in pipelining transfers inter-stage activations sequentially, whereas tensor parallelism synchronizes tensor shards across identical replicas.

Frequently Asked Questions

Can pipeline and tensor parallelism be used together in ds4?

No. As enforced in ds4_tp.c (lines 98–104), the --tensor-parallel flag is mutually exclusive with all --dist-* options including pipeline parallelism. You must choose either vertical layer distribution or horizontal tensor splitting, but not both simultaneously.

What hardware backend is required for tensor parallelism?

Tensor parallelism strictly requires the Metal backend. The initialization code in ds4_tp.c validates Metal availability during startup and will exit with an error if Metal is not present. Pipeline parallelism has no such restriction and works with CUDA, Metal, or CPU backends.

How does pipelined prefill improve inference throughput?

Pipelined prefill improves throughput by overlapping the computation of different input chunks across sequential stages. While stage N processes the current token batch, stage N+1 can begin processing the next chunk, keeping all GPUs active rather than idle-waiting for full layer-stack completion. This micro-batching approach is logged in ds4_distributed.c during coordinator initialization.

When should I choose pipeline parallelism over tensor parallelism?

Choose pipeline parallelism when your model exceeds single-device memory capacity or when scaling beyond two machines. Choose tensor parallelism when your model fits in memory on each of two identical nodes but you need to double the matrix-multiplication throughput per layer. Pipeline parallelism excels for memory-bound deep models; tensor parallelism targets compute-bound scenarios with symmetric hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →