# How to Set Up Tensor Parallelism Over RDMA Between Two Macs Using ds4

> Set up tensor parallelism over RDMA between two Macs using ds4. Split weights for a single inference engine with low latency and TCP fallback.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Tensor parallelism in ds4 enables two identical Apple-silicon Macs to function as a single inference engine by splitting routed-expert weights across a low-latency RDMA data channel, with automatic TCP fallback when Thunderbolt RDMA is unavailable.**

Setting up **tensor parallelism over RDMA between two Macs using ds4** allows you to distribute large language model inference across two machines using high-speed Thunderbolt interconnects. The ds4 repository implements a specialized transport layer that splits model weights evenly between a leader (rank 0) and worker (rank 1) while maintaining lock-step execution through a deterministic gate exchange protocol.

## Understanding the ds4 Tensor Parallel Architecture

The ds4 tensor parallel implementation relies on a strict separation between control plane and data plane operations. This design ensures that session management remains robust while inference data travels over the fastest available path.

### Control Channel vs. Data Channel

The architecture uses two distinct communication layers defined in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) and [`ds4_tp.h`](https://github.com/antirez/ds4/blob/main/ds4_tp.h):

- **Control channel**: A blocking TCP socket that mirrors session commands (create, sync, eval) from the leader to the worker. This channel is established through `tp_listen()` and `tp_dial()` ([source L65-L88](/blob/main/ds4_tp.c#L65-L88)).

- **Data channel**: A per-token "gate" payload containing partial vectors produced by each half of the model. When a Thunderbolt-capable RDMA device is present, ds4 transfers this payload via two-sided RDMA SEND/RECV operations; otherwise, it falls back to a full-duplex TCP socket.

Each gate frame carries a small header (`ds4_tp_gate_header`) that allows the peer to detect desynchronization during inference ([source L83-L90](/blob/main/ds4_tp.c#L83-L90)).

### Memory Slab and RDMA Registration

The **slab** is a contiguous GPU-visible memory block whose layout is described in [`ds4_tp.h`](https://github.com/antirez/ds4/blob/main/ds4_tp.h) ([source L107-L118](/blob/main/ds4_tp.h#L107-L118)). Offsets for output vectors, input vectors, and per-token flags are computed by helper functions including `ds4_tp_slab_out_offset()` and `ds4_tp_slab_in_offset()`.

When RDMA support is enabled via the `DS4_TP_HAVE_VERBS` compile-time flag, the library registers this slab with the NIC at runtime, exchanges remote keys, and posts receives in a deterministic receive-window ([source L118-L128](/blob/main/ds4_tp.c#L118-L128)).

## Hardware and Software Prerequisites

Before configuring the cluster, verify that your Macs support the necessary RDMA capabilities.

### Thunderbolt RDMA Requirements

To use RDMA transport between two Macs, you must ensure:

1. Both machines have Thunderbolt interfaces that report a verbs device (typically named `rdma_enX`).
2. The Thunderbolt cable directly connects the two machines or routes through a compatible switch.
3. Both systems run macOS with accessible `infiniband/verbs` headers (included in standard macOS SDKs).

### Building ds4 with RDMA Support

The ds4 build system automatically detects RDMA capabilities on macOS. Compile the project with standard make:

```bash
make -j$(nproc)

```

The build includes the `DS4_TP_HAVE_VERBS` code paths that enable runtime RDMA device loading.

## Configuring the Tensor Parallel Cluster

Launching a distributed inference session requires starting the leader node first, then connecting the worker node with matching transport settings.

### Leader Node Configuration (Rank 0)

Start the leader on the first Mac using the `--tensor-parallel` flag combined with RDMA-specific options. The `--tensor-parallel` flag automatically sets the distributed role and enables the TP pair:

```bash
./ds4-server \
  --tensor-parallel \
  --role leader \
  --listen 0.0.0.0:12345 \
  --transport rdma \
  --rdma-device rdma_en0 \
  --rdma-gid-index 0 \
  model.gguf

```

The `--role`, `--listen`, and `--coordinator` options are implemented in `ds4_tp_parse_cli_arg()` ([source L75-L92](/blob/main/ds4_tp.c#L75-L92)).

### Worker Node Configuration (Rank 1)

On the second Mac, start the worker pointing to the leader's IP address:

```bash
./ds4-server \
  --tensor-parallel \
  --role worker \
  --coordinator 10.0.0.1:12345 \
  --transport rdma \
  --rdma-device rdma_en0 \
  --rdma-gid-index 0 \
  model.gguf

```

Both machines must use identical model files (`model.gguf`) as weights are split deterministically based on rank.

### Transport Selection and Fallback Behavior

You can query the active transport mode programmatically using the `ds4_tp_is_rdma()` accessor, which returns the `rdma_active` flag set after the hello exchange ([source L141-L145](/blob/main/ds4_tp.c#L141-L145)):

```c
bool using_rdma = ds4_tp_is_rdma(tp);
printf("Tensor-parallel transport: %s\n", using_rdma ? "RDMA" : "TCP");

```

If the RDMA device cannot be opened or the connection fails, ds4 automatically falls back to TCP for the data channel without requiring manual intervention.

## Implementation Details in ds4_tp.c

The core tensor parallel logic resides in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c), which handles the deterministic exchange of inference tokens.

During inference, each token triggers a **gate exchange** via `ds4_tp_gate_exchange()`. This function either posts an RDMA SEND/RECV pair or falls back to TCP send/recv depending on the transport state ([source L30-31](/blob/main/ds4_tp.c#L30-L31)). The exchange blocks until the remote side's partial vector is fully received, guaranteeing lock-step execution across both Macs.

The protocol initialization sequence creates the slab memory region, establishes the control socket through `ds4_tp_create()`, and negotiates transport capabilities before the first inference batch.

## Summary

- **Tensor parallelism in ds4** splits routed-expert weights across two Apple-silicon Macs using a leader-worker architecture with deterministic gate exchange.
- **Two-channel design** separates TCP control traffic from RDMA data traffic, with the data channel implemented in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) and [`ds4_tp.h`](https://github.com/antirez/ds4/blob/main/ds4_tp.h).
- **RDMA setup** requires Thunderbolt verbs devices (`rdma_enX`), specified via `--rdma-device` and `--rdma-gid-index` CLI arguments.
- **Automatic fallback** to TCP occurs when RDMA devices are unavailable, ensuring reliable inference without manual transport selection.
- **Memory management** uses a GPU-visible slab with offsets calculated by `ds4_tp_slab_out_offset()` and related helpers to coordinate partial vector exchanges.

## Frequently Asked Questions

### What hardware do I need for RDMA between two Macs?

You need two Apple-silicon Macs with Thunderbolt ports that expose RDMA verbs devices (typically appearing as `rdma_en0` or similar). The Thunderbolt cable must provide a direct peer-to-peer connection or route through a compatible fabric. Standard Ethernet adapters do not provide the necessary RDMA capabilities for this implementation.

### How does ds4 handle network failures during tensor parallelism?

The `ds4_tp_gate_exchange()` function blocks until the remote side's partial vector is fully received, creating a natural synchronization point. If the control channel (TCP) disconnects, the session terminates immediately. For the data channel, ds4 relies on the underlying RDMA or TCP transport error handling; if RDMA fails during operation, the system does not dynamically switch transports mid-session—reliability depends on the initial connection setup succeeding.

### Can I mix RDMA and TCP transports in the same cluster?

No, both the leader and worker must use the same transport type for the data channel. While the control channel always uses TCP, the data channel (gate exchange) must be either RDMA on both nodes or TCP on both nodes. The `ds4_tp_is_rdma()` function verifies that both sides successfully negotiated the same transport mode during the initial handshake.

### Where is the tensor parallelism protocol defined in the source code?

The protocol is defined in [`ds4_tp.h`](https://github.com/antirez/ds4/blob/main/ds4_tp.h) (public API and slab layout) and implemented in [`ds4_tp.c`](https://github.com/antirez/ds4/blob/main/ds4_tp.c) (transport logic). Key structures include `ds4_tp_gate_header` for frame headers and `ds4_tp_identity` for node roles. The distributed options parsing reuses code from [`ds4_distributed.h`](https://github.com/antirez/ds4/blob/main/ds4_distributed.h) and [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c), which handle address and port configurations for both leader and worker nodes.