How to Set Up Tensor Parallelism Over RDMA Between Two Macs Using ds4

Tensor parallelism in ds4 enables two identical Apple-silicon Macs to function as a single inference engine by splitting routed-expert weights across a low-latency RDMA data channel, with automatic TCP fallback when Thunderbolt RDMA is unavailable.

Setting up tensor parallelism over RDMA between two Macs using ds4 allows you to distribute large language model inference across two machines using high-speed Thunderbolt interconnects. The ds4 repository implements a specialized transport layer that splits model weights evenly between a leader (rank 0) and worker (rank 1) while maintaining lock-step execution through a deterministic gate exchange protocol.

Understanding the ds4 Tensor Parallel Architecture

The ds4 tensor parallel implementation relies on a strict separation between control plane and data plane operations. This design ensures that session management remains robust while inference data travels over the fastest available path.

Control Channel vs. Data Channel

The architecture uses two distinct communication layers defined in ds4_tp.c and ds4_tp.h:

  • Control channel: A blocking TCP socket that mirrors session commands (create, sync, eval) from the leader to the worker. This channel is established through tp_listen() and tp_dial() (source L65-L88).

  • Data channel: A per-token "gate" payload containing partial vectors produced by each half of the model. When a Thunderbolt-capable RDMA device is present, ds4 transfers this payload via two-sided RDMA SEND/RECV operations; otherwise, it falls back to a full-duplex TCP socket.

Each gate frame carries a small header (ds4_tp_gate_header) that allows the peer to detect desynchronization during inference (source L83-L90).

Memory Slab and RDMA Registration

The slab is a contiguous GPU-visible memory block whose layout is described in ds4_tp.h (source L107-L118). Offsets for output vectors, input vectors, and per-token flags are computed by helper functions including ds4_tp_slab_out_offset() and ds4_tp_slab_in_offset().

When RDMA support is enabled via the DS4_TP_HAVE_VERBS compile-time flag, the library registers this slab with the NIC at runtime, exchanges remote keys, and posts receives in a deterministic receive-window (source L118-L128).

Hardware and Software Prerequisites

Before configuring the cluster, verify that your Macs support the necessary RDMA capabilities.

Thunderbolt RDMA Requirements

To use RDMA transport between two Macs, you must ensure:

  1. Both machines have Thunderbolt interfaces that report a verbs device (typically named rdma_enX).
  2. The Thunderbolt cable directly connects the two machines or routes through a compatible switch.
  3. Both systems run macOS with accessible infiniband/verbs headers (included in standard macOS SDKs).

Building ds4 with RDMA Support

The ds4 build system automatically detects RDMA capabilities on macOS. Compile the project with standard make:

make -j$(nproc)

The build includes the DS4_TP_HAVE_VERBS code paths that enable runtime RDMA device loading.

Configuring the Tensor Parallel Cluster

Launching a distributed inference session requires starting the leader node first, then connecting the worker node with matching transport settings.

Leader Node Configuration (Rank 0)

Start the leader on the first Mac using the --tensor-parallel flag combined with RDMA-specific options. The --tensor-parallel flag automatically sets the distributed role and enables the TP pair:

./ds4-server \
  --tensor-parallel \
  --role leader \
  --listen 0.0.0.0:12345 \
  --transport rdma \
  --rdma-device rdma_en0 \
  --rdma-gid-index 0 \
  model.gguf

The --role, --listen, and --coordinator options are implemented in ds4_tp_parse_cli_arg() (source L75-L92).

Worker Node Configuration (Rank 1)

On the second Mac, start the worker pointing to the leader's IP address:

./ds4-server \
  --tensor-parallel \
  --role worker \
  --coordinator 10.0.0.1:12345 \
  --transport rdma \
  --rdma-device rdma_en0 \
  --rdma-gid-index 0 \
  model.gguf

Both machines must use identical model files (model.gguf) as weights are split deterministically based on rank.

Transport Selection and Fallback Behavior

You can query the active transport mode programmatically using the ds4_tp_is_rdma() accessor, which returns the rdma_active flag set after the hello exchange (source L141-L145):

bool using_rdma = ds4_tp_is_rdma(tp);
printf("Tensor-parallel transport: %s\n", using_rdma ? "RDMA" : "TCP");

If the RDMA device cannot be opened or the connection fails, ds4 automatically falls back to TCP for the data channel without requiring manual intervention.

Implementation Details in ds4_tp.c

The core tensor parallel logic resides in ds4_tp.c, which handles the deterministic exchange of inference tokens.

During inference, each token triggers a gate exchange via ds4_tp_gate_exchange(). This function either posts an RDMA SEND/RECV pair or falls back to TCP send/recv depending on the transport state (source L30-31). The exchange blocks until the remote side's partial vector is fully received, guaranteeing lock-step execution across both Macs.

The protocol initialization sequence creates the slab memory region, establishes the control socket through ds4_tp_create(), and negotiates transport capabilities before the first inference batch.

Summary

  • Tensor parallelism in ds4 splits routed-expert weights across two Apple-silicon Macs using a leader-worker architecture with deterministic gate exchange.
  • Two-channel design separates TCP control traffic from RDMA data traffic, with the data channel implemented in ds4_tp.c and ds4_tp.h.
  • RDMA setup requires Thunderbolt verbs devices (rdma_enX), specified via --rdma-device and --rdma-gid-index CLI arguments.
  • Automatic fallback to TCP occurs when RDMA devices are unavailable, ensuring reliable inference without manual transport selection.
  • Memory management uses a GPU-visible slab with offsets calculated by ds4_tp_slab_out_offset() and related helpers to coordinate partial vector exchanges.

Frequently Asked Questions

What hardware do I need for RDMA between two Macs?

You need two Apple-silicon Macs with Thunderbolt ports that expose RDMA verbs devices (typically appearing as rdma_en0 or similar). The Thunderbolt cable must provide a direct peer-to-peer connection or route through a compatible fabric. Standard Ethernet adapters do not provide the necessary RDMA capabilities for this implementation.

How does ds4 handle network failures during tensor parallelism?

The ds4_tp_gate_exchange() function blocks until the remote side's partial vector is fully received, creating a natural synchronization point. If the control channel (TCP) disconnects, the session terminates immediately. For the data channel, ds4 relies on the underlying RDMA or TCP transport error handling; if RDMA fails during operation, the system does not dynamically switch transports mid-session—reliability depends on the initial connection setup succeeding.

Can I mix RDMA and TCP transports in the same cluster?

No, both the leader and worker must use the same transport type for the data channel. While the control channel always uses TCP, the data channel (gate exchange) must be either RDMA on both nodes or TCP on both nodes. The ds4_tp_is_rdma() function verifies that both sides successfully negotiated the same transport mode during the initial handshake.

Where is the tensor parallelism protocol defined in the source code?

The protocol is defined in ds4_tp.h (public API and slab layout) and implemented in ds4_tp.c (transport logic). Key structures include ds4_tp_gate_header for frame headers and ds4_tp_identity for node roles. The distributed options parsing reuses code from ds4_distributed.h and ds4_distributed.c, which handle address and port configurations for both leader and worker nodes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →