# P2P RDMA Weight Transfer Performance in Miles: Benchmarks, Trade‑offs, and Implementation

> Discover P2P RDMA weight transfer performance in Miles. Slash transfer times up to 86% for MoE models and understand benchmarks trade-offs. Optimize your deep learning.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: performance
- Published: 2026-09-06

---

**P2P RDMA weight transfer in Miles replaces NCCL broadcast with direct rank‑to‑rank RDMA writes, cutting transfer time by up to 86% for large MoE models while adding modest overhead for smaller or dense models.**

The Miles training framework supports two modes for moving updated model weights from training ranks to rollout engines: the default **NCCL broadcast** and the **P2P (RDMA) weight transfer** mode activated with `--update-weight-transfer-mode p2p`. This article breaks down the performance implications of the P2P path, grounded in the actual implementation in `radixark/miles`.

## How P2P RDMA Weight Transfer Works

The P2P mode fundamentally restructures data movement. Instead of broadcasting weight shards to every rollout rank, each training rank writes only the specific shards required by its assigned target rollout rank(s) directly into remote GPU memory using RDMA.

In [`miles/backends/training_utils/weight_update/protocols/p2p.py`](https://github.com/radixark/miles/blob/main/miles/backends/training_utils/weight_update/protocols/p2p.py) lines 38‑40, the core write operation occurs:

```python
transfer_engine.batch_transfer_sync_write(
    remote_session.session_id,
    source_ptrs,    # CPU data pointers

    target_ptrs,    # remote GPU addresses

    source_lens     # byte sizes

)

```

This direct write eliminates the redundant copies inherent in NCCL broadcast, reducing network traffic roughly by the number of rollout ranks.

## Memory and Bandwidth Characteristics

### Memory Footprint: O(1) Regardless of Scale

The P2P implementation maintains constant memory usage through a **single pinned CPU buffer** registered once and reused for all targets. The `register_cpu_memory` function in [`p2p_transfer_utils.py`](https://github.com/radixark/miles/blob/main/p2p_transfer_utils.py) lines 86‑99 handles this registration:

```python
from miles.backends.training_utils.weight_update.protocols.p2p_transfer_utils import register_cpu_memory

weight_registry = register_cpu_memory(params_dict, transfer_engine)

```

This design ensures that adding more rollout engines does not increase the memory footprint on the training rank.

### Bandwidth Scaling

Source bandwidth scales with the number of training ranks (`M // src_pp`), while each target receives proportionally less data based on expert parallelism configuration (`sgl_ep`). The transfer plan in [`p2p_transfer_utils.py`](https://github.com/radixark/miles/blob/main/p2p_transfer_utils.py) lines 64‑81 uses round‑robin load balancing to assign sources to targets, keeping the number of RDMA sessions proportional to the minimum of source and target ranks.

## Pipelining and CPU‑GPU Coordination

The P2P path introduces sophisticated pipelining to overlap operations. In [`p2p.py`](https://github.com/radixark/miles/blob/main/p2p.py) lines 98‑108, writes to non‑last engines are serialized (the manager waits for completion), while writes to the last engine are fire‑and‑forgotten to a background thread pool:

```python
from miles.backends.training_utils.weight_update.protocols.p2p_transfer_utils import P2PTransferManager

manager = P2PTransferManager(num_workers=4, transfer_timeout=30.0)

# Block for non‑last engines

future = manager.submit_returning_future(_do_p2p_write_one_session, remote_session, ready_param_names)
future.result()

# Fire‑and‑forget for last engine

manager.submit(_do_p2p_write_one_session, remote_session, ready_param_names)

```

This overlap allows the next bucket's loading phase to proceed while the final RDMA transfer completes in the background.

The CPU handles RDMA memory registration and writes, while the rollout engine still executes `post_load_weights` on GPU after transfer. For large models like Kimi‑K2, this GPU‑side step adds approximately **884 ms**—a significant portion of total latency that the RDMA optimization cannot eliminate.

## Quantitative Performance Results

Profiling data from [`docs/advanced/p2p-weight-transfer.md`](https://github.com/radixark/miles/blob/main/docs/advanced/p2p-weight-transfer.md) lines 109‑118 reveals model‑dependent outcomes:

| Model | NCCL (ms) | RDMA (ms) | Improvement |
|-------|-----------|-----------|-------------|
| GLM‑4.7‑9B‑Flash | 2,508.6 | 4,229.0 | **+68.6% slower** |
| DeepSeek‑V3 (4‑layer) | 732.2 | 1,260.8 | **+72.2% slower** |
| Qwen3‑30B‑A3B | 2,670.0 | 2,160.2 | **19.1% faster** |
| GLM‑4.5‑Air (106B) | 5,001.1 | 2,637.2 | **47.3% faster** |
| Qwen3‑235B‑A22B | 10,753.6 | 3,162.0 | **70.6% faster** |
| DeepSeek‑V3p2 (744B) | 58,301.5 | 8,479.7 | **85.5% faster** |
| Kimi‑K2 (1T) | 53,279.1 | 7,227.3 | **86.4% faster** |

### Pattern Analysis

- **Small/dense models** (GLM‑4.7‑9B‑Flash, DeepSeek‑V3 4‑layer): RDMA plumbing overhead exceeds benefits; NCCL broadcast remains faster.
- **Large MoE models**: The reduction in network traffic dominates, yielding near‑linear scaling improvements with model size.
- **GPU post‑processing bottleneck**: For Kimi‑K2, even with 86% faster network transfer, the total weight update includes substantial GPU‑side work that limits end‑to‑end gains.

## Enabling and Configuring P2P RDMA

Activate P2P weight transfer at training launch:

```bash
python -m miles.main.train \
    --update-weight-transfer-mode p2p \
    --check-weight-update-equal   # optional validation

```

The `TransferEngine` initialization in [`p2p_transfer_utils.py`](https://github.com/radixark/miles/blob/main/p2p_transfer_utils.py) lines 4‑9 prepares the RDMA device:

```python
from miles.backends.training_utils.weight_update.protocols.p2p_transfer_utils import create_transfer_engine

transfer_engine = create_transfer_engine()

```

Each write is wrapped in a `Future` with configurable timeout (`transfer_timeout`), and failures are logged through `P2PTransferManager.wait_transfers` to prevent cluster‑wide stalls from single RDMA write failures.

## Key Source Files

- **[`miles/backends/training_utils/weight_update/protocols/p2p.py`](https://github.com/radixark/miles/blob/main/miles/backends/training_utils/weight_update/protocols/p2p.py)**: Core protocol with CPU replica, transfer manager, and RDMA write logic.
- **[`miles/backends/training_utils/weight_update/protocols/p2p_transfer_utils.py`](https://github.com/radixark/miles/blob/main/miles/backends/training_utils/weight_update/protocols/p2p_transfer_utils.py)**: Transfer plan construction, pinned memory registration, and `TransferEngine` utilities.
- **[`docs/advanced/p2p-weight-transfer.md`](https://github.com/radixark/miles/blob/main/docs/advanced/p2p-weight-transfer.md)**: Architecture documentation and profiling methodology.
- **[`examples/infra_features/p2p_weight_transfer/README.md`](https://github.com/radixark/miles/blob/main/examples/infra_features/p2p_weight_transfer/README.md)**: Example scripts and GPU post‑processing cost notes.

## Summary

- **P2P RDMA weight transfer achieves up to 86% reduction in inter‑node transfer time** for large MoE models by eliminating broadcast redundancy and leveraging direct RDMA writes.
- **Memory footprint remains constant** regardless of rollout engine count through reused pinned CPU buffers.
- **Pipelining overlaps** the final RDMA transfer with subsequent loading phases, reducing idle time.
- **GPU‑side `post_load_weights` can dominate latency** for specific models, limiting end‑to‑end gains even when network transfer is optimized.
- **Smaller or dense models may perform worse** with P2P due to RDMA setup overhead exceeding the benefit of reduced network traffic.

## Frequently Asked Questions

### When should I use P2P RDMA over NCCL broadcast?

Use P2P RDMA when training **large MoE models** (100B+ parameters with expert parallelism) across many rollout nodes. The network traffic reduction outweighs overhead starting around the 30B‑A3B scale. For smaller dense models, NCCL broadcast remains competitive or faster.

### Does P2P RDMA increase GPU memory usage on rollout engines?

No—the memory optimization occurs on the **training rank side**. The single pinned CPU buffer design in `register_cpu_memory` keeps training rank memory O(1). Rollout engine GPU memory usage is unchanged; the engine still performs `post_load_weights` to move weights into GPU‑optimal layouts after RDMA delivery.

### What causes P2P RDMA to be slower than NCCL for small models?

The **fixed costs** of RDMA session establishment, memory registration, and the CPU‑coordinated transfer protocol exceed the benefit of reduced network traffic when the total data volume is small. GLM‑4.7‑9B‑Flash shows 69% slower performance because its compact size makes NCCL's optimized broadcast more efficient than the per‑rank RDMA coordination overhead.

### How does Miles handle RDMA transfer failures?

Each transfer submits to a `P2PTransferManager` with configurable `transfer_timeout` and returns a `Future`. The manager logs failures and can surface exceptions without stalling the entire weight update across the cluster.