P2P RDMA Weight Transfer Performance in Miles: Benchmarks, Trade‑offs, and Implementation

P2P RDMA weight transfer in Miles replaces NCCL broadcast with direct rank‑to‑rank RDMA writes, cutting transfer time by up to 86% for large MoE models while adding modest overhead for smaller or dense models.

The Miles training framework supports two modes for moving updated model weights from training ranks to rollout engines: the default NCCL broadcast and the P2P (RDMA) weight transfer mode activated with --update-weight-transfer-mode p2p. This article breaks down the performance implications of the P2P path, grounded in the actual implementation in radixark/miles.

How P2P RDMA Weight Transfer Works

The P2P mode fundamentally restructures data movement. Instead of broadcasting weight shards to every rollout rank, each training rank writes only the specific shards required by its assigned target rollout rank(s) directly into remote GPU memory using RDMA.

In miles/backends/training_utils/weight_update/protocols/p2p.py lines 38‑40, the core write operation occurs:

transfer_engine.batch_transfer_sync_write(
    remote_session.session_id,
    source_ptrs,    # CPU data pointers

    target_ptrs,    # remote GPU addresses

    source_lens     # byte sizes

)

This direct write eliminates the redundant copies inherent in NCCL broadcast, reducing network traffic roughly by the number of rollout ranks.

Memory and Bandwidth Characteristics

Memory Footprint: O(1) Regardless of Scale

The P2P implementation maintains constant memory usage through a single pinned CPU buffer registered once and reused for all targets. The register_cpu_memory function in p2p_transfer_utils.py lines 86‑99 handles this registration:

from miles.backends.training_utils.weight_update.protocols.p2p_transfer_utils import register_cpu_memory

weight_registry = register_cpu_memory(params_dict, transfer_engine)

This design ensures that adding more rollout engines does not increase the memory footprint on the training rank.

Bandwidth Scaling

Source bandwidth scales with the number of training ranks (M // src_pp), while each target receives proportionally less data based on expert parallelism configuration (sgl_ep). The transfer plan in p2p_transfer_utils.py lines 64‑81 uses round‑robin load balancing to assign sources to targets, keeping the number of RDMA sessions proportional to the minimum of source and target ranks.

Pipelining and CPU‑GPU Coordination

The P2P path introduces sophisticated pipelining to overlap operations. In p2p.py lines 98‑108, writes to non‑last engines are serialized (the manager waits for completion), while writes to the last engine are fire‑and‑forgotten to a background thread pool:

from miles.backends.training_utils.weight_update.protocols.p2p_transfer_utils import P2PTransferManager

manager = P2PTransferManager(num_workers=4, transfer_timeout=30.0)

# Block for non‑last engines

future = manager.submit_returning_future(_do_p2p_write_one_session, remote_session, ready_param_names)
future.result()

# Fire‑and‑forget for last engine

manager.submit(_do_p2p_write_one_session, remote_session, ready_param_names)

This overlap allows the next bucket's loading phase to proceed while the final RDMA transfer completes in the background.

The CPU handles RDMA memory registration and writes, while the rollout engine still executes post_load_weights on GPU after transfer. For large models like Kimi‑K2, this GPU‑side step adds approximately 884 ms—a significant portion of total latency that the RDMA optimization cannot eliminate.

Quantitative Performance Results

Profiling data from docs/advanced/p2p-weight-transfer.md lines 109‑118 reveals model‑dependent outcomes:

Model NCCL (ms) RDMA (ms) Improvement
GLM‑4.7‑9B‑Flash 2,508.6 4,229.0 +68.6% slower
DeepSeek‑V3 (4‑layer) 732.2 1,260.8 +72.2% slower
Qwen3‑30B‑A3B 2,670.0 2,160.2 19.1% faster
GLM‑4.5‑Air (106B) 5,001.1 2,637.2 47.3% faster
Qwen3‑235B‑A22B 10,753.6 3,162.0 70.6% faster
DeepSeek‑V3p2 (744B) 58,301.5 8,479.7 85.5% faster
Kimi‑K2 (1T) 53,279.1 7,227.3 86.4% faster

Pattern Analysis

  • Small/dense models (GLM‑4.7‑9B‑Flash, DeepSeek‑V3 4‑layer): RDMA plumbing overhead exceeds benefits; NCCL broadcast remains faster.
  • Large MoE models: The reduction in network traffic dominates, yielding near‑linear scaling improvements with model size.
  • GPU post‑processing bottleneck: For Kimi‑K2, even with 86% faster network transfer, the total weight update includes substantial GPU‑side work that limits end‑to‑end gains.

Enabling and Configuring P2P RDMA

Activate P2P weight transfer at training launch:

python -m miles.main.train \
    --update-weight-transfer-mode p2p \
    --check-weight-update-equal   # optional validation

The TransferEngine initialization in p2p_transfer_utils.py lines 4‑9 prepares the RDMA device:

from miles.backends.training_utils.weight_update.protocols.p2p_transfer_utils import create_transfer_engine

transfer_engine = create_transfer_engine()

Each write is wrapped in a Future with configurable timeout (transfer_timeout), and failures are logged through P2PTransferManager.wait_transfers to prevent cluster‑wide stalls from single RDMA write failures.

Key Source Files

Summary

  • P2P RDMA weight transfer achieves up to 86% reduction in inter‑node transfer time for large MoE models by eliminating broadcast redundancy and leveraging direct RDMA writes.
  • Memory footprint remains constant regardless of rollout engine count through reused pinned CPU buffers.
  • Pipelining overlaps the final RDMA transfer with subsequent loading phases, reducing idle time.
  • GPU‑side post_load_weights can dominate latency for specific models, limiting end‑to‑end gains even when network transfer is optimized.
  • Smaller or dense models may perform worse with P2P due to RDMA setup overhead exceeding the benefit of reduced network traffic.

Frequently Asked Questions

When should I use P2P RDMA over NCCL broadcast?

Use P2P RDMA when training large MoE models (100B+ parameters with expert parallelism) across many rollout nodes. The network traffic reduction outweighs overhead starting around the 30B‑A3B scale. For smaller dense models, NCCL broadcast remains competitive or faster.

Does P2P RDMA increase GPU memory usage on rollout engines?

No—the memory optimization occurs on the training rank side. The single pinned CPU buffer design in register_cpu_memory keeps training rank memory O(1). Rollout engine GPU memory usage is unchanged; the engine still performs post_load_weights to move weights into GPU‑optimal layouts after RDMA delivery.

What causes P2P RDMA to be slower than NCCL for small models?

The fixed costs of RDMA session establishment, memory registration, and the CPU‑coordinated transfer protocol exceed the benefit of reduced network traffic when the total data volume is small. GLM‑4.7‑9B‑Flash shows 69% slower performance because its compact size makes NCCL's optimized broadcast more efficient than the per‑rank RDMA coordination overhead.

How does Miles handle RDMA transfer failures?

Each transfer submits to a P2PTransferManager with configurable transfer_timeout and returns a Future. The manager logs failures and can surface exceptions without stalling the entire weight update across the cluster.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →