How Ray Placement Groups Manage Distributed Rollout and Training in Miles

Miles uses Ray placement groups to partition GPU resources into isolated actor (training) and rollout (inference) bundles, enabling deterministic, co-located or distributed execution of RL-HF pipelines.

The Miles framework, developed by Radixark, relies on Ray's placement group API to orchestrate distributed reinforcement learning from human feedback (RL-HF). The system splits a cluster's GPUs between model training and response generation, ensuring each component runs on precisely allocated hardware. This design supports both single-node colocation and large-scale multi-node deployments.

Core Architecture: Two Logical Components

Miles defines two primary components that share a cluster's GPU pool:

Component Purpose Placement Group Assignment
Actor (Trainer) Holds and updates the policy model; optionally hosts a critic when --use_critic is enabled First rollout_offset bundles
Rollout (Inference) Runs SGLang engines for response generation and evaluation Remaining bundles from rollout_offset onward

Both components derive from a single placement group created at startup, then logically partitioned based on command-line configuration.

Placement Group Creation and Layout

The orchestration logic lives in miles/ray/placement_group.py. The entry function create_placement_groups(args) follows a three-phase pipeline:

  1. Calculate GPU layout via _get_placement_group_layout(args) (lines 92-105)
  2. Build the placement group via _create_placement_group(num_gpus) (lines 51-90)
  3. Split bundles into actor and rollout slices (lines 108-124)

Determining GPU Layout

The _get_placement_group_layout function inspects CLI arguments to compute:

  • total_gpus: Total bundles to reserve
  • rollout_offset: Index where rollout-specific bundles begin

# From miles/ray/placement_group.py, lines 92-105

def _get_placement_group_layout(args):
    # Handles --debug_train_only, --debug_rollout_only, --colocate, --rollout_external

    # Returns (total_gpus, rollout_offset)

    ...

Key flags alter the layout:

  • --colocate: Uses max(actor_gpus, rollout_gpus) for shared execution
  • --rollout_external: Places rollout GPUs after actor GPUs for physical isolation
  • --debug_train_only / --debug_rollout_only: Allocates zero GPUs to the unused component

Building the Placement Group

The _create_placement_group function constructs bundles requesting {"GPU": 1, "CPU": 1} each, then launches an InfoActor per bundle to discover node IP and physical GPU ID. Bundles are sorted by (node_ip, gpu_id) to ensure deterministic logical ordering.


# Simplified structure from lines 51-90

def _create_placement_group(num_gpus):
    bundles = [{"GPU": 1, "CPU": 1} for _ in range(num_gpus)]
    pg = ray.util.placement_group(bundles, strategy="STRICT_SPREAD")
    # Launch InfoActor per bundle, gather (node_ip, gpu_id), sort

    return PlacementGroupInfo(pg, sorted_indices, sorted_gpu_ids)

Splitting Bundles: Actor vs. Rollout Placement Groups

The create_placement_groups function slices the ordered bundle list:


# From miles/ray/placement_group.py, lines 108-124

def create_placement_groups(args):
    total, offset = _get_placement_group_layout(args)
    pg, actor_idx, actor_gpu = _create_placement_group(total)
    
    # Training slice: first 'offset' bundles

    actor_pg_info = PlacementGroupInfo(pg, actor_idx[:offset], actor_gpu[:offset])
    
    # Rollout slice: remaining bundles

    rollout_pg_info = PlacementGroupInfo(pg, actor_idx[offset:], actor_gpu[offset:])
    
    # Critic reuses actor placement group when --use_critic is set

    return {"actor": actor_pg_info, "rollout": rollout_pg_info}

This returns a dictionary mapping component names to PlacementGroupInfo objects containing:

  • The shared PlacementGroup handle
  • Indices of relevant bundles
  • Physical GPU IDs for each bundle

Instantiating Distributed Components

With placement groups prepared, Miles launches actual workers using Ray actor options to bind to specific bundles.

Rollout Components: Inference and Generation

The create_rollout_components(args) function (also in placement_group.py) builds:

  • InferenceController: Manages SGLang inference engines
  • RolloutExecutor: Ray actor that orchestrates rollout generation

Both are launched with placement group scheduling:


# Conceptual pattern from the implementation

rollout_executor = RolloutExecutor.options(
    num_cpus=1,
    num_gpus=1,
    scheduling_strategy=PlacementGroupSchedulingStrategy(
        placement_group=pg_info.placement_group,
        placement_group_bundle_index=bundle_index
    )
).remote(args, inference_controller)

The InferenceController.init() method creates SGLang servers, each pinned to specific GPU bundles using the indices stored in PlacementGroupInfo.

Training Components: Actor and Critic Models

The create_training_models(args, inference_controller, rollout_executor) function launches TrainerController actors in the actor placement group:


# From train_async.py, lines 33-40

inference_controller, rollout_executor, _ = await create_rollout_components(args)
actor_model, critic_model = await create_training_models(
    args, inference_controller, rollout_executor
)
await update_weights(actor_model, rollout_executor)  # Initial sync

The trainer receives handles to the inference controller and rollout executor, enabling:

  • On-demand rollout requests during training
  • Asynchronous weight updates to the inference engines

Weight Synchronization Between Components

After each training interval, update_weights(actor_model, rollout_executor, rollout_id) propagates the newest model parameters from the trainer to the rollout side. This ensures generation always uses the latest policy.


# From miles/ray/placement_group.py, lines 62-65

async def update_weights(actor_model, rollout_executor, rollout_id=None):
    weights = await actor_model.get_weights.remote()
    await rollout_executor.update_weights.remote(weights, rollout_id)

The rollout executor distributes these weights to all SGLang engines in its placement group slice.

Deployment Modes: Colocation and External Rollout

Miles supports two primary GPU allocation strategies via _get_placement_group_layout:

Mode Flag Behavior
Colocated --colocate Single placement group sized to max(actor_gpus, rollout_gpus); both components share physical GPUs. Ideal for single-node debugging.
External Rollout --rollout_external Distinct GPU sets: actor GPUs first, rollout GPUs after. Enables dedicated inference machines without training interference.

These modes use the same placement group creation code but alter the total_gpus and rollout_offset calculations to achieve different physical mappings.

Complete Usage Example

from miles.ray.placement_group import (
    create_placement_groups,
    create_rollout_components,
    create_training_models,
    update_weights
)

# Parse configuration

args = parse_args()  # Contains actor_num_nodes, actor_num_gpus_per_node, etc.

# Phase 1: Create and partition placement groups

pg_info = create_placement_groups(args)

# pg_info["actor"] -> PlacementGroupInfo for training

# pg_info["rollout"] -> PlacementGroupInfo for inference

# Phase 2: Launch rollout components in rollout placement group

inference_controller, rollout_executor, _ = await create_rollout_components(args)

# Phase 3: Launch training components in actor placement group

actor_model, critic_model = await create_training_models(
    args, inference_controller, rollout_executor
)

# Phase 4: Synchronize initial weights

await update_weights(actor_model, rollout_executor)

# Training loop proceeds with async rollouts and weight updates

Key Source Files

File Responsibility
miles/ray/placement_group.py Placement group creation, layout calculation, bundle splitting, and component factory functions
miles/ray/rollout/inference_controller.py SGLang engine management within rollout placement group
miles/ray/rollout/rollout_executor.py Ray actor for executing rollouts; receives weight updates
miles/ray/train/group.py TrainerController implementation for actor placement group
train_async.py, train.py, train_multi_lora_async.py Entry points orchestrating placement group setup and training loop

Summary

  • Ray placement groups in Miles provide deterministic GPU-to-component mapping for distributed RL-HF training
  • _get_placement_group_layout computes bundle allocation based on CLI flags for colocation or external rollout
  • _create_placement_group builds sorted, InfoActor-discovered bundles for reproducible execution
  • create_placement_groups splits the unified placement group into actor (training) and rollout (inference) slices
  • Weight synchronization via update_weights keeps inference engines current with the training policy
  • Flexible deployment supports both single-node debugging (--colocate) and large-scale distributed execution (--rollout_external)

Frequently Asked Questions

How does Miles ensure deterministic GPU assignment across restarts?

Miles launches an InfoActor in each bundle to discover the physical (node_ip, gpu_id) pair, then sorts bundles by these values before assigning logical indices. This produces consistent ordering regardless of Ray's default scheduling, as implemented in _create_placement_group (lines 51-90).

Can training and rollout run on completely separate machines?

Yes. When --rollout_external is specified, _get_placement_group_layout places rollout bundles after actor bundles in the placement group. If combined with Ray's node resource constraints, this maps to physically distinct machines. The single placement group spans nodes, with bundles allocated to specific node types via Ray's resource requirements.

What happens when --colocate is enabled with fewer GPUs than actor plus rollout need?

The colocation mode uses max(actor_gpus, rollout_gpus) total GPUs, not the sum. Both components time-share the same physical GPUs. This requires careful scheduling to prevent memory exhaustion and is intended for development and debugging rather than production throughput.

How does the critic model fit into the placement group structure?

When --use_critic is enabled, the critic TrainerController simply reuses the actor's PlacementGroupInfo. Both policy and value networks occupy the same GPU bundles, with Ray actor scheduling distinguishing their execution. See create_placement_groups return value handling (lines 108-124).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →