# How Ray Placement Groups Manage Distributed Rollout and Training in Miles

> Learn how Miles leverages Ray placement groups to manage distributed rollout and training. Achieve deterministic execution of RL-HF pipelines with partitioned GPU resources.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: architecture
- Published: 2026-09-06

---

**Miles uses Ray placement groups to partition GPU resources into isolated actor (training) and rollout (inference) bundles, enabling deterministic, co-located or distributed execution of RL-HF pipelines.**

The Miles framework, developed by Radixark, relies on Ray's placement group API to orchestrate distributed reinforcement learning from human feedback (RL-HF). The system splits a cluster's GPUs between model training and response generation, ensuring each component runs on precisely allocated hardware. This design supports both single-node colocation and large-scale multi-node deployments.

## Core Architecture: Two Logical Components

Miles defines two primary components that share a cluster's GPU pool:

| Component | Purpose | Placement Group Assignment |
|-----------|---------|---------------------------|
| **Actor (Trainer)** | Holds and updates the policy model; optionally hosts a critic when `--use_critic` is enabled | First `rollout_offset` bundles |
| **Rollout (Inference)** | Runs SGLang engines for response generation and evaluation | Remaining bundles from `rollout_offset` onward |

Both components derive from a single placement group created at startup, then logically partitioned based on command-line configuration.

## Placement Group Creation and Layout

The orchestration logic lives in [`miles/ray/placement_group.py`](https://github.com/radixark/miles/blob/main/miles/ray/placement_group.py). The entry function `create_placement_groups(args)` follows a three-phase pipeline:

1. **Calculate GPU layout** via `_get_placement_group_layout(args)` (lines 92-105)
2. **Build the placement group** via `_create_placement_group(num_gpus)` (lines 51-90)
3. **Split bundles** into actor and rollout slices (lines 108-124)

### Determining GPU Layout

The `_get_placement_group_layout` function inspects CLI arguments to compute:

- `total_gpus`: Total bundles to reserve
- `rollout_offset`: Index where rollout-specific bundles begin

```python

# From miles/ray/placement_group.py, lines 92-105

def _get_placement_group_layout(args):
    # Handles --debug_train_only, --debug_rollout_only, --colocate, --rollout_external

    # Returns (total_gpus, rollout_offset)

    ...

```

Key flags alter the layout:
- `--colocate`: Uses max(actor_gpus, rollout_gpus) for shared execution
- `--rollout_external`: Places rollout GPUs after actor GPUs for physical isolation
- `--debug_train_only` / `--debug_rollout_only`: Allocates zero GPUs to the unused component

### Building the Placement Group

The `_create_placement_group` function constructs bundles requesting `{"GPU": 1, "CPU": 1}` each, then launches an `InfoActor` per bundle to discover node IP and physical GPU ID. Bundles are sorted by `(node_ip, gpu_id)` to ensure deterministic logical ordering.

```python

# Simplified structure from lines 51-90

def _create_placement_group(num_gpus):
    bundles = [{"GPU": 1, "CPU": 1} for _ in range(num_gpus)]
    pg = ray.util.placement_group(bundles, strategy="STRICT_SPREAD")
    # Launch InfoActor per bundle, gather (node_ip, gpu_id), sort

    return PlacementGroupInfo(pg, sorted_indices, sorted_gpu_ids)

```

## Splitting Bundles: Actor vs. Rollout Placement Groups

The `create_placement_groups` function slices the ordered bundle list:

```python

# From miles/ray/placement_group.py, lines 108-124

def create_placement_groups(args):
    total, offset = _get_placement_group_layout(args)
    pg, actor_idx, actor_gpu = _create_placement_group(total)
    
    # Training slice: first 'offset' bundles

    actor_pg_info = PlacementGroupInfo(pg, actor_idx[:offset], actor_gpu[:offset])
    
    # Rollout slice: remaining bundles

    rollout_pg_info = PlacementGroupInfo(pg, actor_idx[offset:], actor_gpu[offset:])
    
    # Critic reuses actor placement group when --use_critic is set

    return {"actor": actor_pg_info, "rollout": rollout_pg_info}

```

This returns a dictionary mapping component names to `PlacementGroupInfo` objects containing:
- The shared `PlacementGroup` handle
- Indices of relevant bundles
- Physical GPU IDs for each bundle

## Instantiating Distributed Components

With placement groups prepared, Miles launches actual workers using Ray actor options to bind to specific bundles.

### Rollout Components: Inference and Generation

The `create_rollout_components(args)` function (also in [`placement_group.py`](https://github.com/radixark/miles/blob/main/placement_group.py)) builds:

- **`InferenceController`**: Manages SGLang inference engines
- **`RolloutExecutor`**: Ray actor that orchestrates rollout generation

Both are launched with placement group scheduling:

```python

# Conceptual pattern from the implementation

rollout_executor = RolloutExecutor.options(
    num_cpus=1,
    num_gpus=1,
    scheduling_strategy=PlacementGroupSchedulingStrategy(
        placement_group=pg_info.placement_group,
        placement_group_bundle_index=bundle_index
    )
).remote(args, inference_controller)

```

The `InferenceController.init()` method creates SGLang servers, each pinned to specific GPU bundles using the indices stored in `PlacementGroupInfo`.

### Training Components: Actor and Critic Models

The `create_training_models(args, inference_controller, rollout_executor)` function launches `TrainerController` actors in the actor placement group:

```python

# From train_async.py, lines 33-40

inference_controller, rollout_executor, _ = await create_rollout_components(args)
actor_model, critic_model = await create_training_models(
    args, inference_controller, rollout_executor
)
await update_weights(actor_model, rollout_executor)  # Initial sync

```

The trainer receives handles to the inference controller and rollout executor, enabling:
- On-demand rollout requests during training
- Asynchronous weight updates to the inference engines

## Weight Synchronization Between Components

After each training interval, `update_weights(actor_model, rollout_executor, rollout_id)` propagates the newest model parameters from the trainer to the rollout side. This ensures generation always uses the latest policy.

```python

# From miles/ray/placement_group.py, lines 62-65

async def update_weights(actor_model, rollout_executor, rollout_id=None):
    weights = await actor_model.get_weights.remote()
    await rollout_executor.update_weights.remote(weights, rollout_id)

```

The rollout executor distributes these weights to all SGLang engines in its placement group slice.

## Deployment Modes: Colocation and External Rollout

Miles supports two primary GPU allocation strategies via `_get_placement_group_layout`:

| Mode | Flag | Behavior |
|------|------|----------|
| **Colocated** | `--colocate` | Single placement group sized to `max(actor_gpus, rollout_gpus)`; both components share physical GPUs. Ideal for single-node debugging. |
| **External Rollout** | `--rollout_external` | Distinct GPU sets: actor GPUs first, rollout GPUs after. Enables dedicated inference machines without training interference. |

These modes use the same placement group creation code but alter the `total_gpus` and `rollout_offset` calculations to achieve different physical mappings.

## Complete Usage Example

```python
from miles.ray.placement_group import (
    create_placement_groups,
    create_rollout_components,
    create_training_models,
    update_weights
)

# Parse configuration

args = parse_args()  # Contains actor_num_nodes, actor_num_gpus_per_node, etc.

# Phase 1: Create and partition placement groups

pg_info = create_placement_groups(args)

# pg_info["actor"] -> PlacementGroupInfo for training

# pg_info["rollout"] -> PlacementGroupInfo for inference

# Phase 2: Launch rollout components in rollout placement group

inference_controller, rollout_executor, _ = await create_rollout_components(args)

# Phase 3: Launch training components in actor placement group

actor_model, critic_model = await create_training_models(
    args, inference_controller, rollout_executor
)

# Phase 4: Synchronize initial weights

await update_weights(actor_model, rollout_executor)

# Training loop proceeds with async rollouts and weight updates

```

## Key Source Files

| File | Responsibility |
|------|---------------|
| [`miles/ray/placement_group.py`](https://github.com/radixark/miles/blob/main/miles/ray/placement_group.py) | Placement group creation, layout calculation, bundle splitting, and component factory functions |
| [`miles/ray/rollout/inference_controller.py`](https://github.com/radixark/miles/blob/main/miles/ray/rollout/inference_controller.py) | SGLang engine management within rollout placement group |
| [`miles/ray/rollout/rollout_executor.py`](https://github.com/radixark/miles/blob/main/miles/ray/rollout/rollout_executor.py) | Ray actor for executing rollouts; receives weight updates |
| [`miles/ray/train/group.py`](https://github.com/radixark/miles/blob/main/miles/ray/train/group.py) | `TrainerController` implementation for actor placement group |
| [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py), [`train.py`](https://github.com/radixark/miles/blob/main/train.py), [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) | Entry points orchestrating placement group setup and training loop |

## Summary

- **Ray placement groups in Miles** provide deterministic GPU-to-component mapping for distributed RL-HF training
- **`_get_placement_group_layout`** computes bundle allocation based on CLI flags for colocation or external rollout
- **`_create_placement_group`** builds sorted, InfoActor-discovered bundles for reproducible execution
- **`create_placement_groups`** splits the unified placement group into actor (training) and rollout (inference) slices
- **Weight synchronization** via `update_weights` keeps inference engines current with the training policy
- **Flexible deployment** supports both single-node debugging (`--colocate`) and large-scale distributed execution (`--rollout_external`)

## Frequently Asked Questions

### How does Miles ensure deterministic GPU assignment across restarts?

Miles launches an `InfoActor` in each bundle to discover the physical `(node_ip, gpu_id)` pair, then sorts bundles by these values before assigning logical indices. This produces consistent ordering regardless of Ray's default scheduling, as implemented in `_create_placement_group` (lines 51-90).

### Can training and rollout run on completely separate machines?

Yes. When `--rollout_external` is specified, `_get_placement_group_layout` places rollout bundles after actor bundles in the placement group. If combined with Ray's node resource constraints, this maps to physically distinct machines. The single placement group spans nodes, with bundles allocated to specific node types via Ray's resource requirements.

### What happens when `--colocate` is enabled with fewer GPUs than actor plus rollout need?

The colocation mode uses `max(actor_gpus, rollout_gpus)` total GPUs, not the sum. Both components time-share the same physical GPUs. This requires careful scheduling to prevent memory exhaustion and is intended for development and debugging rather than production throughput.

### How does the critic model fit into the placement group structure?

When `--use_critic` is enabled, the critic `TrainerController` simply reuses the actor's `PlacementGroupInfo`. Both policy and value networks occupy the same GPU bundles, with Ray actor scheduling distinguishing their execution. See `create_placement_groups` return value handling (lines 108-124).