# How the InferenceController Coordinates Rollout and Training in Miles

> Discover how the InferenceController orchestrates continuous training-while-serving in Miles. Learn about its role in exclusive weight updates, checkpoint synchronization, and RolloutServer coordination.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: internals
- Published: 2026-09-06

---

**The `InferenceController` serves as the central orchestrator in Miles, managing a `ContextLock` for exclusive weight updates, synchronizing model checkpoints between training and inference, and coordinating `RolloutServer` instances to enable continuous training-while-serving.**

The `InferenceController` in the Miles reinforcement learning framework bridges the gap between model training pipelines and live inference services. Located in [`miles/ray/rollout/inference_controller.py`](https://github.com/radixark/miles/blob/main/miles/ray/rollout/inference_controller.py), this component ensures that rollout servers always serve consistent model versions while allowing training to push updates without interruption.

## Core Responsibilities of the InferenceController

The controller's architecture can be broken into five interconnected responsibilities that together enable seamless coordination between rollout and training operations.

### Lock Management for Safe Weight Updates

The `InferenceController` maintains a **`ContextLock("InferenceController")`** that guarantees exclusive access during state-changing operations. This lock prevents race conditions when the controller swaps model weights or reconfigures rollout servers.

Training processes must acquire this same lock before replacing weights. Rollout servers, meanwhile, can safely read the current model knowing no partial updates are in progress. This mutual exclusion mechanism is critical for avoiding corrupted inference states.

### Weight Synchronization Via Checksum Events

The controller receives **`InferenceEngineWeightChecksumEvent`** notifications from the training pipeline, defined in [`miles/utils/audit_utils/event_logger/models.py`](https://github.com/radixark/miles/blob/main/miles/utils/audit_utils/event_logger/models.py). These events contain:

- `model_name`: Identifier for the model variant
- `checksum`: Hash of the weight checkpoint for validation
- `step`: Training step number for versioning

The controller stores the latest checksum map and validates that running rollout servers match expected values. When a checkpoint completes, training pushes a new checksum; the controller then propagates this to rollout servers, which reload and emit confirmation events.

### Rollout Server Coordination

The `InferenceController` **instantiates and owns** one or more `RolloutServer` objects such as `SFTRollout` ([`miles/ray/rollout/sft_rollout.py`](https://github.com/radixark/miles/blob/main/miles/ray/rollout/sft_rollout.py)) or `SGLangRollout`. Its coordination duties include:

- Routing inference requests to appropriate servers
- Monitoring server health and responsiveness
- Restarting servers when new checkpoints are ready
- Ensuring correct device placement for GPU allocation

This ownership model gives the controller complete lifecycle management over inference capabilities.

### Health and Status Reporting

A unified **`/health` endpoint** aggregates status across all rollout servers and lock states. Training queries this endpoint to determine inference readiness before pushing checkpoints, eliminating race conditions between the two pipelines.

### Runtime Configuration Management

The controller reads environment variables governing rollout behavior—ports, random seeds, gated launch flags—allowing dynamic adjustment without full process restarts. Training can modify rollout parameters through this configuration layer.

## How Rollout and Training Stay Synchronized

The synchronization protocol follows a clear sequence grounded in the source code implementation:

1. **Training completes a checkpoint** and computes its checksum
2. **Checksum event is logged** via `InferenceEngineWeightChecksumEvent`
3. **Controller acquires the `ContextLock`** and validates the new checksum
4. **Rollout servers receive update signal** and reload weights
5. **Servers emit confirmation checksums** back to the controller
6. **Lock releases** and normal inference resumes

This handshake ensures **no inference request observes partially-loaded weights**. The `InferenceController` acts as the single source of truth for model versioning throughout.

## Code Examples: Working with the InferenceController

### Instantiating the Controller

```python
from miles.ray.rollout.inference_controller import InferenceController

controller = InferenceController(
    rollout_server_cls="miles.rollout.sft_rollout.SFTRollout",
    model_checkpoint="/path/to/checkpoint",
    port=8000,
)

```

This creates a controller with an SFT rollout server bound to port 8000.

### Logging a Weight Checksum from Training

```python
from miles.utils.audit_utils.event_logger.models import InferenceEngineWeightChecksumEvent

event = InferenceEngineWeightChecksumEvent(
    model_name="gpt-xyz",
    checksum="abc123def456",
    step=42,
)
controller.event_logger.log(
    InferenceEngineWeightChecksumEvent, 
    event, 
    print_log=False
)

```

Training pipelines use this pattern to report checkpoint availability.

### Triggering Safe Weight Updates

```python
controller.update_rollout_weights(new_checkpoint="/new/checkpoint")

```

The `update_rollout_weights` method encapsulates lock acquisition, server coordination, and validation—all critical for maintaining serving consistency.

## Key Source Files

- **[`miles/ray/rollout/inference_controller.py`](https://github.com/radixark/miles/blob/main/miles/ray/rollout/inference_controller.py)** — Core implementation with lock handling and server orchestration
- **[`miles/ray/rollout/sft_rollout.py`](https://github.com/radixark/miles/blob/main/miles/ray/rollout/sft_rollout.py)** — Reference rollout server implementation
- **[`miles/utils/audit_utils/event_logger/models.py`](https://github.com/radixark/miles/blob/main/miles/utils/audit_utils/event_logger/models.py)** — Event definitions for weight synchronization
- **[`tests/fast/ray/rollout/test_rollout_server_wait_cells.py`](https://github.com/radixark/miles/blob/main/tests/fast/ray/rollout/test_rollout_server_wait_cells.py)** — Lock acquisition tests
- **[`tests/fast/ray/test_update_weights_ordering.py`](https://github.com/radixark/miles/blob/main/tests/fast/ray/test_update_weights_ordering.py)** — Weight update ordering verification

## Summary

- The **`InferenceController`** is Miles' central hub for coordinating distributed inference with ongoing training
- **`ContextLock("InferenceController")`** ensures atomic weight updates without serving interruptions
- **Checksum events** provide verifiable, versioned communication between training and rollout
- **Direct ownership of `RolloutServer` instances** enables complete lifecycle and health management
- The **`/health` endpoint** exposes unified status for pipeline synchronization decisions

## Frequently Asked Questions

### What prevents race conditions when weights update during active inference?

The `ContextLock("InferenceController")` enforces exclusive access. Training acquires this lock before pushing checkpoints; the controller holds it during server reloads. Inference requests either complete against the old version or wait for the new one—never observing intermediate states.

### How does training know when rollout servers are ready for a new checkpoint?

Training polls the controller's **`/health` endpoint**, which aggregates readiness across all `RolloutServer` instances. The status response indicates whether servers are actively serving, reloading, or in error states—enabling checkpoint timing decisions.

### Can multiple rollout server types run under one InferenceController?

Yes. The `rollout_server_cls` parameter accepts server class specifications, and the controller architecture supports heterogeneous deployments. Different model variants or inference optimizations (SFT, SGLang) can coexist with unified coordination.

### What happens if a checksum validation fails?

The controller detects mismatches between expected and reported checksums. Failed validations trigger logging events and can halt the update pipeline, preventing corrupted weights from reaching production inference paths.