How the InferenceController Coordinates Rollout and Training in Miles

The InferenceController serves as the central orchestrator in Miles, managing a ContextLock for exclusive weight updates, synchronizing model checkpoints between training and inference, and coordinating RolloutServer instances to enable continuous training-while-serving.

The InferenceController in the Miles reinforcement learning framework bridges the gap between model training pipelines and live inference services. Located in miles/ray/rollout/inference_controller.py, this component ensures that rollout servers always serve consistent model versions while allowing training to push updates without interruption.

Core Responsibilities of the InferenceController

The controller's architecture can be broken into five interconnected responsibilities that together enable seamless coordination between rollout and training operations.

Lock Management for Safe Weight Updates

The InferenceController maintains a ContextLock("InferenceController") that guarantees exclusive access during state-changing operations. This lock prevents race conditions when the controller swaps model weights or reconfigures rollout servers.

Training processes must acquire this same lock before replacing weights. Rollout servers, meanwhile, can safely read the current model knowing no partial updates are in progress. This mutual exclusion mechanism is critical for avoiding corrupted inference states.

Weight Synchronization Via Checksum Events

The controller receives InferenceEngineWeightChecksumEvent notifications from the training pipeline, defined in miles/utils/audit_utils/event_logger/models.py. These events contain:

  • model_name: Identifier for the model variant
  • checksum: Hash of the weight checkpoint for validation
  • step: Training step number for versioning

The controller stores the latest checksum map and validates that running rollout servers match expected values. When a checkpoint completes, training pushes a new checksum; the controller then propagates this to rollout servers, which reload and emit confirmation events.

Rollout Server Coordination

The InferenceController instantiates and owns one or more RolloutServer objects such as SFTRollout (miles/ray/rollout/sft_rollout.py) or SGLangRollout. Its coordination duties include:

  • Routing inference requests to appropriate servers
  • Monitoring server health and responsiveness
  • Restarting servers when new checkpoints are ready
  • Ensuring correct device placement for GPU allocation

This ownership model gives the controller complete lifecycle management over inference capabilities.

Health and Status Reporting

A unified /health endpoint aggregates status across all rollout servers and lock states. Training queries this endpoint to determine inference readiness before pushing checkpoints, eliminating race conditions between the two pipelines.

Runtime Configuration Management

The controller reads environment variables governing rollout behavior—ports, random seeds, gated launch flags—allowing dynamic adjustment without full process restarts. Training can modify rollout parameters through this configuration layer.

How Rollout and Training Stay Synchronized

The synchronization protocol follows a clear sequence grounded in the source code implementation:

  1. Training completes a checkpoint and computes its checksum
  2. Checksum event is logged via InferenceEngineWeightChecksumEvent
  3. Controller acquires the ContextLock and validates the new checksum
  4. Rollout servers receive update signal and reload weights
  5. Servers emit confirmation checksums back to the controller
  6. Lock releases and normal inference resumes

This handshake ensures no inference request observes partially-loaded weights. The InferenceController acts as the single source of truth for model versioning throughout.

Code Examples: Working with the InferenceController

Instantiating the Controller

from miles.ray.rollout.inference_controller import InferenceController

controller = InferenceController(
    rollout_server_cls="miles.rollout.sft_rollout.SFTRollout",
    model_checkpoint="/path/to/checkpoint",
    port=8000,
)

This creates a controller with an SFT rollout server bound to port 8000.

Logging a Weight Checksum from Training

from miles.utils.audit_utils.event_logger.models import InferenceEngineWeightChecksumEvent

event = InferenceEngineWeightChecksumEvent(
    model_name="gpt-xyz",
    checksum="abc123def456",
    step=42,
)
controller.event_logger.log(
    InferenceEngineWeightChecksumEvent, 
    event, 
    print_log=False
)

Training pipelines use this pattern to report checkpoint availability.

Triggering Safe Weight Updates

controller.update_rollout_weights(new_checkpoint="/new/checkpoint")

The update_rollout_weights method encapsulates lock acquisition, server coordination, and validation—all critical for maintaining serving consistency.

Key Source Files

Summary

  • The InferenceController is Miles' central hub for coordinating distributed inference with ongoing training
  • ContextLock("InferenceController") ensures atomic weight updates without serving interruptions
  • Checksum events provide verifiable, versioned communication between training and rollout
  • Direct ownership of RolloutServer instances enables complete lifecycle and health management
  • The /health endpoint exposes unified status for pipeline synchronization decisions

Frequently Asked Questions

What prevents race conditions when weights update during active inference?

The ContextLock("InferenceController") enforces exclusive access. Training acquires this lock before pushing checkpoints; the controller holds it during server reloads. Inference requests either complete against the old version or wait for the new one—never observing intermediate states.

How does training know when rollout servers are ready for a new checkpoint?

Training polls the controller's /health endpoint, which aggregates readiness across all RolloutServer instances. The status response indicates whether servers are actively serving, reloading, or in error states—enabling checkpoint timing decisions.

Can multiple rollout server types run under one InferenceController?

Yes. The rollout_server_cls parameter accepts server class specifications, and the controller architecture supports heterogeneous deployments. Different model variants or inference optimizations (SFT, SGLang) can coexist with unified coordination.

What happens if a checksum validation fails?

The controller detects mismatches between expected and reported checksums. Failed validations trigger logging events and can halt the update pipeline, preventing corrupted weights from reaching production inference paths.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →