How the Rollout Executor (Worker Manager) Connects to SGLang Engines in Miles

The RolloutExecutor is a Ray remote actor that stores the SGLang router address and sends all generation and evaluation requests via HTTP to the router, which forwards them to the underlying SGLang engine processes.

In the Miles reinforcement learning framework, the rollout executor—also referred to as the worker manager—serves as the critical bridge between the training loop and distributed SGLang inference engines. This component handles data generation, evaluation, and checkpointing by routing requests through a centralized SGLang router. Understanding this connection is essential for debugging, scaling, or modifying Miles inference pipelines.

Creating the Rollout Manager

The rollout infrastructure is initialized through a single factory function that orchestrates multiple Ray actors.

In miles/ray/placement_group.py, the create_rollout_components function spins up both the RolloutExecutor actor and the SGLang engine pool:


# train_async.py, line 33

inference_controller, rollout_executor, num_rollout_per_epoch = await create_rollout_components(args)

This call appears in the main training scripts:

The create_rollout_components function returns three values: an inference controller, the rollout executor handle, and the number of rollouts to generate per epoch.

RolloutExecutor Initialization and SGLang Router Connection

The RolloutExecutor class is decorated as a Ray remote actor (@ray.remote). Its constructor captures the full argument namespace and establishes the HTTP connection to the SGLang router.

From miles/ray/rollout/rollout_executor.py, lines 46-64:

class RolloutExecutor:
    def __init__(self, args: "Args"):
        self.args = args
        # Line 59-60: derive router address from CLI arguments

        router_addr = f"http://{args.sglang_router_ip}:{args.sglang_router_port}"
        
        # Line 62-64: initialize HTTP client unless in debug mode

        if not self.args.debug_train_only:
            init_http_client(args)  # from miles/utils/http_utils.py

            init_tracking(args, primary=False, router_addr=router_addr)

The router address construction pulls directly from args.sglang_router_ip and args.sglang_router_port. These values originate from command-line configuration and point to the SGLang router process that sits in front of the actual engine replicas.

Sending Rollout Requests to SGLang Engines

When the training loop needs fresh rollout data, it invokes the executor's get method:


# Inside training loop

rollout_data = await rollout_executor.get.remote(rollout_id)

The get method internally calls call_rollout_function (or legacy call_rollout_fn), which performs an HTTP POST to the router address established during initialization. The SGLang router then:

  1. Receives the request
  2. Selects an available SGLang engine process
  3. Forwards the inference call
  4. Returns generated tokens, KV-cache metadata, and logprobs

The executor post-processes this response into training-compatible tensors before returning to the trainer.

Evaluation Path Through the Same Router

Evaluation requests follow an identical pattern. The RolloutExecutor.eval method (also in miles/ray/rollout/rollout_executor.py) uses the same HTTP client to reach the router:


# Optional checkpoint evaluation

await rollout_executor.eval.remote(rollout_id, hf_dir=hf_path)

For checkpoint evaluation, Miles can optionally spin up a separate "checkpoint-eval fleet" of SGLang engines, but the communication mechanism remains unchanged—HTTP requests routed through the same address.

Lifecycle Management and Cleanup

The create_rollout_components factory registers the executor with the router so the router maintains awareness of active workloads. When training completes, proper shutdown is critical to release GPU resources:


# train_async.py, line 137

await rollout_executor.dispose.remote()

This dispose call terminates the SGLang engine processes and cleans up the Ray actor.

Complete Usage Pattern


# 1. Initialize rollout infrastructure

inference_ctl, rollout_executor, n_per_epoch = await create_rollout_components(args)

# 2. Generate rollouts during training

for rollout_id in range(n_per_epoch):
    data = await rollout_executor.get.remote(rollout_id)
    # ... training step ...

# 3. Optional evaluation

await rollout_executor.eval.remote(checkpoint_id, hf_dir="/path/to/checkpoint")

# 4. Shutdown

await rollout_executor.dispose.remote()

Key Source Files

File Purpose
miles/ray/placement_group.py create_rollout_components() factory; orchestrates Ray actors and SGLang engines
miles/ray/rollout/rollout_executor.py RolloutExecutor class; Ray remote actor that manages HTTP communication with SGLang router
train_async.py, train.py, train_multi_lora_async.py Entry points demonstrating executor usage
miles/utils/http_utils.py init_http_client() and request helpers
miles/backends/sglang_utils/sglang_engine.py SGLang server process management (implicit in router setup)

Summary

  • The RolloutExecutor is a Ray remote actor that acts as the worker manager for Miles inference operations.
  • It stores the SGLang router address (sglang_router_ip:sglang_router_port) derived from CLI arguments.
  • All generation and evaluation requests are sent as HTTP RPCs to this router, which load-balances across SGLang engine processes.
  • The executor lifecycle is managed through create_rollout_components() initialization and dispose.remote() cleanup.
  • This architecture decouples the training loop from engine scaling, allowing dynamic adjustment of inference capacity.

Frequently Asked Questions

How does the RolloutExecutor discover the SGLang router address?

The router address is passed explicitly through the args namespace. At initialization, the executor constructs the full URL as f"http://{args.sglang_router_ip}:{args.sglang_router_port}" and passes it to init_tracking() in miles/ray/rollout/rollout_executor.py, lines 59-60.

What happens if debug_train_only is set to True?

When args.debug_train_only is enabled, the executor skips init_http_client() entirely (line 62-64). This mode allows training logic verification without launching actual SGLang engines or making inference requests.

Can multiple RolloutExecutors connect to the same SGLang router?

Yes. The SGLang router is designed as a shared entry point that can accept connections from multiple clients. However, Miles typically creates one RolloutExecutor per training job through create_rollout_components(), which coordinates the full fleet.

Where is the HTTP client implementation located?

The HTTP client initialization lives in miles/utils/http_utils.py via init_http_client(). The actual request logic for rollout calls is implemented in the call_rollout_function helper, which performs POST requests to the stored router address.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →