Asynchronous Rollout and Evaluation Modes in Miles: Beyond the Default Pipeline
Miles implements three distinct asynchronous strategies—async rollout, fully‑async rollout, and multi‑LoRA async—that overlap or decouple rollout generation from training to reduce wall‑time and enable parallel evaluation.
The radixark/miles repository provides multiple execution strategies for reinforcement learning training pipelines. While the default synchronous mode is simplest, the asynchronous rollout and evaluation modes offer significant performance improvements for large‑scale workloads. This article examines each mode's architecture, implementation files, and activation methods based on the actual source code.
Synchronous Rollout: The Default Baseline
The synchronous rollout mode is the classic blocking pipeline. The training loop waits entirely for the rollout manager to generate a full batch of samples before proceeding.
In train.py, the driver creates rollout components via create_rollout_components, prepares a rollout with inference_controller.prepare_rollout, and blocks on rollout_executor.get.remote to retrieve data. Training occurs only after this call returns, then the cycle repeats.
Because rollout and training execute sequentially, wall‑time per iteration equals rollout_time + train_time. This mode requires no special flags:
python -m miles.main.scripts.run_qwen3_5_0.5B_gsm8k_short.py \
--model qwen3-5b \
--num-rollout 1000
Key implementation files: train.py, miles/rollout/base_types.py, miles/rollout/data_source.py.
Async Rollout: Overlapping Generation and Training
Async rollout eliminates the blocking bottleneck by starting the next rollout while the current batch is being trained. This mode is activated by running train_async.py without the --fully-async flag.
The driver uses eager_create_task from miles/utils/async_utils.py to spawn a future for the upcoming rollout immediately before entering the training step:
rollout_data_next_future = await eager_create_task(prepare_and_generate(args.start_rollout_id))
for rollout_id in range(args.start_rollout_id, args.num_rollout):
if rollout_data_next_future is not None:
rollout_data_curr_ref = await rollout_data_next_future # consume previous rollout
# start the next rollout early
if rollout_id + 1 < args.num_rollout:
rollout_data_next_future = await eager_create_task(
prepare_and_generate(rollout_id + 1))
# ... training on rollout_data_curr_ref ...
This pattern reduces per‑iteration wall‑time to approximately max(rollout_time, train_time). Evaluation runs concurrently through EvalDispatcher in miles/rollout/eval_dispatch.py, rather than pausing training.
Activate this mode with:
python -m miles.main.train_async \
--model qwen3-5b \
--num-rollout 1000 \
--eval-interval 100
Key implementation files: train_async.py, miles/utils/async_utils.py, miles/rollout/fully_async_rollout.py (async API reference), miles/rollout/eval_dispatch.py.
Fully‑Async Rollout: Persistent Background Production
Fully‑async rollout decouples rollout production from consumption entirely using a persistent background worker. This mode requires the class‑based rollout API and the --fully-async flag.
The FullyAsyncRolloutFn class in miles/rollout/fully_async_rollout.py maintains a _worker_loop that continuously submits rollout groups up to async_max_concurrent_samples. Finished groups land in a DefaultDataBuffer (or user‑provided implementation) from miles/rollout/fully_async_data_buffer.py. Training pulls groups via _next_group, which blocks until data is available or warns after NO_PROGRESS_WARN_SECS.
Key architectural components:
- Submission scheduler (
make_submission_schedulerinmiles/rollout/submission_scheduler.py): Tracks in‑flight groups to enforce capacity limits - Data buffer: FIFO storage with staleness detection; aborted groups can be recycled via
--async-unused-samples-handler - Evaluation handling: The producer pauses (
_producer_resumed.clear()) so the same engines can execute a blocking eval pass, or evaluation runs on a dedicatedGenerateState
Launch fully‑async mode with buffer size control:
python -m miles.main.train_async \
--model qwen3-5b \
--num-rollout 1000 \
--fully-async \
--async-max-concurrent-samples 256
Key implementation files: miles/rollout/fully_async_rollout.py, miles/rollout/fully_async_data_buffer.py, miles/rollout/submission_scheduler.py.
Multi‑LoRA Async: Parallel Adapter Training
Multi‑LoRA async extends the async pattern to train multiple LoRA adapters simultaneously without multiplying rollout costs. The driver in train_multi_lora_async.py creates rollout components once, then repeatedly calls update_weights for each LoRA model while re‑using the same rollout data.
This specialized async variant leverages miles/ray/placement_group.py for resource allocation and miles/utils/data.py for data handling. It is ideal when scaling to dozens or hundreds of LoRA adapters.
Run multi‑LoRA training with:
python -m miles.main.train_multi_lora_async \
--model qwen3-5b \
--num-rollout 500 \
--lora-count 8
Key implementation files: train_multi_lora_async.py, miles/ray/placement_group.py, miles/utils/data.py.
Choosing Between Asynchronous Rollout and Evaluation Modes
| Mode | Wall‑time characteristic | Best for | Activation |
|---|---|---|---|
| Synchronous | rollout_time + train_time |
Simplicity, debugging, small scale | train.py (default) |
| Async | max(rollout_time, train_time) |
Moderate scale, first async speedup | train_async.py |
| Fully‑async | Training never waits for rollout | Maximum throughput, continuous training | train_async.py --fully-async |
| Multi‑LoRA async | Shared rollout cost across adapters | Many LoRA experiments in parallel | train_multi_lora_async.py |
Summary
Miles provides four distinct rollout and evaluation execution strategies:
- Synchronous rollout in
train.pyoffers the simplest blocking pipeline - Async rollout in
train_async.pyoverlaps generation and training usingeager_create_task - Fully‑async rollout with
--fully-asyncdecouples production viaFullyAsyncRolloutFnandDefaultDataBuffer - Multi‑LoRA async in
train_multi_lora_async.pytrains multiple adapters on shared rollout data
All modes preserve identical model semantics while trading implementation complexity for wall‑time efficiency.
Frequently Asked Questions
What is the performance difference between async and fully‑async rollout in Miles?
Async rollout reduces per‑iteration time from rollout_time + train_time to max(rollout_time, train_time) by overlapping adjacent rollouts. Fully‑async rollout eliminates this dependency entirely—a persistent background worker ensures training never blocks on generation, yielding the lowest possible wall‑time when rollout is slower than training.
How does evaluation work during asynchronous training?
In standard async mode, EvalDispatcher from miles/rollout/eval_dispatch.py runs evaluation concurrently with training. In fully‑async mode, you can either pause the producer to reuse engines for a blocking eval pass, or maintain a separate GenerateState for dedicated evaluation workers.
Can I use fully‑async rollout with custom data handling?
Yes. The FullyAsyncRolloutFn constructor accepts a user‑provided data buffer implementation. The default DefaultDataBuffer in miles/rollout/fully_async_data_buffer.py handles FIFO ordering, staleness detection, and optional sample recycling, but you can substitute your own buffer for specialized requirements.
When should I use multi‑LoRA async instead of standard async?
Use multi‑LoRA async when training multiple LoRA adapters that share the same base model. Because rollouts are generated once and reused across all adapters, you avoid the linear cost growth that would occur from running separate training jobs. The train_multi_lora_async.py driver manages weight updates and resource placement automatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →