# Asynchronous Rollout and Evaluation Modes in Miles: Beyond the Default Pipeline

> Discover Miles' async rollout strategies: async rollout, fully-async rollout, and multi-LoRA async. Decouple generation from training to cut wall-time and enable parallel evaluation.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: deep-dive
- Published: 2026-09-06

---

**Miles implements three distinct asynchronous strategies—async rollout, fully‑async rollout, and multi‑LoRA async—that overlap or decouple rollout generation from training to reduce wall‑time and enable parallel evaluation.**

The [radixark/miles](https://github.com/radixark/miles) repository provides multiple execution strategies for reinforcement learning training pipelines. While the default synchronous mode is simplest, the asynchronous rollout and evaluation modes offer significant performance improvements for large‑scale workloads. This article examines each mode's architecture, implementation files, and activation methods based on the actual source code.

---

## Synchronous Rollout: The Default Baseline

The **synchronous rollout** mode is the classic blocking pipeline. The training loop waits entirely for the rollout manager to generate a full batch of samples before proceeding.

In [`train.py`](https://github.com/radixark/miles/blob/main/train.py), the driver creates rollout components via `create_rollout_components`, prepares a rollout with `inference_controller.prepare_rollout`, and blocks on `rollout_executor.get.remote` to retrieve data. Training occurs only after this call returns, then the cycle repeats.

Because rollout and training execute sequentially, wall‑time per iteration equals `rollout_time + train_time`. This mode requires no special flags:

```bash
python -m miles.main.scripts.run_qwen3_5_0.5B_gsm8k_short.py \
    --model qwen3-5b \
    --num-rollout 1000

```

Key implementation files: [`train.py`](https://github.com/radixark/miles/blob/main/train.py), [`miles/rollout/base_types.py`](https://github.com/radixark/miles/blob/main/miles/rollout/base_types.py), [`miles/rollout/data_source.py`](https://github.com/radixark/miles/blob/main/miles/rollout/data_source.py).

---

## Async Rollout: Overlapping Generation and Training

**Async rollout** eliminates the blocking bottleneck by starting the next rollout **while** the current batch is being trained. This mode is activated by running [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py) without the `--fully-async` flag.

The driver uses `eager_create_task` from [`miles/utils/async_utils.py`](https://github.com/radixark/miles/blob/main/miles/utils/async_utils.py) to spawn a future for the upcoming rollout immediately before entering the training step:

```python
rollout_data_next_future = await eager_create_task(prepare_and_generate(args.start_rollout_id))
for rollout_id in range(args.start_rollout_id, args.num_rollout):
    if rollout_data_next_future is not None:
        rollout_data_curr_ref = await rollout_data_next_future   # consume previous rollout

    
    # start the next rollout early

    if rollout_id + 1 < args.num_rollout:
        rollout_data_next_future = await eager_create_task(
            prepare_and_generate(rollout_id + 1))
    # ... training on rollout_data_curr_ref ...

```

This pattern reduces per‑iteration wall‑time to approximately `max(rollout_time, train_time)`. Evaluation runs concurrently through `EvalDispatcher` in [`miles/rollout/eval_dispatch.py`](https://github.com/radixark/miles/blob/main/miles/rollout/eval_dispatch.py), rather than pausing training.

Activate this mode with:

```bash
python -m miles.main.train_async \
    --model qwen3-5b \
    --num-rollout 1000 \
    --eval-interval 100

```

Key implementation files: [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py), [`miles/utils/async_utils.py`](https://github.com/radixark/miles/blob/main/miles/utils/async_utils.py), [`miles/rollout/fully_async_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/fully_async_rollout.py) (async API reference), [`miles/rollout/eval_dispatch.py`](https://github.com/radixark/miles/blob/main/miles/rollout/eval_dispatch.py).

---

## Fully‑Async Rollout: Persistent Background Production

**Fully‑async rollout** decouples rollout production from consumption entirely using a persistent background worker. This mode requires the class‑based rollout API and the `--fully-async` flag.

The `FullyAsyncRolloutFn` class in [`miles/rollout/fully_async_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/fully_async_rollout.py) maintains a `_worker_loop` that continuously submits rollout groups up to `async_max_concurrent_samples`. Finished groups land in a `DefaultDataBuffer` (or user‑provided implementation) from [`miles/rollout/fully_async_data_buffer.py`](https://github.com/radixark/miles/blob/main/miles/rollout/fully_async_data_buffer.py). Training pulls groups via `_next_group`, which blocks until data is available or warns after `NO_PROGRESS_WARN_SECS`.

Key architectural components:

- **Submission scheduler** (`make_submission_scheduler` in [`miles/rollout/submission_scheduler.py`](https://github.com/radixark/miles/blob/main/miles/rollout/submission_scheduler.py)): Tracks in‑flight groups to enforce capacity limits
- **Data buffer**: FIFO storage with staleness detection; aborted groups can be recycled via `--async-unused-samples-handler`
- **Evaluation handling**: The producer pauses (`_producer_resumed.clear()`) so the same engines can execute a blocking eval pass, or evaluation runs on a dedicated `GenerateState`

Launch fully‑async mode with buffer size control:

```bash
python -m miles.main.train_async \
    --model qwen3-5b \
    --num-rollout 1000 \
    --fully-async \
    --async-max-concurrent-samples 256

```

Key implementation files: [`miles/rollout/fully_async_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/fully_async_rollout.py), [`miles/rollout/fully_async_data_buffer.py`](https://github.com/radixark/miles/blob/main/miles/rollout/fully_async_data_buffer.py), [`miles/rollout/submission_scheduler.py`](https://github.com/radixark/miles/blob/main/miles/rollout/submission_scheduler.py).

---

## Multi‑LoRA Async: Parallel Adapter Training

**Multi‑LoRA async** extends the async pattern to train **multiple LoRA adapters simultaneously** without multiplying rollout costs. The driver in [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) creates rollout components once, then repeatedly calls `update_weights` for each LoRA model while re‑using the same rollout data.

This specialized async variant leverages [`miles/ray/placement_group.py`](https://github.com/radixark/miles/blob/main/miles/ray/placement_group.py) for resource allocation and [`miles/utils/data.py`](https://github.com/radixark/miles/blob/main/miles/utils/data.py) for data handling. It is ideal when scaling to dozens or hundreds of LoRA adapters.

Run multi‑LoRA training with:

```bash
python -m miles.main.train_multi_lora_async \
    --model qwen3-5b \
    --num-rollout 500 \
    --lora-count 8

```

Key implementation files: [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py), [`miles/ray/placement_group.py`](https://github.com/radixark/miles/blob/main/miles/ray/placement_group.py), [`miles/utils/data.py`](https://github.com/radixark/miles/blob/main/miles/utils/data.py).

---

## Choosing Between Asynchronous Rollout and Evaluation Modes

| Mode | Wall‑time characteristic | Best for | Activation |
|------|-------------------------|----------|------------|
| Synchronous | `rollout_time + train_time` | Simplicity, debugging, small scale | [`train.py`](https://github.com/radixark/miles/blob/main/train.py) (default) |
| Async | `max(rollout_time, train_time)` | Moderate scale, first async speedup | [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py) |
| Fully‑async | Training never waits for rollout | Maximum throughput, continuous training | `train_async.py --fully-async` |
| Multi‑LoRA async | Shared rollout cost across adapters | Many LoRA experiments in parallel | [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) |

---

## Summary

Miles provides four distinct rollout and evaluation execution strategies:

- **Synchronous rollout** in [`train.py`](https://github.com/radixark/miles/blob/main/train.py) offers the simplest blocking pipeline
- **Async rollout** in [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py) overlaps generation and training using `eager_create_task`
- **Fully‑async rollout** with `--fully-async` decouples production via `FullyAsyncRolloutFn` and `DefaultDataBuffer`
- **Multi‑LoRA async** in [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) trains multiple adapters on shared rollout data

All modes preserve identical model semantics while trading implementation complexity for wall‑time efficiency.

---

## Frequently Asked Questions

### What is the performance difference between async and fully‑async rollout in Miles?

Async rollout reduces per‑iteration time from `rollout_time + train_time` to `max(rollout_time, train_time)` by overlapping adjacent rollouts. Fully‑async rollout eliminates this dependency entirely—a persistent background worker ensures training never blocks on generation, yielding the lowest possible wall‑time when rollout is slower than training.

### How does evaluation work during asynchronous training?

In standard async mode, `EvalDispatcher` from [`miles/rollout/eval_dispatch.py`](https://github.com/radixark/miles/blob/main/miles/rollout/eval_dispatch.py) runs evaluation concurrently with training. In fully‑async mode, you can either pause the producer to reuse engines for a blocking eval pass, or maintain a separate `GenerateState` for dedicated evaluation workers.

### Can I use fully‑async rollout with custom data handling?

Yes. The `FullyAsyncRolloutFn` constructor accepts a user‑provided data buffer implementation. The default `DefaultDataBuffer` in [`miles/rollout/fully_async_data_buffer.py`](https://github.com/radixark/miles/blob/main/miles/rollout/fully_async_data_buffer.py) handles FIFO ordering, staleness detection, and optional sample recycling, but you can substitute your own buffer for specialized requirements.

### When should I use multi‑LoRA async instead of standard async?

Use multi‑LoRA async when training **multiple LoRA adapters** that share the same base model. Because rollouts are generated once and reused across all adapters, you avoid the linear cost growth that would occur from running separate training jobs. The [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) driver manages weight updates and resource placement automatically.