# How the Multi-LoRA Async Trainer Differs from the Standard Async Trainer in Miles

> Discover how the Miles multi-LoRA async trainer enhances standard async training with dynamic adapter management for seamless on-the-fly registration, retirement, and weight propagation.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: deep-dive
- Published: 2026-09-06

---

**The multi-LoRA async trainer ([`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py)) adds dynamic LoRA-adapter lifecycle management on top of Miles' standard async training pipeline, enabling on-the-fly registration, retirement, and weight propagation of multiple adapters without stopping training.**

The Miles reinforcement learning framework provides two distinct paths for asynchronous training. While [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py) handles single-model training, [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) introduces a sophisticated controller-based architecture designed specifically for multi-adapter Low-Rank Adaptation (LoRA) workloads. This article examines the architectural differences between these two trainers based on the radixark/miles source code.

## Core Architectural Difference: The MultiLoRAController

The most fundamental distinction lies in how training orchestration is handled. The standard async trainer keeps all logic within a single `train` coroutine. In contrast, the multi-LoRA async trainer instantiates a dedicated **MultiLoRAController** via `create_multilora_controller()` at startup:

```python

# train_multi_lora_async.py lines 44-48

controller = create_multilora_controller(args)
controller.start.remote()

```

This Ray actor wraps both a backend service and an HTTP server, exposing APIs for adapter lifecycle management. The controller also logs its endpoint for external clients: `http://{host}:{api_port}` (lines 49-50).

## Adapter Registration and Configuration

The standard trainer loads a single static model configuration. The multi-LoRA version instead parses multiple adapter specifications from YAML files:

```python

# train_multi_lora_async.py lines 55-58

for adapter_run_yaml in args.multi_lora_adapters:
    config = parse_adapter_run_yaml(Path(adapter_run_yaml))
    await controller.register_adapter.remote(adapter_run_name, config)

```

The `parse_adapter_run_yaml` function from [`miles/utils/adapter_config.py`](https://github.com/radixark/miles/blob/main/miles/utils/adapter_config.py) converts each adapter's configuration into runtime objects that the controller manages independently.

## The Dynamic Snapshot Loop vs. Simple Rollout Iteration

Training loop structure diverges dramatically between the two implementations.

**Standard trainer ([`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py)):**

```python

# Simple finite rollout loop

for rollout_id in range(args.start_rollout_id, args.num_rollout):
    rollout_data_curr_ref = await rollout_data_next_future
    rollout_data_next_future = await eager_create_task(prepare_and_generate(rollout_id + 1))
    await actor_model.train(rollout_id, rollout_data_curr_ref)

```

**Multi-LoRA trainer ([`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) lines 60-71):**

```python

# Infinite loop with adapter state monitoring

while True:
    advisors_pending, rolledout, retired_adapters = (
        get_multi_lora_controller().snapshot.remote()
    )
    if rolledout.all_adapters_cleaned_up:
        await asyncio.sleep(args.multi_lora_idle_poll_s)
        continue

```

The multi-LoRA loop queries controller snapshots to determine adapter states: **pending**, **active**, **retiring**, or **cleaned up**. When no adapters are present, the loop deliberately pauses using `multi_lora_idle_poll_s` rather than terminating.

## Weight Reconciliation and Generation Guards

Before generating rollouts, the multi-LoRA trainer performs additional synchronization steps missing from the standard implementation:

```python

# train_multi_lora_async.py lines 73-82

actor_model.reconcile_adapters()
await actor_model.update_weights.remote(advisors_pending)  # Push pending/running

await controller.signal_weights_updated.remote()

# Secondary guard before generation

post_update = await get_multi_lora_controller().snapshot.remote()
if post_update.rolledout.active_adapters is None or len(post_update.rolledout.active_adapters) == 0:
    continue  # Skip generation if no active adapters remain

```

The `reconcile_adapters()` call ensures newly registered adapters are materialized, while `update_weights` propagates weights to the inference service. The post-update snapshot acts as a generation guard—if all adapters became inactive during weight propagation, the iteration skips rather than fails.

## Specialized Error Handling

The multi-LoRA trainer catches a specific failure mode absent from standard training:

```python

# train_multi_lora_async.py lines 84-91

except EmptyBatchTimeoutError as e:
    if len(active_advisors) == 0:
        logger.info(f"No trainable groups after update, retrying: {e}")
        continue

```

**EmptyBatchTimeoutError** indicates no trainable adapter groups are available. Instead of aborting, the multi-LoRA trainer logs the condition and retries the reconcile/update cycle. This resilience accommodates adapters that may finish or retire between snapshots.

## Saving and Shutdown Behavior

Both trainers use similar checkpointing mechanisms, but with key operational differences:

| Aspect | Standard Async Trainer | Multi-LoRA Async Trainer |
|--------|------------------------|--------------------------|
| Save trigger | `save_interval`, `save_trigger_sentinel` | Per-adapter cadence via backend |
| Save call location | Configurable intervals | After each training step (lines 95-98) |
| Shutdown sequence | Dispose rollout executor, inference controller, models | Plus explicit `controller.stop.remote()` (lines 101-104) |

The additional controller shutdown ensures proper cleanup of the HTTP server and backend actor.

## Launching the Multi-LoRA Trainer

Typical invocation requires adapter specifications and routing configuration:

```bash
python -m miles.main.train_multi_lora_async \
    --multi-lora-n-adapters 3 \
    --multi-lora-adapters adapter1.yaml adapter2.yaml adapter3.yaml \
    --sglang-router-ip 10.0.0.1 \
    --sglang-router-port 12345 \
    [other standard args]

```

Runtime adapter registration remains possible through the controller's remote API:

```python
config = parse_adapter_run_yaml(Path("/path/to/new_adapter.yaml"))
await controller.register_adapter.remote("new_adapter_name", config)

```

## Summary

The multi-LoRA async trainer in Miles extends the standard async pipeline through four critical additions:

- **MultiLoRAController** — A Ray actor managing adapter lifecycles via HTTP and internal APIs
- **Dynamic adapter registration** — YAML-parsed configurations loaded at startup or runtime
- **Snapshot-driven loop** — Continuous polling of adapter states with idle pausing and generation guards
- **Resilient weight propagation** — Reconciliation, update, and retry logic for adapter availability

These capabilities enable training scenarios requiring multiple specialized LoRA adapters—such as domain-specific fine-tuning or ensemble approaches—without interrupting the underlying async RL flow.

## Frequently Asked Questions

### How does adapter registration work at runtime in the multi-LoRA trainer?

The controller exposes a `register_adapter.remote(name, config)` API that accepts parsed YAML configurations. New adapters enter a **pending** state, then transition through reconciliation and weight propagation before becoming **active** for generation and training.

### What happens when all adapters finish training?

The snapshot loop detects `all_adapters_cleaned_up` and pauses execution using `multi_lora_idle_poll_s` (default configurable). The trainer remains alive, polling for new adapter registrations, rather than terminating.

### Why does the multi-LoRA trainer catch EmptyBatchTimeoutError specifically?

This error signals that no adapter groups have pending training data after weight updates. Since adapters may asynchronously complete between snapshots, the trainer treats this as a transient condition and retries reconciliation rather than failing.

### Can the standard async trainer be upgraded to support multi-LoRA?

Not without substantial modification. The standard trainer lacks the controller abstraction, snapshot protocol, and adapter-aware weight propagation. Migration requires implementing the `MultiLoRAController` integration pattern found in [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py).