How the Multi-LoRA Async Trainer Differs from the Standard Async Trainer in Miles
The multi-LoRA async trainer (train_multi_lora_async.py) adds dynamic LoRA-adapter lifecycle management on top of Miles' standard async training pipeline, enabling on-the-fly registration, retirement, and weight propagation of multiple adapters without stopping training.
The Miles reinforcement learning framework provides two distinct paths for asynchronous training. While train_async.py handles single-model training, train_multi_lora_async.py introduces a sophisticated controller-based architecture designed specifically for multi-adapter Low-Rank Adaptation (LoRA) workloads. This article examines the architectural differences between these two trainers based on the radixark/miles source code.
Core Architectural Difference: The MultiLoRAController
The most fundamental distinction lies in how training orchestration is handled. The standard async trainer keeps all logic within a single train coroutine. In contrast, the multi-LoRA async trainer instantiates a dedicated MultiLoRAController via create_multilora_controller() at startup:
# train_multi_lora_async.py lines 44-48
controller = create_multilora_controller(args)
controller.start.remote()
This Ray actor wraps both a backend service and an HTTP server, exposing APIs for adapter lifecycle management. The controller also logs its endpoint for external clients: http://{host}:{api_port} (lines 49-50).
Adapter Registration and Configuration
The standard trainer loads a single static model configuration. The multi-LoRA version instead parses multiple adapter specifications from YAML files:
# train_multi_lora_async.py lines 55-58
for adapter_run_yaml in args.multi_lora_adapters:
config = parse_adapter_run_yaml(Path(adapter_run_yaml))
await controller.register_adapter.remote(adapter_run_name, config)
The parse_adapter_run_yaml function from miles/utils/adapter_config.py converts each adapter's configuration into runtime objects that the controller manages independently.
The Dynamic Snapshot Loop vs. Simple Rollout Iteration
Training loop structure diverges dramatically between the two implementations.
Standard trainer (train_async.py):
# Simple finite rollout loop
for rollout_id in range(args.start_rollout_id, args.num_rollout):
rollout_data_curr_ref = await rollout_data_next_future
rollout_data_next_future = await eager_create_task(prepare_and_generate(rollout_id + 1))
await actor_model.train(rollout_id, rollout_data_curr_ref)
Multi-LoRA trainer (train_multi_lora_async.py lines 60-71):
# Infinite loop with adapter state monitoring
while True:
advisors_pending, rolledout, retired_adapters = (
get_multi_lora_controller().snapshot.remote()
)
if rolledout.all_adapters_cleaned_up:
await asyncio.sleep(args.multi_lora_idle_poll_s)
continue
The multi-LoRA loop queries controller snapshots to determine adapter states: pending, active, retiring, or cleaned up. When no adapters are present, the loop deliberately pauses using multi_lora_idle_poll_s rather than terminating.
Weight Reconciliation and Generation Guards
Before generating rollouts, the multi-LoRA trainer performs additional synchronization steps missing from the standard implementation:
# train_multi_lora_async.py lines 73-82
actor_model.reconcile_adapters()
await actor_model.update_weights.remote(advisors_pending) # Push pending/running
await controller.signal_weights_updated.remote()
# Secondary guard before generation
post_update = await get_multi_lora_controller().snapshot.remote()
if post_update.rolledout.active_adapters is None or len(post_update.rolledout.active_adapters) == 0:
continue # Skip generation if no active adapters remain
The reconcile_adapters() call ensures newly registered adapters are materialized, while update_weights propagates weights to the inference service. The post-update snapshot acts as a generation guard—if all adapters became inactive during weight propagation, the iteration skips rather than fails.
Specialized Error Handling
The multi-LoRA trainer catches a specific failure mode absent from standard training:
# train_multi_lora_async.py lines 84-91
except EmptyBatchTimeoutError as e:
if len(active_advisors) == 0:
logger.info(f"No trainable groups after update, retrying: {e}")
continue
EmptyBatchTimeoutError indicates no trainable adapter groups are available. Instead of aborting, the multi-LoRA trainer logs the condition and retries the reconcile/update cycle. This resilience accommodates adapters that may finish or retire between snapshots.
Saving and Shutdown Behavior
Both trainers use similar checkpointing mechanisms, but with key operational differences:
| Aspect | Standard Async Trainer | Multi-LoRA Async Trainer |
|---|---|---|
| Save trigger | save_interval, save_trigger_sentinel |
Per-adapter cadence via backend |
| Save call location | Configurable intervals | After each training step (lines 95-98) |
| Shutdown sequence | Dispose rollout executor, inference controller, models | Plus explicit controller.stop.remote() (lines 101-104) |
The additional controller shutdown ensures proper cleanup of the HTTP server and backend actor.
Launching the Multi-LoRA Trainer
Typical invocation requires adapter specifications and routing configuration:
python -m miles.main.train_multi_lora_async \
--multi-lora-n-adapters 3 \
--multi-lora-adapters adapter1.yaml adapter2.yaml adapter3.yaml \
--sglang-router-ip 10.0.0.1 \
--sglang-router-port 12345 \
[other standard args]
Runtime adapter registration remains possible through the controller's remote API:
config = parse_adapter_run_yaml(Path("/path/to/new_adapter.yaml"))
await controller.register_adapter.remote("new_adapter_name", config)
Summary
The multi-LoRA async trainer in Miles extends the standard async pipeline through four critical additions:
- MultiLoRAController — A Ray actor managing adapter lifecycles via HTTP and internal APIs
- Dynamic adapter registration — YAML-parsed configurations loaded at startup or runtime
- Snapshot-driven loop — Continuous polling of adapter states with idle pausing and generation guards
- Resilient weight propagation — Reconciliation, update, and retry logic for adapter availability
These capabilities enable training scenarios requiring multiple specialized LoRA adapters—such as domain-specific fine-tuning or ensemble approaches—without interrupting the underlying async RL flow.
Frequently Asked Questions
How does adapter registration work at runtime in the multi-LoRA trainer?
The controller exposes a register_adapter.remote(name, config) API that accepts parsed YAML configurations. New adapters enter a pending state, then transition through reconciliation and weight propagation before becoming active for generation and training.
What happens when all adapters finish training?
The snapshot loop detects all_adapters_cleaned_up and pauses execution using multi_lora_idle_poll_s (default configurable). The trainer remains alive, polling for new adapter registrations, rather than terminating.
Why does the multi-LoRA trainer catch EmptyBatchTimeoutError specifically?
This error signals that no adapter groups have pending training data after weight updates. Since adapters may asynchronously complete between snapshots, the trainer treats this as a transient condition and retries reconciliation rather than failing.
Can the standard async trainer be upgraded to support multi-LoRA?
Not without substantial modification. The standard trainer lacks the controller abstraction, snapshot protocol, and adapter-aware weight propagation. Migration requires implementing the MultiLoRAController integration pattern found in train_multi_lora_async.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →