How LoRA and Multi-LoRA Training Work with SGLang Rollout Integration in Miles
Miles integrates LoRA and Multi-LoRA training with SGLang through a dual-mode architecture: single-adapter rollouts inject LoRA weights directly into HTTP payloads, while Multi-LoRA rollouts use an async producer-consumer pattern with per-adapter buffering, version tracking, and dynamic staleness filtering.
The radixark/miles repository implements parameter-efficient fine-tuning at scale by combining Low-Rank Adaptation (LoRA) with the SGLang inference engine. This integration enables high-throughput rollout generation during reinforcement learning training, supporting both single-adapter and concurrent multi-adapter scenarios with correct weight versioning and safe adapter lifecycle management.
Single-LoRA Rollout Mode
When lora_rollout_enabled(args) returns true, the system operates in single-adapter mode. Each generated sample carries routing information that directs SGLang to load the correct LoRA weights.
Payload Construction in generate()
In miles/rollout/sglang_rollout.py, the generate() method builds HTTP requests with LoRA-specific fields when a sample has an associated adapter:
# Inside miles/rollout/sglang_rollout.py → generate()
if sample.adapter is not None:
# Resolve the current adapter from the controller cache
adapter = await AdaptersCache().get(sample.adapter.name)
if adapter is None:
# Adapter has been deregistered → abort generation
sample.status = Sample.Status.ABORTED
return sample
payload["lora_path"] = slot_lora_name(sample.adapter.slot)
payload["rid"] = make_rid(sample.adapter.name)
payload["extra_key"] = f"{sample.adapter.name}:v{adapter.version}"
Three critical fields enable correct routing:
lora_path– The filesystem path to the LoRA weights, derived fromslot_lora_name(sample.adapter.slot)rid– A routing-id generated bymake_rid()that uniquely identifies the adapter across the distributed systemextra_key– A version string (name:v{version}) that prevents stale weight usage during adapter updates
The payload is dispatched via post(url, payload, headers) to the SGLang router, which forwards it to the appropriate engine instance.
Multi-LoRA Async Rollout Mode
For training multiple adapters concurrently, Miles uses AsyncMultiLoRAWorker in miles/rollout/multi_lora/async_rollout.py. This implements a producer-consumer pattern with per-adapter buffering and round-robin batch collection.
Producer Thread and Group Buffering
The worker spawns a background thread via run_loop() that continuously:
- Fetches samples from the data source
- Groups them via
process_and_enqueue() - Stamps each group with the current
slot_versionandregistration_id - Applies dynamic filtering
- Stores results in per-adapter
GroupBufferinstances
# Pseudo-flow of process_and_enqueue():
process_group(group) # Stamps version/registration
apply_dynamic_filter(group) # Drops invalid groups
self.buffers[adapter_name].put(result) # Buffer per adapter
Batch Collection with Round-Robin Selection
The trainer retrieves batches through collect_batch(), which calls worker.get_groups(). This method:
- Iterates live adapters in round-robin order
- Respects
min_groups_per_dp_splitper adapter - Enforces per-adapter rollout batch sizes
- Discards groups from retired adapters or stale registrations
# Key invocation in training loop:
batch = await collect_batch(args, worker, snapshot)
Adapter Registry and Lifecycle Management
Central coordination happens through MultiLoRAController in miles/ray/multi_lora/controller.py, a Ray actor that maintains the canonical view of all adapters.
Registration Flow
MultiLoRAController.register_adapter(name, config)delegates toMultiLoRABackend.resolve_adapter_config()for validation- Backend records
slot,version, anddata_pathfor the adapter AdaptersCache(TTL singleton) provides fast lookup during rollout
Safe Retirement and Abort Handling
When adapters retire, MultiLoRABackend.abort_adapter_requests() cancels in-flight requests to prevent stale generation:
# In miles/ray/multi_lora/backend.py
async def abort_adapter_requests(self, adapter_name: str) -> None:
prefix = f"{adapter_name}{RID_SEPARATOR}"
urls = await self.worker_urls()
await asyncio.gather(
*(self.client.post(f"{url}/abort_request",
json={"rid": prefix, "prefix": True})
for url in urls), return_exceptions=True
)
The prefix=True parameter ensures scoped cancellation—only requests matching the adapter's routing-id prefix are aborted, leaving other adapters undisturbed even when sharing the same GPU slot.
Version Consistency and Staleness Filtering
Miles implements multiple mechanisms to ensure weight-version consistency:
- Registration-scoped buffering:
GroupBuffer.drop_foreign()discards groups from previous registrations when an adapter name is reused - Explicit version stamping: Each sample carries
slot_versionenabling post-hoc staleness detection - Controller snapshots:
snapshotobjects passed tocollect_batch()capture the adapter registry state at batch assembly time
These defenses prevent the silent correctness bug where training proceeds on generations produced with outdated LoRA weights.
Configuration and Throughput Considerations
Key parameters controlling the integration:
| Parameter | Effect |
|---|---|
args.rollout_batch_size |
Determines producer concurrency |
args.sglang_server_concurrency |
Upper bound on GPU-side parallel generation |
min_groups_per_dp_split |
Minimum groups per adapter per batch |
empty_batch_timeout |
Maximum wait for incomplete batches |
The architecture separates throughput-bound concerns (async producer, SGLang engine) from correctness-bound concerns (version tracking, staleness filtering, abort handling).
Summary
- Single-LoRA rollouts inject
lora_path,rid, andextra_keyinto SGLang HTTP payloads viagenerate()insglang_rollout.py - Multi-LoRA rollouts use
AsyncMultiLoRAWorkerwith per-adapterGroupBufferstorage and round-robincollect_batch()retrieval - Adapter lifecycle is managed by
MultiLoRAControllerandMultiLoRABackend, with safe retirement via scoped request abort - Version safety is ensured through registration IDs, explicit version stamping, and dynamic staleness filtering
- All components integrate with SGLang's native LoRA serving capability while adding Miles-specific correctness guarantees
Frequently Asked Questions
How does Miles prevent using stale LoRA weights during Multi-LoRA training?
Each generated group is stamped with the adapter's current slot_version and registration_id. The collect_batch() function in async_rollout.py filters out groups whose registration IDs don't match the current adapter state. Additionally, when an adapter is re-registered (e.g., after checkpoint reload), GroupBuffer.drop_foreign() automatically drops buffered groups from the previous registration.
What happens when an adapter is retired mid-rollout?
The MultiLoRABackend.abort_adapter_requests() method sends abort requests to all SGLang workers with a prefix matching the adapter's routing-id. This cancels in-flight generations for that specific adapter without affecting concurrent adapters, even when they share GPU slots. The backend then deregisters the adapter from the controller.
Why does single-LoRA mode use three routing fields (lora_path, rid, extra_key)?
These fields serve distinct purposes: lora_path tells SGLang which weight file to load; rid enables request routing and scoped aborts across the distributed system; extra_key encodes version information for debugging and verification. Together they support both functional correctness (loading right weights) and operational visibility (tracking which version produced each generation).
How does the async worker maintain throughput without blocking training?
AsyncMultiLoRAWorker runs a dedicated producer thread that operates independently of the training loop. It pre-generates and buffers prompt groups, allowing collect_batch() to retrieve ready data with minimal latency. The concurrency level is controlled by args.rollout_batch_size, while args.sglang_server_concurrency limits GPU-side parallelism to prevent resource exhaustion.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →