How LoRA and Multi-LoRA Training Work with SGLang Rollout Integration in Miles

Miles integrates LoRA and Multi-LoRA training with SGLang through a dual-mode architecture: single-adapter rollouts inject LoRA weights directly into HTTP payloads, while Multi-LoRA rollouts use an async producer-consumer pattern with per-adapter buffering, version tracking, and dynamic staleness filtering.

The radixark/miles repository implements parameter-efficient fine-tuning at scale by combining Low-Rank Adaptation (LoRA) with the SGLang inference engine. This integration enables high-throughput rollout generation during reinforcement learning training, supporting both single-adapter and concurrent multi-adapter scenarios with correct weight versioning and safe adapter lifecycle management.


Single-LoRA Rollout Mode

When lora_rollout_enabled(args) returns true, the system operates in single-adapter mode. Each generated sample carries routing information that directs SGLang to load the correct LoRA weights.

Payload Construction in generate()

In miles/rollout/sglang_rollout.py, the generate() method builds HTTP requests with LoRA-specific fields when a sample has an associated adapter:


# Inside miles/rollout/sglang_rollout.py → generate()

if sample.adapter is not None:
    # Resolve the current adapter from the controller cache

    adapter = await AdaptersCache().get(sample.adapter.name)
    if adapter is None:
        # Adapter has been deregistered → abort generation

        sample.status = Sample.Status.ABORTED
        return sample
    payload["lora_path"] = slot_lora_name(sample.adapter.slot)
    payload["rid"] = make_rid(sample.adapter.name)
    payload["extra_key"] = f"{sample.adapter.name}:v{adapter.version}"

Three critical fields enable correct routing:

  • lora_path – The filesystem path to the LoRA weights, derived from slot_lora_name(sample.adapter.slot)
  • rid – A routing-id generated by make_rid() that uniquely identifies the adapter across the distributed system
  • extra_key – A version string (name:v{version}) that prevents stale weight usage during adapter updates

The payload is dispatched via post(url, payload, headers) to the SGLang router, which forwards it to the appropriate engine instance.


Multi-LoRA Async Rollout Mode

For training multiple adapters concurrently, Miles uses AsyncMultiLoRAWorker in miles/rollout/multi_lora/async_rollout.py. This implements a producer-consumer pattern with per-adapter buffering and round-robin batch collection.

Producer Thread and Group Buffering

The worker spawns a background thread via run_loop() that continuously:

  1. Fetches samples from the data source
  2. Groups them via process_and_enqueue()
  3. Stamps each group with the current slot_version and registration_id
  4. Applies dynamic filtering
  5. Stores results in per-adapter GroupBuffer instances

# Pseudo-flow of process_and_enqueue():

process_group(group)  # Stamps version/registration

apply_dynamic_filter(group)  # Drops invalid groups

self.buffers[adapter_name].put(result)  # Buffer per adapter

Batch Collection with Round-Robin Selection

The trainer retrieves batches through collect_batch(), which calls worker.get_groups(). This method:

  • Iterates live adapters in round-robin order
  • Respects min_groups_per_dp_split per adapter
  • Enforces per-adapter rollout batch sizes
  • Discards groups from retired adapters or stale registrations

# Key invocation in training loop:

batch = await collect_batch(args, worker, snapshot)

Adapter Registry and Lifecycle Management

Central coordination happens through MultiLoRAController in miles/ray/multi_lora/controller.py, a Ray actor that maintains the canonical view of all adapters.

Registration Flow

  1. MultiLoRAController.register_adapter(name, config) delegates to MultiLoRABackend.resolve_adapter_config() for validation
  2. Backend records slot, version, and data_path for the adapter
  3. AdaptersCache (TTL singleton) provides fast lookup during rollout

Safe Retirement and Abort Handling

When adapters retire, MultiLoRABackend.abort_adapter_requests() cancels in-flight requests to prevent stale generation:


# In miles/ray/multi_lora/backend.py

async def abort_adapter_requests(self, adapter_name: str) -> None:
    prefix = f"{adapter_name}{RID_SEPARATOR}"
    urls = await self.worker_urls()
    await asyncio.gather(
        *(self.client.post(f"{url}/abort_request", 
           json={"rid": prefix, "prefix": True})
          for url in urls), return_exceptions=True
    )

The prefix=True parameter ensures scoped cancellation—only requests matching the adapter's routing-id prefix are aborted, leaving other adapters undisturbed even when sharing the same GPU slot.


Version Consistency and Staleness Filtering

Miles implements multiple mechanisms to ensure weight-version consistency:

  • Registration-scoped buffering: GroupBuffer.drop_foreign() discards groups from previous registrations when an adapter name is reused
  • Explicit version stamping: Each sample carries slot_version enabling post-hoc staleness detection
  • Controller snapshots: snapshot objects passed to collect_batch() capture the adapter registry state at batch assembly time

These defenses prevent the silent correctness bug where training proceeds on generations produced with outdated LoRA weights.


Configuration and Throughput Considerations

Key parameters controlling the integration:

Parameter Effect
args.rollout_batch_size Determines producer concurrency
args.sglang_server_concurrency Upper bound on GPU-side parallel generation
min_groups_per_dp_split Minimum groups per adapter per batch
empty_batch_timeout Maximum wait for incomplete batches

The architecture separates throughput-bound concerns (async producer, SGLang engine) from correctness-bound concerns (version tracking, staleness filtering, abort handling).


Summary

  • Single-LoRA rollouts inject lora_path, rid, and extra_key into SGLang HTTP payloads via generate() in sglang_rollout.py
  • Multi-LoRA rollouts use AsyncMultiLoRAWorker with per-adapter GroupBuffer storage and round-robin collect_batch() retrieval
  • Adapter lifecycle is managed by MultiLoRAController and MultiLoRABackend, with safe retirement via scoped request abort
  • Version safety is ensured through registration IDs, explicit version stamping, and dynamic staleness filtering
  • All components integrate with SGLang's native LoRA serving capability while adding Miles-specific correctness guarantees

Frequently Asked Questions

How does Miles prevent using stale LoRA weights during Multi-LoRA training?

Each generated group is stamped with the adapter's current slot_version and registration_id. The collect_batch() function in async_rollout.py filters out groups whose registration IDs don't match the current adapter state. Additionally, when an adapter is re-registered (e.g., after checkpoint reload), GroupBuffer.drop_foreign() automatically drops buffered groups from the previous registration.

What happens when an adapter is retired mid-rollout?

The MultiLoRABackend.abort_adapter_requests() method sends abort requests to all SGLang workers with a prefix matching the adapter's routing-id. This cancels in-flight generations for that specific adapter without affecting concurrent adapters, even when they share GPU slots. The backend then deregisters the adapter from the controller.

Why does single-LoRA mode use three routing fields (lora_path, rid, extra_key)?

These fields serve distinct purposes: lora_path tells SGLang which weight file to load; rid enables request routing and scoped aborts across the distributed system; extra_key encodes version information for debugging and verification. Together they support both functional correctness (loading right weights) and operational visibility (tracking which version produced each generation).

How does the async worker maintain throughput without blocking training?

AsyncMultiLoRAWorker runs a dedicated producer thread that operates independently of the training loop. It pre-generates and buffers prompt groups, allowing collect_batch() to retrieve ready data with minimal latency. The concurrency level is controlled by args.rollout_batch_size, while args.sglang_server_concurrency limits GPU-side parallelism to prevent resource exhaustion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →