# How LoRA and Multi-LoRA Training Work with SGLang Rollout Integration in Miles

> Discover how Miles integrates LoRA and Multi-LoRA training with SGLang. Learn about dual-mode architecture, async producer-consumer patterns, and dynamic staleness filtering for efficient rollouts.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: deep-dive
- Published: 2026-09-06

---

**Miles integrates LoRA and Multi-LoRA training with SGLang through a dual-mode architecture: single-adapter rollouts inject LoRA weights directly into HTTP payloads, while Multi-LoRA rollouts use an async producer-consumer pattern with per-adapter buffering, version tracking, and dynamic staleness filtering.**

The `radixark/miles` repository implements parameter-efficient fine-tuning at scale by combining **Low-Rank Adaptation (LoRA)** with the **[SGLang](https://github.com/sgl-project/sglang) inference engine**. This integration enables high-throughput rollout generation during reinforcement learning training, supporting both single-adapter and concurrent multi-adapter scenarios with correct weight versioning and safe adapter lifecycle management.

---

## Single-LoRA Rollout Mode

When `lora_rollout_enabled(args)` returns true, the system operates in **single-adapter mode**. Each generated sample carries routing information that directs SGLang to load the correct LoRA weights.

### Payload Construction in `generate()`

In [`miles/rollout/sglang_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/sglang_rollout.py), the `generate()` method builds HTTP requests with LoRA-specific fields when a sample has an associated adapter:

```python

# Inside miles/rollout/sglang_rollout.py → generate()

if sample.adapter is not None:
    # Resolve the current adapter from the controller cache

    adapter = await AdaptersCache().get(sample.adapter.name)
    if adapter is None:
        # Adapter has been deregistered → abort generation

        sample.status = Sample.Status.ABORTED
        return sample
    payload["lora_path"] = slot_lora_name(sample.adapter.slot)
    payload["rid"] = make_rid(sample.adapter.name)
    payload["extra_key"] = f"{sample.adapter.name}:v{adapter.version}"

```

Three critical fields enable correct routing:

- **`lora_path`** – The filesystem path to the LoRA weights, derived from `slot_lora_name(sample.adapter.slot)`
- **`rid`** – A **routing-id** generated by `make_rid()` that uniquely identifies the adapter across the distributed system
- **`extra_key`** – A version string (`name:v{version}`) that prevents stale weight usage during adapter updates

The payload is dispatched via `post(url, payload, headers)` to the SGLang router, which forwards it to the appropriate engine instance.

---

## Multi-LoRA Async Rollout Mode

For training multiple adapters concurrently, Miles uses `AsyncMultiLoRAWorker` in [`miles/rollout/multi_lora/async_rollout.py`](https://github.com/radixark/miles/blob/main/miles/rollout/multi_lora/async_rollout.py). This implements a **producer-consumer pattern** with per-adapter buffering and round-robin batch collection.

### Producer Thread and Group Buffering

The worker spawns a background thread via `run_loop()` that continuously:

1. Fetches samples from the data source
2. Groups them via `process_and_enqueue()`
3. Stamps each group with the current `slot_version` and `registration_id`
4. Applies dynamic filtering
5. Stores results in per-adapter `GroupBuffer` instances

```python

# Pseudo-flow of process_and_enqueue():

process_group(group)  # Stamps version/registration

apply_dynamic_filter(group)  # Drops invalid groups

self.buffers[adapter_name].put(result)  # Buffer per adapter

```

### Batch Collection with Round-Robin Selection

The trainer retrieves batches through `collect_batch()`, which calls `worker.get_groups()`. This method:

- Iterates live adapters in **round-robin order**
- Respects `min_groups_per_dp_split` per adapter
- Enforces per-adapter rollout batch sizes
- Discards groups from **retired adapters** or **stale registrations**

```python

# Key invocation in training loop:

batch = await collect_batch(args, worker, snapshot)

```

---

## Adapter Registry and Lifecycle Management

Central coordination happens through `MultiLoRAController` in [`miles/ray/multi_lora/controller.py`](https://github.com/radixark/miles/blob/main/miles/ray/multi_lora/controller.py), a **Ray actor** that maintains the canonical view of all adapters.

### Registration Flow

1. `MultiLoRAController.register_adapter(name, config)` delegates to `MultiLoRABackend.resolve_adapter_config()` for validation
2. Backend records `slot`, `version`, and `data_path` for the adapter
3. `AdaptersCache` (TTL singleton) provides fast lookup during rollout

### Safe Retirement and Abort Handling

When adapters retire, `MultiLoRABackend.abort_adapter_requests()` cancels in-flight requests to prevent stale generation:

```python

# In miles/ray/multi_lora/backend.py

async def abort_adapter_requests(self, adapter_name: str) -> None:
    prefix = f"{adapter_name}{RID_SEPARATOR}"
    urls = await self.worker_urls()
    await asyncio.gather(
        *(self.client.post(f"{url}/abort_request", 
           json={"rid": prefix, "prefix": True})
          for url in urls), return_exceptions=True
    )

```

The `prefix=True` parameter ensures scoped cancellation—only requests matching the adapter's routing-id prefix are aborted, leaving other adapters undisturbed even when sharing the same GPU slot.

---

## Version Consistency and Staleness Filtering

Miles implements multiple mechanisms to ensure **weight-version consistency**:

- **Registration-scoped buffering**: `GroupBuffer.drop_foreign()` discards groups from previous registrations when an adapter name is reused
- **Explicit version stamping**: Each sample carries `slot_version` enabling post-hoc staleness detection
- **Controller snapshots**: `snapshot` objects passed to `collect_batch()` capture the adapter registry state at batch assembly time

These defenses prevent the **silent correctness bug** where training proceeds on generations produced with outdated LoRA weights.

---

## Configuration and Throughput Considerations

Key parameters controlling the integration:

| Parameter | Effect |
|-----------|--------|
| `args.rollout_batch_size` | Determines producer concurrency |
| `args.sglang_server_concurrency` | Upper bound on GPU-side parallel generation |
| `min_groups_per_dp_split` | Minimum groups per adapter per batch |
| `empty_batch_timeout` | Maximum wait for incomplete batches |

The architecture separates **throughput-bound** concerns (async producer, SGLang engine) from **correctness-bound** concerns (version tracking, staleness filtering, abort handling).

---

## Summary

- **Single-LoRA rollouts** inject `lora_path`, `rid`, and `extra_key` into SGLang HTTP payloads via `generate()` in [`sglang_rollout.py`](https://github.com/radixark/miles/blob/main/sglang_rollout.py)
- **Multi-LoRA rollouts** use `AsyncMultiLoRAWorker` with per-adapter `GroupBuffer` storage and round-robin `collect_batch()` retrieval
- **Adapter lifecycle** is managed by `MultiLoRAController` and `MultiLoRABackend`, with safe retirement via scoped request abort
- **Version safety** is ensured through registration IDs, explicit version stamping, and dynamic staleness filtering
- All components integrate with SGLang's native LoRA serving capability while adding Miles-specific correctness guarantees

---

## Frequently Asked Questions

### How does Miles prevent using stale LoRA weights during Multi-LoRA training?

Each generated group is stamped with the adapter's current `slot_version` and `registration_id`. The `collect_batch()` function in [`async_rollout.py`](https://github.com/radixark/miles/blob/main/async_rollout.py) filters out groups whose registration IDs don't match the current adapter state. Additionally, when an adapter is re-registered (e.g., after checkpoint reload), `GroupBuffer.drop_foreign()` automatically drops buffered groups from the previous registration.

### What happens when an adapter is retired mid-rollout?

The `MultiLoRABackend.abort_adapter_requests()` method sends abort requests to all SGLang workers with a prefix matching the adapter's routing-id. This cancels in-flight generations for that specific adapter without affecting concurrent adapters, even when they share GPU slots. The backend then deregisters the adapter from the controller.

### Why does single-LoRA mode use three routing fields (`lora_path`, `rid`, `extra_key`)?

These fields serve distinct purposes: `lora_path` tells SGLang which weight file to load; `rid` enables request routing and scoped aborts across the distributed system; `extra_key` encodes version information for debugging and verification. Together they support both functional correctness (loading right weights) and operational visibility (tracking which version produced each generation).

### How does the async worker maintain throughput without blocking training?

`AsyncMultiLoRAWorker` runs a dedicated producer thread that operates independently of the training loop. It pre-generates and buffers prompt groups, allowing `collect_batch()` to retrieve ready data with minimal latency. The concurrency level is controlled by `args.rollout_batch_size`, while `args.sglang_server_concurrency` limits GPU-side parallelism to prevent resource exhaustion.