# How to Configure SGLang for Efficient Rollout in the Miles Framework

> Configure SGLang for efficient rollout in Miles by passing YAML config and tuning GPU allocation, memory offloading, and LoRA via per-group overrides for optimal performance.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: how-to-guide
- Published: 2026-09-06

---

**To configure SGLang for efficient rollout in Miles, pass a YAML config via `--sglang-config` and tune GPU allocation, prefill/decode splitting, memory offloading, and LoRA settings through per-group overrides.**

Miles uses a dedicated SGLang backend to serve models during training, making **SGLang rollout configuration** a critical bottleneck for throughput. The configuration system in `radixark/miles` combines command-line arguments, structured YAML, and per-group overrides to control how inference engines are placed across GPUs. This guide walks through the source code architecture—found primarily in [`miles/backends/sglang_utils/sglang_config.py`](https://github.com/radixark/miles/blob/main/miles/backends/sglang_utils/sglang_config.py)—and provides concrete patterns for maximizing efficiency.

## How Miles Builds the SGLang Configuration

The configuration pipeline follows four distinct phases, starting at `resolve_sglang_config()` in [`sglang_config.py`](https://github.com/radixark/miles/blob/main/sglang_config.py).

### Entry Point and Source Selection

`resolve_sglang_config(args)` (lines L82-L86) selects the raw configuration source:

- A user-provided YAML file (`--sglang-config`)
- A legacy `--prefill-num-servers` flag for backward compatibility
- A default single-group configuration

### Raw Model Descriptors

Two internal dataclasses capture the initial specification:

- **`_RawModelConfig`** (lines L20-L38): Stores `model_path`, `num_gpus_per_engine`, and `update_weights`
- **`_RawServerGroupConfig`** (lines L40-L65): Stores `worker_type` (prefill/decode/regular), GPU counts, and an `overrides` dictionary that maps directly to SGLang's `ServerArgs`

### Resolution to Concrete Objects

`ModelConfig.resolve()` (lines L96-L124) transforms raw specs into executable configurations:

1. Applies defaults for `model_path` and `num_gpus_per_engine`
2. Builds `ServerGroupConfig` objects with computed GPU offsets
3. Sets `update_weights` to determine trainer integration

### GPU Offset Tracking

The `_OffsetCursor` helper (lines L76-L80) maintains cumulative GPU and engine offsets. `ServerGroupConfig.resolve()` (lines L57-L90) applies these to ensure each group receives a unique placement slice without overlap.

### Final Aggregated Config

`SglangConfig` (lines L60-L74) collects resolved `ModelConfig`s and exposes utilities like `has_pd_disaggregation()` to detect prefill/decode splits.

## Core Efficiency Knobs

| Area | Configuration | Effect | Source Location |
|------|-------------|--------|-----------------|
| **Tensor parallelism** | `num_gpus_per_engine` | Controls TP size; larger values reduce inter-GPU traffic but consume more GPUs per engine | `_RawModelConfig.num_gpus_per_engine` |
| **Prefill/decode split** | `worker_type: prefill` / `decode` | Separates prompt processing from token generation to reduce latency for long contexts | `ServerGroupConfig.worker_type` |
| **Memory offloading** | `--offload-rollout` | Enables SGLang's memory-saver mode, moving weights to CPU when idle | `_compute_server_args()` at line L95 |
| **LoRA support** | `enable_lora`, `max_loras_per_batch`, `lora_target_modules` | Activates adapter-based inference; multi-LoRA gated by `is_multi_lora_enabled(args)` | `_compute_server_args()` lines L42-L55 |
| **Eval isolation** | `--eval-num-gpus`, `--eval-sglang-*` | Creates separate eval engines that skip training replay, reducing bandwidth | `_compute_raw_sglang_config()` lines L92-L108 |
| **Fine-grained overrides** | `overrides` dict per server group | Directly sets any `ServerArgs` field (e.g., `tp_size`, `dp_size`) | `ServerGroupConfig.resolve()` lines L84-L87 |
| **Disaggregation bootstrap** | `disaggregation_bootstrap_port` | Enables prefill worker discovery without centralized scheduling | `compute_engine_launch_cmd()` lines L22-L28 |

## Complete Configuration Example

This YAML configures an actor model with split prefill/decode workers plus a frozen reference model:

```yaml
sglang:
  - name: actor
    model_path: /mnt/checkpoints/actor
    update_weights: true
    num_gpus_per_engine: 4
    server_groups:
      - worker_type: prefill
        num_gpus: 8
        overrides:
          tp_size: 4
          load_balance_method: round_robin
      - worker_type: decode
        num_gpus: 8
        overrides:
          tp_size: 4
          prefill_round_robin_balance: true
  - name: ref
    model_path: /mnt/checkpoints/ref
    update_weights: false
    server_groups:
      - worker_type: regular
        num_gpus: 4

```

Launch with:

```bash
python -m miles.main.train \
    --sglang-config=sglang_config.yaml \
    --rollout-num-gpus=20 \
    --rollout-num-gpus-per-engine=4 \
    --offload-rollout \
    --use-rollout-routing-replay \
    --lora-rollout-enabled \
    --lora-adapter-path=./lora_adapter

```

Miles executes this through `resolve_sglang_config()` → `ServerGroupConfig.resolve()` → `compute_engine_launch_cmd()` (lines L33-L70 in [`sglang_engine.py`](https://github.com/radixark/miles/blob/main/sglang_engine.py)), producing commands like:

```bash
python -m sglang.launch_server \
  --model_path=/mnt/checkpoints/actor \
  --host=0.0.0.0 --port=31000 \
  --tp_size=4 --dp_size=1 --pp_size=1 --ep_size=1 \
  --enable_memory_saver \
  --enable_lora \
  --max_loras_per_batch=1 \
  --max_lora_rank=8 \
  --lora_paths=adapter=./lora_adapter

```

## Programmatic Configuration Access

### Inspect Resolved Configurations

```python
from miles.backends.sglang_utils.sglang_config import resolve_sglang_config

sglang_cfg = resolve_sglang_config(args)

for model in sglang_cfg.models:
    print(f"Model {model.name}:")
    for grp in model.server_groups:
        print(
            f"  {grp.worker_type} — GPUs: {grp.num_gpus}, "
            f"per engine: {grp.num_gpus_per_engine}, "
            f"offset: {grp.gpu_offset}"
        )

```

### Build Launch Commands Directly

```python
from miles.backends.sglang_utils.sglang_engine import compute_engine_launch_cmd

cmd = compute_engine_launch_cmd(
    args,
    node_rank=0,
    worker_type="prefill",
    base_gpu_id=0,
    sglang_overrides={"tp_size": 4},
    num_gpus_per_engine=4,
    dist_init_addr="10.1.0.1",
    nccl_port=12345,
    host="0.0.0.0",
    port=31000,
    disaggregation_bootstrap_port=31100,
    engine_info_bootstrap_port=31200,
    gated_launch_port=31300,
    random_seed=42,
)

print("Launch command:", cmd)

```

### Debug Server Arguments

```python
from miles.backends.sglang_utils.sglang_engine import _compute_server_args

server_args = _compute_server_args(
    args,
    node_rank=0,
    dist_init_addr="10.1.0.1",
    nccl_port=12345,
    host="0.0.0.0",
    port=31000,
    worker_type="decode",
    disaggregation_bootstrap_port=None,
    base_gpu_id=0,
    engine_info_bootstrap_port=None,
    sglang_overrides={"tp_size": 4},
    num_gpus_per_engine=4,
    gated_launch_port=31300,
    random_seed=42,
)

print(server_args)  # Inspect final ServerArgs dict

```

## Performance Optimization Strategies

| Strategy | Implementation | When to Apply |
|----------|---------------|-------------|
| **Minimum TP of 2** | Set `num_gpus_per_engine: 2` or higher | Always; keeps tensor-parallel pipelines full |
| **Memory-saver mode** | Add `--offload-rollout` | GPU memory constrained; larger batch sizes needed |
| **Prefill-decode split** | Define separate `worker_type: prefill` and `worker_type: decode` groups | Long prompts, high concurrency |
| **Routing replay** | Add `--use-rollout-routing-replay` | Multi-node setups to reduce expert data movement |
| **Right-size LoRA batches** | Set `max_loras_per_batch` to actual adapter count | Mixed LoRA workloads; prevents memory waste or runtime errors |
| **Multi-LoRA mode** | Add `--multi-lora-enabled` | Many active adapters; improves cache locality |
| **Isolated eval fleet** | Use `--eval-num-gpus` | Continuous eval needed; avoids replay overhead |

## Key Source Files

| File | Purpose |
|------|---------|
| [`miles/backends/sglang_utils/sglang_config.py`](https://github.com/radixark/miles/blob/main/miles/backends/sglang_utils/sglang_config.py) | Core dataclasses (`_RawModelConfig`, `ServerGroupConfig`, `SglangConfig`) and `resolve_sglang_config()` entry point |
| [`miles/backends/sglang_utils/sglang_engine.py`](https://github.com/radixark/miles/blob/main/miles/backends/sglang_utils/sglang_engine.py) | `compute_engine_launch_cmd()` and `_compute_server_args()` for command generation |
| [`miles/backends/sglang_utils/sglang_api_client.py`](https://github.com/radixark/miles/blob/main/miles/backends/sglang_utils/sglang_api_client.py) | Runtime client for weight sync and health checks |
| [`miles/backends/sglang_utils/sglang_router_api_client.py`](https://github.com/radixark/miles/blob/main/miles/backends/sglang_utils/sglang_router_api_client.py) | Request routing and engine aggregation |
| [`tests/fast/utils/mock_sglang_server.py`](https://github.com/radixark/miles/blob/main/tests/fast/utils/mock_sglang_server.py) | Test fixtures showing expected protocol shapes |

## Summary

- **SGLang rollout configuration in Miles** centers on `resolve_sglang_config()` in [`sglang_config.py`](https://github.com/radixark/miles/blob/main/sglang_config.py), which merges CLI flags, YAML specs, and per-group overrides into executable `ServerGroupConfig` objects
- **GPU efficiency** comes from tuning `num_gpus_per_engine` for tensor parallelism and optionally splitting prefill/decode workers
- **Memory efficiency** uses `--offload-rollout` to enable SGLang's memory-saver mode
- **LoRA efficiency** requires correct `max_loras_per_batch` sizing and optional multi-LoRA mode
- **Isolation** via separate eval fleets prevents unnecessary replay traffic

## Frequently Asked Questions

### What file format does Miles use for SGLang configuration?

Miles accepts YAML configuration files passed via `--sglang-config`. The schema defines a top-level `sglang` key containing a list of model entries, each with `name`, `model_path`, `update_weights`, `num_gpus_per_engine`, and `server_groups`. See lines L78-L96 in [`sglang_config.py`](https://github.com/radixark/miles/blob/main/sglang_config.py) for the full parser implementation.

### How do I enable memory-efficient inference when GPU memory is limited?

Add `--offload-rollout` to your command line. This sets `enable_memory_saver` in the generated SGLang arguments (line L95 of [`sglang_engine.py`](https://github.com/radixark/miles/blob/main/sglang_engine.py)), which moves model weights to CPU between forward passes, freeing GPU memory for larger batch sizes at the cost of weight-transfer overhead.

### What is prefill-decode disaggregation and when should I use it?

Prefill-decode disaggregation splits inference into dedicated **prefill workers** that process input prompts and **decode workers** that generate output tokens. Configure this by setting `worker_type: prefill` and `worker_type: decode` in separate server groups. Use it when serving long contexts or high request concurrency, as it prevents slow prefill operations from blocking low-latency decode generation.

### How do per-group overrides interact with command-line defaults?

Per-group `overrides` dictionaries take precedence over CLI-derived defaults. In `ServerGroupConfig.resolve()` (lines L84-L87), values from `sglang_overrides` are merged last into the `ServerArgs` structure, allowing fine-grained control without modifying global flags.

### Can I use LoRA adapters during rollout without restarting engines?

Yes. Enable with `--lora-rollout-enabled` or `--multi-lora-enabled`. The framework automatically populates `enable_lora`, `max_loras_per_batch`, and `lora_target_modules` through `_compute_server_args()` (lines L42-L55). Ensure `max_loras_per_batch` matches your expected concurrent adapter count to avoid runtime errors or memory waste.