How to Configure SGLang for Efficient Rollout in the Miles Framework
To configure SGLang for efficient rollout in Miles, pass a YAML config via --sglang-config and tune GPU allocation, prefill/decode splitting, memory offloading, and LoRA settings through per-group overrides.
Miles uses a dedicated SGLang backend to serve models during training, making SGLang rollout configuration a critical bottleneck for throughput. The configuration system in radixark/miles combines command-line arguments, structured YAML, and per-group overrides to control how inference engines are placed across GPUs. This guide walks through the source code architecture—found primarily in miles/backends/sglang_utils/sglang_config.py—and provides concrete patterns for maximizing efficiency.
How Miles Builds the SGLang Configuration
The configuration pipeline follows four distinct phases, starting at resolve_sglang_config() in sglang_config.py.
Entry Point and Source Selection
resolve_sglang_config(args) (lines L82-L86) selects the raw configuration source:
- A user-provided YAML file (
--sglang-config) - A legacy
--prefill-num-serversflag for backward compatibility - A default single-group configuration
Raw Model Descriptors
Two internal dataclasses capture the initial specification:
_RawModelConfig(lines L20-L38): Storesmodel_path,num_gpus_per_engine, andupdate_weights_RawServerGroupConfig(lines L40-L65): Storesworker_type(prefill/decode/regular), GPU counts, and anoverridesdictionary that maps directly to SGLang'sServerArgs
Resolution to Concrete Objects
ModelConfig.resolve() (lines L96-L124) transforms raw specs into executable configurations:
- Applies defaults for
model_pathandnum_gpus_per_engine - Builds
ServerGroupConfigobjects with computed GPU offsets - Sets
update_weightsto determine trainer integration
GPU Offset Tracking
The _OffsetCursor helper (lines L76-L80) maintains cumulative GPU and engine offsets. ServerGroupConfig.resolve() (lines L57-L90) applies these to ensure each group receives a unique placement slice without overlap.
Final Aggregated Config
SglangConfig (lines L60-L74) collects resolved ModelConfigs and exposes utilities like has_pd_disaggregation() to detect prefill/decode splits.
Core Efficiency Knobs
| Area | Configuration | Effect | Source Location |
|---|---|---|---|
| Tensor parallelism | num_gpus_per_engine |
Controls TP size; larger values reduce inter-GPU traffic but consume more GPUs per engine | _RawModelConfig.num_gpus_per_engine |
| Prefill/decode split | worker_type: prefill / decode |
Separates prompt processing from token generation to reduce latency for long contexts | ServerGroupConfig.worker_type |
| Memory offloading | --offload-rollout |
Enables SGLang's memory-saver mode, moving weights to CPU when idle | _compute_server_args() at line L95 |
| LoRA support | enable_lora, max_loras_per_batch, lora_target_modules |
Activates adapter-based inference; multi-LoRA gated by is_multi_lora_enabled(args) |
_compute_server_args() lines L42-L55 |
| Eval isolation | --eval-num-gpus, --eval-sglang-* |
Creates separate eval engines that skip training replay, reducing bandwidth | _compute_raw_sglang_config() lines L92-L108 |
| Fine-grained overrides | overrides dict per server group |
Directly sets any ServerArgs field (e.g., tp_size, dp_size) |
ServerGroupConfig.resolve() lines L84-L87 |
| Disaggregation bootstrap | disaggregation_bootstrap_port |
Enables prefill worker discovery without centralized scheduling | compute_engine_launch_cmd() lines L22-L28 |
Complete Configuration Example
This YAML configures an actor model with split prefill/decode workers plus a frozen reference model:
sglang:
- name: actor
model_path: /mnt/checkpoints/actor
update_weights: true
num_gpus_per_engine: 4
server_groups:
- worker_type: prefill
num_gpus: 8
overrides:
tp_size: 4
load_balance_method: round_robin
- worker_type: decode
num_gpus: 8
overrides:
tp_size: 4
prefill_round_robin_balance: true
- name: ref
model_path: /mnt/checkpoints/ref
update_weights: false
server_groups:
- worker_type: regular
num_gpus: 4
Launch with:
python -m miles.main.train \
--sglang-config=sglang_config.yaml \
--rollout-num-gpus=20 \
--rollout-num-gpus-per-engine=4 \
--offload-rollout \
--use-rollout-routing-replay \
--lora-rollout-enabled \
--lora-adapter-path=./lora_adapter
Miles executes this through resolve_sglang_config() → ServerGroupConfig.resolve() → compute_engine_launch_cmd() (lines L33-L70 in sglang_engine.py), producing commands like:
python -m sglang.launch_server \
--model_path=/mnt/checkpoints/actor \
--host=0.0.0.0 --port=31000 \
--tp_size=4 --dp_size=1 --pp_size=1 --ep_size=1 \
--enable_memory_saver \
--enable_lora \
--max_loras_per_batch=1 \
--max_lora_rank=8 \
--lora_paths=adapter=./lora_adapter
Programmatic Configuration Access
Inspect Resolved Configurations
from miles.backends.sglang_utils.sglang_config import resolve_sglang_config
sglang_cfg = resolve_sglang_config(args)
for model in sglang_cfg.models:
print(f"Model {model.name}:")
for grp in model.server_groups:
print(
f" {grp.worker_type} — GPUs: {grp.num_gpus}, "
f"per engine: {grp.num_gpus_per_engine}, "
f"offset: {grp.gpu_offset}"
)
Build Launch Commands Directly
from miles.backends.sglang_utils.sglang_engine import compute_engine_launch_cmd
cmd = compute_engine_launch_cmd(
args,
node_rank=0,
worker_type="prefill",
base_gpu_id=0,
sglang_overrides={"tp_size": 4},
num_gpus_per_engine=4,
dist_init_addr="10.1.0.1",
nccl_port=12345,
host="0.0.0.0",
port=31000,
disaggregation_bootstrap_port=31100,
engine_info_bootstrap_port=31200,
gated_launch_port=31300,
random_seed=42,
)
print("Launch command:", cmd)
Debug Server Arguments
from miles.backends.sglang_utils.sglang_engine import _compute_server_args
server_args = _compute_server_args(
args,
node_rank=0,
dist_init_addr="10.1.0.1",
nccl_port=12345,
host="0.0.0.0",
port=31000,
worker_type="decode",
disaggregation_bootstrap_port=None,
base_gpu_id=0,
engine_info_bootstrap_port=None,
sglang_overrides={"tp_size": 4},
num_gpus_per_engine=4,
gated_launch_port=31300,
random_seed=42,
)
print(server_args) # Inspect final ServerArgs dict
Performance Optimization Strategies
| Strategy | Implementation | When to Apply |
|---|---|---|
| Minimum TP of 2 | Set num_gpus_per_engine: 2 or higher |
Always; keeps tensor-parallel pipelines full |
| Memory-saver mode | Add --offload-rollout |
GPU memory constrained; larger batch sizes needed |
| Prefill-decode split | Define separate worker_type: prefill and worker_type: decode groups |
Long prompts, high concurrency |
| Routing replay | Add --use-rollout-routing-replay |
Multi-node setups to reduce expert data movement |
| Right-size LoRA batches | Set max_loras_per_batch to actual adapter count |
Mixed LoRA workloads; prevents memory waste or runtime errors |
| Multi-LoRA mode | Add --multi-lora-enabled |
Many active adapters; improves cache locality |
| Isolated eval fleet | Use --eval-num-gpus |
Continuous eval needed; avoids replay overhead |
Key Source Files
| File | Purpose |
|---|---|
miles/backends/sglang_utils/sglang_config.py |
Core dataclasses (_RawModelConfig, ServerGroupConfig, SglangConfig) and resolve_sglang_config() entry point |
miles/backends/sglang_utils/sglang_engine.py |
compute_engine_launch_cmd() and _compute_server_args() for command generation |
miles/backends/sglang_utils/sglang_api_client.py |
Runtime client for weight sync and health checks |
miles/backends/sglang_utils/sglang_router_api_client.py |
Request routing and engine aggregation |
tests/fast/utils/mock_sglang_server.py |
Test fixtures showing expected protocol shapes |
Summary
- SGLang rollout configuration in Miles centers on
resolve_sglang_config()insglang_config.py, which merges CLI flags, YAML specs, and per-group overrides into executableServerGroupConfigobjects - GPU efficiency comes from tuning
num_gpus_per_enginefor tensor parallelism and optionally splitting prefill/decode workers - Memory efficiency uses
--offload-rolloutto enable SGLang's memory-saver mode - LoRA efficiency requires correct
max_loras_per_batchsizing and optional multi-LoRA mode - Isolation via separate eval fleets prevents unnecessary replay traffic
Frequently Asked Questions
What file format does Miles use for SGLang configuration?
Miles accepts YAML configuration files passed via --sglang-config. The schema defines a top-level sglang key containing a list of model entries, each with name, model_path, update_weights, num_gpus_per_engine, and server_groups. See lines L78-L96 in sglang_config.py for the full parser implementation.
How do I enable memory-efficient inference when GPU memory is limited?
Add --offload-rollout to your command line. This sets enable_memory_saver in the generated SGLang arguments (line L95 of sglang_engine.py), which moves model weights to CPU between forward passes, freeing GPU memory for larger batch sizes at the cost of weight-transfer overhead.
What is prefill-decode disaggregation and when should I use it?
Prefill-decode disaggregation splits inference into dedicated prefill workers that process input prompts and decode workers that generate output tokens. Configure this by setting worker_type: prefill and worker_type: decode in separate server groups. Use it when serving long contexts or high request concurrency, as it prevents slow prefill operations from blocking low-latency decode generation.
How do per-group overrides interact with command-line defaults?
Per-group overrides dictionaries take precedence over CLI-derived defaults. In ServerGroupConfig.resolve() (lines L84-L87), values from sglang_overrides are merged last into the ServerArgs structure, allowing fine-grained control without modifying global flags.
Can I use LoRA adapters during rollout without restarting engines?
Yes. Enable with --lora-rollout-enabled or --multi-lora-enabled. The framework automatically populates enable_lora, max_loras_per_batch, and lora_target_modules through _compute_server_args() (lines L42-L55). Ensure max_loras_per_batch matches your expected concurrent adapter count to avoid runtime errors or memory waste.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →