How to Configure SGLang for Efficient Rollout in the Miles Framework

To configure SGLang for efficient rollout in Miles, pass a YAML config via --sglang-config and tune GPU allocation, prefill/decode splitting, memory offloading, and LoRA settings through per-group overrides.

Miles uses a dedicated SGLang backend to serve models during training, making SGLang rollout configuration a critical bottleneck for throughput. The configuration system in radixark/miles combines command-line arguments, structured YAML, and per-group overrides to control how inference engines are placed across GPUs. This guide walks through the source code architecture—found primarily in miles/backends/sglang_utils/sglang_config.py—and provides concrete patterns for maximizing efficiency.

How Miles Builds the SGLang Configuration

The configuration pipeline follows four distinct phases, starting at resolve_sglang_config() in sglang_config.py.

Entry Point and Source Selection

resolve_sglang_config(args) (lines L82-L86) selects the raw configuration source:

  • A user-provided YAML file (--sglang-config)
  • A legacy --prefill-num-servers flag for backward compatibility
  • A default single-group configuration

Raw Model Descriptors

Two internal dataclasses capture the initial specification:

  • _RawModelConfig (lines L20-L38): Stores model_path, num_gpus_per_engine, and update_weights
  • _RawServerGroupConfig (lines L40-L65): Stores worker_type (prefill/decode/regular), GPU counts, and an overrides dictionary that maps directly to SGLang's ServerArgs

Resolution to Concrete Objects

ModelConfig.resolve() (lines L96-L124) transforms raw specs into executable configurations:

  1. Applies defaults for model_path and num_gpus_per_engine
  2. Builds ServerGroupConfig objects with computed GPU offsets
  3. Sets update_weights to determine trainer integration

GPU Offset Tracking

The _OffsetCursor helper (lines L76-L80) maintains cumulative GPU and engine offsets. ServerGroupConfig.resolve() (lines L57-L90) applies these to ensure each group receives a unique placement slice without overlap.

Final Aggregated Config

SglangConfig (lines L60-L74) collects resolved ModelConfigs and exposes utilities like has_pd_disaggregation() to detect prefill/decode splits.

Core Efficiency Knobs

Area Configuration Effect Source Location
Tensor parallelism num_gpus_per_engine Controls TP size; larger values reduce inter-GPU traffic but consume more GPUs per engine _RawModelConfig.num_gpus_per_engine
Prefill/decode split worker_type: prefill / decode Separates prompt processing from token generation to reduce latency for long contexts ServerGroupConfig.worker_type
Memory offloading --offload-rollout Enables SGLang's memory-saver mode, moving weights to CPU when idle _compute_server_args() at line L95
LoRA support enable_lora, max_loras_per_batch, lora_target_modules Activates adapter-based inference; multi-LoRA gated by is_multi_lora_enabled(args) _compute_server_args() lines L42-L55
Eval isolation --eval-num-gpus, --eval-sglang-* Creates separate eval engines that skip training replay, reducing bandwidth _compute_raw_sglang_config() lines L92-L108
Fine-grained overrides overrides dict per server group Directly sets any ServerArgs field (e.g., tp_size, dp_size) ServerGroupConfig.resolve() lines L84-L87
Disaggregation bootstrap disaggregation_bootstrap_port Enables prefill worker discovery without centralized scheduling compute_engine_launch_cmd() lines L22-L28

Complete Configuration Example

This YAML configures an actor model with split prefill/decode workers plus a frozen reference model:

sglang:
  - name: actor
    model_path: /mnt/checkpoints/actor
    update_weights: true
    num_gpus_per_engine: 4
    server_groups:
      - worker_type: prefill
        num_gpus: 8
        overrides:
          tp_size: 4
          load_balance_method: round_robin
      - worker_type: decode
        num_gpus: 8
        overrides:
          tp_size: 4
          prefill_round_robin_balance: true
  - name: ref
    model_path: /mnt/checkpoints/ref
    update_weights: false
    server_groups:
      - worker_type: regular
        num_gpus: 4

Launch with:

python -m miles.main.train \
    --sglang-config=sglang_config.yaml \
    --rollout-num-gpus=20 \
    --rollout-num-gpus-per-engine=4 \
    --offload-rollout \
    --use-rollout-routing-replay \
    --lora-rollout-enabled \
    --lora-adapter-path=./lora_adapter

Miles executes this through resolve_sglang_config() → ServerGroupConfig.resolve() → compute_engine_launch_cmd() (lines L33-L70 in sglang_engine.py), producing commands like:

python -m sglang.launch_server \
  --model_path=/mnt/checkpoints/actor \
  --host=0.0.0.0 --port=31000 \
  --tp_size=4 --dp_size=1 --pp_size=1 --ep_size=1 \
  --enable_memory_saver \
  --enable_lora \
  --max_loras_per_batch=1 \
  --max_lora_rank=8 \
  --lora_paths=adapter=./lora_adapter

Programmatic Configuration Access

Inspect Resolved Configurations

from miles.backends.sglang_utils.sglang_config import resolve_sglang_config

sglang_cfg = resolve_sglang_config(args)

for model in sglang_cfg.models:
    print(f"Model {model.name}:")
    for grp in model.server_groups:
        print(
            f"  {grp.worker_type} — GPUs: {grp.num_gpus}, "
            f"per engine: {grp.num_gpus_per_engine}, "
            f"offset: {grp.gpu_offset}"
        )

Build Launch Commands Directly

from miles.backends.sglang_utils.sglang_engine import compute_engine_launch_cmd

cmd = compute_engine_launch_cmd(
    args,
    node_rank=0,
    worker_type="prefill",
    base_gpu_id=0,
    sglang_overrides={"tp_size": 4},
    num_gpus_per_engine=4,
    dist_init_addr="10.1.0.1",
    nccl_port=12345,
    host="0.0.0.0",
    port=31000,
    disaggregation_bootstrap_port=31100,
    engine_info_bootstrap_port=31200,
    gated_launch_port=31300,
    random_seed=42,
)

print("Launch command:", cmd)

Debug Server Arguments

from miles.backends.sglang_utils.sglang_engine import _compute_server_args

server_args = _compute_server_args(
    args,
    node_rank=0,
    dist_init_addr="10.1.0.1",
    nccl_port=12345,
    host="0.0.0.0",
    port=31000,
    worker_type="decode",
    disaggregation_bootstrap_port=None,
    base_gpu_id=0,
    engine_info_bootstrap_port=None,
    sglang_overrides={"tp_size": 4},
    num_gpus_per_engine=4,
    gated_launch_port=31300,
    random_seed=42,
)

print(server_args)  # Inspect final ServerArgs dict

Performance Optimization Strategies

Strategy Implementation When to Apply
Minimum TP of 2 Set num_gpus_per_engine: 2 or higher Always; keeps tensor-parallel pipelines full
Memory-saver mode Add --offload-rollout GPU memory constrained; larger batch sizes needed
Prefill-decode split Define separate worker_type: prefill and worker_type: decode groups Long prompts, high concurrency
Routing replay Add --use-rollout-routing-replay Multi-node setups to reduce expert data movement
Right-size LoRA batches Set max_loras_per_batch to actual adapter count Mixed LoRA workloads; prevents memory waste or runtime errors
Multi-LoRA mode Add --multi-lora-enabled Many active adapters; improves cache locality
Isolated eval fleet Use --eval-num-gpus Continuous eval needed; avoids replay overhead

Key Source Files

File Purpose
miles/backends/sglang_utils/sglang_config.py Core dataclasses (_RawModelConfig, ServerGroupConfig, SglangConfig) and resolve_sglang_config() entry point
miles/backends/sglang_utils/sglang_engine.py compute_engine_launch_cmd() and _compute_server_args() for command generation
miles/backends/sglang_utils/sglang_api_client.py Runtime client for weight sync and health checks
miles/backends/sglang_utils/sglang_router_api_client.py Request routing and engine aggregation
tests/fast/utils/mock_sglang_server.py Test fixtures showing expected protocol shapes

Summary

  • SGLang rollout configuration in Miles centers on resolve_sglang_config() in sglang_config.py, which merges CLI flags, YAML specs, and per-group overrides into executable ServerGroupConfig objects
  • GPU efficiency comes from tuning num_gpus_per_engine for tensor parallelism and optionally splitting prefill/decode workers
  • Memory efficiency uses --offload-rollout to enable SGLang's memory-saver mode
  • LoRA efficiency requires correct max_loras_per_batch sizing and optional multi-LoRA mode
  • Isolation via separate eval fleets prevents unnecessary replay traffic

Frequently Asked Questions

What file format does Miles use for SGLang configuration?

Miles accepts YAML configuration files passed via --sglang-config. The schema defines a top-level sglang key containing a list of model entries, each with name, model_path, update_weights, num_gpus_per_engine, and server_groups. See lines L78-L96 in sglang_config.py for the full parser implementation.

How do I enable memory-efficient inference when GPU memory is limited?

Add --offload-rollout to your command line. This sets enable_memory_saver in the generated SGLang arguments (line L95 of sglang_engine.py), which moves model weights to CPU between forward passes, freeing GPU memory for larger batch sizes at the cost of weight-transfer overhead.

What is prefill-decode disaggregation and when should I use it?

Prefill-decode disaggregation splits inference into dedicated prefill workers that process input prompts and decode workers that generate output tokens. Configure this by setting worker_type: prefill and worker_type: decode in separate server groups. Use it when serving long contexts or high request concurrency, as it prevents slow prefill operations from blocking low-latency decode generation.

How do per-group overrides interact with command-line defaults?

Per-group overrides dictionaries take precedence over CLI-derived defaults. In ServerGroupConfig.resolve() (lines L84-L87), values from sglang_overrides are merged last into the ServerArgs structure, allowing fine-grained control without modifying global flags.

Can I use LoRA adapters during rollout without restarting engines?

Yes. Enable with --lora-rollout-enabled or --multi-lora-enabled. The framework automatically populates enable_lora, max_loras_per_batch, and lora_target_modules through _compute_server_args() (lines L42-L55). Ensure max_loras_per_batch matches your expected concurrent adapter count to avoid runtime errors or memory waste.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →