How `--offload-train` and `--offload-rollout` Memory Offloading Works in Miles

Miles uses --offload-train to move optimizer state and model parameters to CPU or NVMe during training, and --offload-rollout to move inference tensors (KV-cache and weights) off-GPU between rollouts, enabling training of models that exceed GPU memory capacity.

Miles is a distributed training and inference framework that supports memory offloading strategies to fit large language models into limited GPU resources. The --offload-train and --offload-rollout flags work together to swap tensors between GPU memory and slower storage tiers—CPU RAM or NVMe disk—at strategic points in the training loop. This article explains exactly how these mechanisms work, with references to the actual source implementation in radixark/miles.


What the Offloading Flags Control

Flag Purpose Default Target
--offload-train Offload training tensors (optimizer state, parameters) when not actively computing gradients cpu (configurable to disk)
--offload-rollout Offload rollout tensors (KV-cache, model weights) when not generating inference outputs Configurable via --offload-rollout-level

Both flags are enabled automatically when --colocate mode is active, which runs training and inference on the same GPU. Users can disable them with --no-offload-train or --no-offload-rollout.


How --offload-train Works

Flag Definition and Validation

The offloading infrastructure is defined in miles/utils/arguments.py. The --offload-train flag is part of a grouped option documented as "equivalent to --offload-train + --offload-rollout" when both training and rollout memory need management.


# From miles/utils/arguments.py lines 172, 214-224

parser.add_argument("--offload-train", action="store_true", ...)
parser.add_argument("--offload-train-target", choices=["cpu", "disk"], default="cpu")
parser.add_argument("--offload-train-disk-dir", type=str)
parser.add_argument("--offload-train-disk-chunk-mb", type=int, default=64)

Critical validation rules are enforced:

  • --offload-train-target=disk requires --offload-train to be enabled—this is asserted at lines 3434-3435
  • When disk is selected, both --offload-train-disk-dir and a positive chunk size must be provided (lines 342-348)

Training Loop Integration

In train.py, the training orchestrator calls the asynchronous helper offload_train() before each training step when the flag is true (lines 78-85):


# From train.py lines 78-85

if args.offload_train:
    await offload_train(
        actor_model=actor_model,
        critic_model=critic_model,
        target=args.offload_train_target,
        disk_dir=args.offload_train_disk_dir,
    )

The offload_train() coroutine invokes:

  • await actor_model.offload(target) — moves actor model parameters
  • await critic_model.offload(target) — moves critic model parameters (when PPO is used)

Disk Streaming for Optimizer State

When --stream-optimizer-state-to-disk is combined with --offload-train-target=disk, the optimizer tensors are written to NVMe in chunked streams. The chunk size is controlled by TMS_DISK_BACKUP_CHUNK_MB, which is injected into worker environments in miles/ray/specs/train.py (lines 128-144):


# From miles/ray/specs/train.py lines 128-144

if args.offload_train and args.offload_train_target == "disk":
    env_vars["TMS_DISK_BACKUP_CHUNK_MB"] = str(args.offload_train_disk_chunk_mb)

How --offload-rollout Works

Rollout-Level Configuration

The --offload-rollout flag controls which inference tensors leave GPU memory between generation passes. The specific components are selected via --offload-rollout-level:


# From train.py lines 118-123

offload_tags = []
if "kv_cache" in args.offload_rollout_level:
    offload_tags.append(GPU_MEMORY_TYPE_KV_CACHE)
if "weight" in args.offload_rollout_level:
    offload_tags.append(GPU_MEMORY_TYPE_WEIGHTS)

Common values include:

  • kv_cache — offload only the attention key-value cache
  • weight — offload model weights
  • kv_cache weight — offload both (maximum memory savings)

InferenceController Offloading Sequence

After each training step, if --offload-rollout is enabled, Miles delegates to the InferenceController, which performs offloading in two phases:

  1. KV-cache offload — controller.offload_kv() moves attention caches
  2. Weight offload — controller.offload_weights() moves model parameters

This sequencing ensures that active generation contexts are preserved until the rollout fully completes, then aggressively clears GPU memory before the next training step begins.


Backend-Specific Implementation Details

Megatron Backend

Megatron actors implement the actual offload(target) calls in miles/backends/megatron_utils/actor.py. The target parameter ("cpu" or "disk") determines whether tensors move to pinned host memory or are serialized to the NVMe path specified in --offload-train-disk-dir.

FSDP Backend Conflict Resolution

FSDP (Fully Sharded Data Parallel) has its own CPU offloading mechanism, which conflicts with Miles' explicit offloading. In miles/backends/fsdp_utils/actor.py (lines 98-100), Miles automatically detects and disables FSDP's native CPU offload when --offload-train is enabled:


# From miles/backends/fsdp_utils/actor.py lines 98-100

if args.offload_train and self.use_fsdp_cpu_offload:
    logger.warning("Disabling FSDP CPU offload due to Miles offloading")
    self.use_fsdp_cpu_offload = False

This prevents double-offloading overhead and ensures deterministic memory management.


Colocate Mode Requirements

When --colocate is specified—running training and inference on the same GPU—the memory offloading strategy becomes mandatory. The validator at line 35 of train.py enforces that both flags must be true:


# From train.py line 35

if args.colocate:
    assert args.offload_train, "colocate requires --offload-train"
    assert args.offload_rollout, "colocate requires --offload-rollout"

This requirement exists because a single GPU cannot simultaneously hold optimizer state, training parameters, model weights, and KV-cache for large models. The peak memory would exceed capacity without aggressive offloading.


Practical Configuration Examples

CPU-Based Offloading (Default)

python -m miles.main.train \
    --train-backend megatron \
    --colocate \
    --offload-train \
    --offload-train-target cpu \
    --offload-rollout \
    --offload-rollout-level kv_cache weight

NVMe Disk Offloading with Streaming

python -m miles.main.train \
    --train-backend megatron \
    --colocate \
    --offload-train \
    --offload-train-target disk \
    --offload-train-disk-dir /scratch/miles_offload \
    --offload-train-disk-chunk-mb 256 \
    --stream-optimizer-state-to-disk \
    --offload-rollout \
    --offload-rollout-level kv_cache weight

Key parameters:

  • --offload-train-disk-chunk-mb — controls granularity of NVMe writes (larger = fewer files, smaller = better overlap)
  • --stream-optimizer-state-to-disk — enables asynchronous optimizer state serialization to prevent blocking

Summary

  • --offload-train moves training tensors (optimizer state, parameters) to CPU or NVMe when gradients are not being computed, implemented via offload_train() in train.py and model.offload() in Megatron actors
  • --offload-rollout moves inference tensors (KV-cache, weights) off-GPU between rollouts, controlled by --offload-rollout-level and executed through InferenceController
  • Both flags are required in --colocate mode to fit large models on single GPUs, enforced by assertions at train.py:35
  • Target selection (cpu vs disk) and chunk sizing are validated in miles/utils/arguments.py lines 342-348 and 3434-3435
  • FSDP backend automatically disables native CPU offload to avoid conflicts, as shown in miles/backends/fsdp_utils/actor.py lines 98-100

Frequently Asked Questions

What happens if I enable --offload-train-target=disk without specifying --offload-train?

The argument parser raises an error. According to miles/utils/arguments.py lines 3434-3435, disk targeting requires the base offloading flag to be enabled, preventing misconfiguration where NVMe paths would be ignored.

Can I use --offload-train with FSDP's native CPU offloading?

No—Miles detects this conflict and automatically disables FSDP's native CPU offload. As implemented in miles/backends/fsdp_utils/actor.py lines 98-100, enabling both would cause double data movement and undefined memory states.

What is the difference between --offload-rollout-level kv_cache and kv_cache weight?

The kv_cache option offloads only the attention cache (typically 10-20% of model memory), while kv_cache weight also offloads model parameters (50-70% additional savings). The latter maximizes GPU availability for training but increases transfer overhead between steps.

How does --stream-optimizer-state-to-disk improve performance?

Rather than blocking the training loop while writing optimizer tensors to NVMe, streaming writes chunks asynchronously using the chunk size from --offload-train-disk-chunk-mb. This overlaps IO with computation, reducing step-time overhead as configured in miles/ray/specs/train.py lines 128-144.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →