How `--offload-train` and `--offload-rollout` Memory Offloading Works in Miles
Miles uses --offload-train to move optimizer state and model parameters to CPU or NVMe during training, and --offload-rollout to move inference tensors (KV-cache and weights) off-GPU between rollouts, enabling training of models that exceed GPU memory capacity.
Miles is a distributed training and inference framework that supports memory offloading strategies to fit large language models into limited GPU resources. The --offload-train and --offload-rollout flags work together to swap tensors between GPU memory and slower storage tiers—CPU RAM or NVMe disk—at strategic points in the training loop. This article explains exactly how these mechanisms work, with references to the actual source implementation in radixark/miles.
What the Offloading Flags Control
| Flag | Purpose | Default Target |
|---|---|---|
--offload-train |
Offload training tensors (optimizer state, parameters) when not actively computing gradients | cpu (configurable to disk) |
--offload-rollout |
Offload rollout tensors (KV-cache, model weights) when not generating inference outputs | Configurable via --offload-rollout-level |
Both flags are enabled automatically when --colocate mode is active, which runs training and inference on the same GPU. Users can disable them with --no-offload-train or --no-offload-rollout.
How --offload-train Works
Flag Definition and Validation
The offloading infrastructure is defined in miles/utils/arguments.py. The --offload-train flag is part of a grouped option documented as "equivalent to --offload-train + --offload-rollout" when both training and rollout memory need management.
# From miles/utils/arguments.py lines 172, 214-224
parser.add_argument("--offload-train", action="store_true", ...)
parser.add_argument("--offload-train-target", choices=["cpu", "disk"], default="cpu")
parser.add_argument("--offload-train-disk-dir", type=str)
parser.add_argument("--offload-train-disk-chunk-mb", type=int, default=64)
Critical validation rules are enforced:
--offload-train-target=diskrequires--offload-trainto be enabled—this is asserted at lines 3434-3435- When
diskis selected, both--offload-train-disk-dirand a positive chunk size must be provided (lines 342-348)
Training Loop Integration
In train.py, the training orchestrator calls the asynchronous helper offload_train() before each training step when the flag is true (lines 78-85):
# From train.py lines 78-85
if args.offload_train:
await offload_train(
actor_model=actor_model,
critic_model=critic_model,
target=args.offload_train_target,
disk_dir=args.offload_train_disk_dir,
)
The offload_train() coroutine invokes:
await actor_model.offload(target)— moves actor model parametersawait critic_model.offload(target)— moves critic model parameters (when PPO is used)
Disk Streaming for Optimizer State
When --stream-optimizer-state-to-disk is combined with --offload-train-target=disk, the optimizer tensors are written to NVMe in chunked streams. The chunk size is controlled by TMS_DISK_BACKUP_CHUNK_MB, which is injected into worker environments in miles/ray/specs/train.py (lines 128-144):
# From miles/ray/specs/train.py lines 128-144
if args.offload_train and args.offload_train_target == "disk":
env_vars["TMS_DISK_BACKUP_CHUNK_MB"] = str(args.offload_train_disk_chunk_mb)
How --offload-rollout Works
Rollout-Level Configuration
The --offload-rollout flag controls which inference tensors leave GPU memory between generation passes. The specific components are selected via --offload-rollout-level:
# From train.py lines 118-123
offload_tags = []
if "kv_cache" in args.offload_rollout_level:
offload_tags.append(GPU_MEMORY_TYPE_KV_CACHE)
if "weight" in args.offload_rollout_level:
offload_tags.append(GPU_MEMORY_TYPE_WEIGHTS)
Common values include:
kv_cache— offload only the attention key-value cacheweight— offload model weightskv_cache weight— offload both (maximum memory savings)
InferenceController Offloading Sequence
After each training step, if --offload-rollout is enabled, Miles delegates to the InferenceController, which performs offloading in two phases:
- KV-cache offload —
controller.offload_kv()moves attention caches - Weight offload —
controller.offload_weights()moves model parameters
This sequencing ensures that active generation contexts are preserved until the rollout fully completes, then aggressively clears GPU memory before the next training step begins.
Backend-Specific Implementation Details
Megatron Backend
Megatron actors implement the actual offload(target) calls in miles/backends/megatron_utils/actor.py. The target parameter ("cpu" or "disk") determines whether tensors move to pinned host memory or are serialized to the NVMe path specified in --offload-train-disk-dir.
FSDP Backend Conflict Resolution
FSDP (Fully Sharded Data Parallel) has its own CPU offloading mechanism, which conflicts with Miles' explicit offloading. In miles/backends/fsdp_utils/actor.py (lines 98-100), Miles automatically detects and disables FSDP's native CPU offload when --offload-train is enabled:
# From miles/backends/fsdp_utils/actor.py lines 98-100
if args.offload_train and self.use_fsdp_cpu_offload:
logger.warning("Disabling FSDP CPU offload due to Miles offloading")
self.use_fsdp_cpu_offload = False
This prevents double-offloading overhead and ensures deterministic memory management.
Colocate Mode Requirements
When --colocate is specified—running training and inference on the same GPU—the memory offloading strategy becomes mandatory. The validator at line 35 of train.py enforces that both flags must be true:
# From train.py line 35
if args.colocate:
assert args.offload_train, "colocate requires --offload-train"
assert args.offload_rollout, "colocate requires --offload-rollout"
This requirement exists because a single GPU cannot simultaneously hold optimizer state, training parameters, model weights, and KV-cache for large models. The peak memory would exceed capacity without aggressive offloading.
Practical Configuration Examples
CPU-Based Offloading (Default)
python -m miles.main.train \
--train-backend megatron \
--colocate \
--offload-train \
--offload-train-target cpu \
--offload-rollout \
--offload-rollout-level kv_cache weight
NVMe Disk Offloading with Streaming
python -m miles.main.train \
--train-backend megatron \
--colocate \
--offload-train \
--offload-train-target disk \
--offload-train-disk-dir /scratch/miles_offload \
--offload-train-disk-chunk-mb 256 \
--stream-optimizer-state-to-disk \
--offload-rollout \
--offload-rollout-level kv_cache weight
Key parameters:
--offload-train-disk-chunk-mb— controls granularity of NVMe writes (larger = fewer files, smaller = better overlap)--stream-optimizer-state-to-disk— enables asynchronous optimizer state serialization to prevent blocking
Summary
--offload-trainmoves training tensors (optimizer state, parameters) to CPU or NVMe when gradients are not being computed, implemented viaoffload_train()intrain.pyandmodel.offload()in Megatron actors--offload-rolloutmoves inference tensors (KV-cache, weights) off-GPU between rollouts, controlled by--offload-rollout-leveland executed throughInferenceController- Both flags are required in
--colocatemode to fit large models on single GPUs, enforced by assertions attrain.py:35 - Target selection (
cpuvsdisk) and chunk sizing are validated inmiles/utils/arguments.pylines 342-348 and 3434-3435 - FSDP backend automatically disables native CPU offload to avoid conflicts, as shown in
miles/backends/fsdp_utils/actor.pylines 98-100
Frequently Asked Questions
What happens if I enable --offload-train-target=disk without specifying --offload-train?
The argument parser raises an error. According to miles/utils/arguments.py lines 3434-3435, disk targeting requires the base offloading flag to be enabled, preventing misconfiguration where NVMe paths would be ignored.
Can I use --offload-train with FSDP's native CPU offloading?
No—Miles detects this conflict and automatically disables FSDP's native CPU offload. As implemented in miles/backends/fsdp_utils/actor.py lines 98-100, enabling both would cause double data movement and undefined memory states.
What is the difference between --offload-rollout-level kv_cache and kv_cache weight?
The kv_cache option offloads only the attention cache (typically 10-20% of model memory), while kv_cache weight also offloads model parameters (50-70% additional savings). The latter maximizes GPU availability for training but increases transfer overhead between steps.
How does --stream-optimizer-state-to-disk improve performance?
Rather than blocking the training loop while writing optimizer tensors to NVMe, streaming writes chunks asynchronously using the chunk size from --offload-train-disk-chunk-mb. This overlaps IO with computation, reducing step-time overhead as configured in miles/ray/specs/train.py lines 128-144.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →