# How `--offload-train` and `--offload-rollout` Memory Offloading Works in Miles

> Learn how Miles memory offloading with --offload-train and --offload-rollout trains larger models by moving optimizer state, parameters, and inference tensors off-GPU to CPU or NVMe.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: internals
- Published: 2026-09-06

---

**Miles uses `--offload-train` to move optimizer state and model parameters to CPU or NVMe during training, and `--offload-rollout` to move inference tensors (KV-cache and weights) off-GPU between rollouts, enabling training of models that exceed GPU memory capacity.**

Miles is a distributed training and inference framework that supports **memory offloading strategies** to fit large language models into limited GPU resources. The `--offload-train` and `--offload-rollout` flags work together to swap tensors between GPU memory and slower storage tiers—CPU RAM or NVMe disk—at strategic points in the training loop. This article explains exactly how these mechanisms work, with references to the actual source implementation in `radixark/miles`.

---

## What the Offloading Flags Control

| Flag | Purpose | Default Target |
|------|---------|--------------|
| `--offload-train` | Offload **training tensors** (optimizer state, parameters) when not actively computing gradients | `cpu` (configurable to `disk`) |
| `--offload-rollout` | Offload **rollout tensors** (KV-cache, model weights) when not generating inference outputs | Configurable via `--offload-rollout-level` |

Both flags are **enabled automatically** when `--colocate` mode is active, which runs training and inference on the same GPU. Users can disable them with `--no-offload-train` or `--no-offload-rollout`.

---

## How `--offload-train` Works

### Flag Definition and Validation

The offloading infrastructure is defined in **[`miles/utils/arguments.py`](https://github.com/radixark/miles/blob/main/miles/utils/arguments.py)**. The `--offload-train` flag is part of a grouped option documented as "equivalent to `--offload-train + --offload-rollout`" when both training and rollout memory need management.

```python

# From miles/utils/arguments.py lines 172, 214-224

parser.add_argument("--offload-train", action="store_true", ...)
parser.add_argument("--offload-train-target", choices=["cpu", "disk"], default="cpu")
parser.add_argument("--offload-train-disk-dir", type=str)
parser.add_argument("--offload-train-disk-chunk-mb", type=int, default=64)

```

Critical validation rules are enforced:
- `--offload-train-target=disk` **requires** `--offload-train` to be enabled—this is asserted at lines **3434-3435**
- When `disk` is selected, both `--offload-train-disk-dir` and a positive chunk size must be provided (lines **342-348**)

### Training Loop Integration

In **[`train.py`](https://github.com/radixark/miles/blob/main/train.py)**, the training orchestrator calls the asynchronous helper `offload_train()` before each training step when the flag is true (lines **78-85**):

```python

# From train.py lines 78-85

if args.offload_train:
    await offload_train(
        actor_model=actor_model,
        critic_model=critic_model,
        target=args.offload_train_target,
        disk_dir=args.offload_train_disk_dir,
    )

```

The `offload_train()` coroutine invokes:
- `await actor_model.offload(target)` — moves actor model parameters
- `await critic_model.offload(target)` — moves critic model parameters (when PPO is used)

### Disk Streaming for Optimizer State

When `--stream-optimizer-state-to-disk` is combined with `--offload-train-target=disk`, the optimizer tensors are written to NVMe in chunked streams. The chunk size is controlled by `TMS_DISK_BACKUP_CHUNK_MB`, which is injected into worker environments in **[`miles/ray/specs/train.py`](https://github.com/radixark/miles/blob/main/miles/ray/specs/train.py)** (lines **128-144**):

```python

# From miles/ray/specs/train.py lines 128-144

if args.offload_train and args.offload_train_target == "disk":
    env_vars["TMS_DISK_BACKUP_CHUNK_MB"] = str(args.offload_train_disk_chunk_mb)

```

---

## How `--offload-rollout` Works

### Rollout-Level Configuration

The `--offload-rollout` flag controls which inference tensors leave GPU memory between generation passes. The specific components are selected via `--offload-rollout-level`:

```python

# From train.py lines 118-123

offload_tags = []
if "kv_cache" in args.offload_rollout_level:
    offload_tags.append(GPU_MEMORY_TYPE_KV_CACHE)
if "weight" in args.offload_rollout_level:
    offload_tags.append(GPU_MEMORY_TYPE_WEIGHTS)

```

Common values include:
- `kv_cache` — offload only the attention key-value cache
- `weight` — offload model weights
- `kv_cache weight` — offload both (maximum memory savings)

### InferenceController Offloading Sequence

After each training step, if `--offload-rollout` is enabled, Miles delegates to the **`InferenceController`**, which performs offloading in two phases:

1. **KV-cache offload** — `controller.offload_kv()` moves attention caches
2. **Weight offload** — `controller.offload_weights()` moves model parameters

This sequencing ensures that active generation contexts are preserved until the rollout fully completes, then aggressively clears GPU memory before the next training step begins.

---

## Backend-Specific Implementation Details

### Megatron Backend

Megatron actors implement the actual `offload(target)` calls in **[`miles/backends/megatron_utils/actor.py`](https://github.com/radixark/miles/blob/main/miles/backends/megatron_utils/actor.py)**. The target parameter (`"cpu"` or `"disk"`) determines whether tensors move to pinned host memory or are serialized to the NVMe path specified in `--offload-train-disk-dir`.

### FSDP Backend Conflict Resolution

FSDP (Fully Sharded Data Parallel) has its own CPU offloading mechanism, which conflicts with Miles' explicit offloading. In **[`miles/backends/fsdp_utils/actor.py`](https://github.com/radixark/miles/blob/main/miles/backends/fsdp_utils/actor.py)** (lines **98-100**), Miles automatically detects and disables FSDP's native CPU offload when `--offload-train` is enabled:

```python

# From miles/backends/fsdp_utils/actor.py lines 98-100

if args.offload_train and self.use_fsdp_cpu_offload:
    logger.warning("Disabling FSDP CPU offload due to Miles offloading")
    self.use_fsdp_cpu_offload = False

```

This prevents double-offloading overhead and ensures deterministic memory management.

---

## Colocate Mode Requirements

When `--colocate` is specified—running training and inference on the same GPU—the memory offloading strategy becomes **mandatory**. The validator at line **35** of [`train.py`](https://github.com/radixark/miles/blob/main/train.py) enforces that both flags must be true:

```python

# From train.py line 35

if args.colocate:
    assert args.offload_train, "colocate requires --offload-train"
    assert args.offload_rollout, "colocate requires --offload-rollout"

```

This requirement exists because a single GPU cannot simultaneously hold optimizer state, training parameters, model weights, and KV-cache for large models. The peak memory would exceed capacity without aggressive offloading.

---

## Practical Configuration Examples

### CPU-Based Offloading (Default)

```bash
python -m miles.main.train \
    --train-backend megatron \
    --colocate \
    --offload-train \
    --offload-train-target cpu \
    --offload-rollout \
    --offload-rollout-level kv_cache weight

```

### NVMe Disk Offloading with Streaming

```bash
python -m miles.main.train \
    --train-backend megatron \
    --colocate \
    --offload-train \
    --offload-train-target disk \
    --offload-train-disk-dir /scratch/miles_offload \
    --offload-train-disk-chunk-mb 256 \
    --stream-optimizer-state-to-disk \
    --offload-rollout \
    --offload-rollout-level kv_cache weight

```

**Key parameters:**
- `--offload-train-disk-chunk-mb` — controls granularity of NVMe writes (larger = fewer files, smaller = better overlap)
- `--stream-optimizer-state-to-disk` — enables asynchronous optimizer state serialization to prevent blocking

---

## Summary

- **`--offload-train`** moves training tensors (optimizer state, parameters) to CPU or NVMe when gradients are not being computed, implemented via `offload_train()` in [`train.py`](https://github.com/radixark/miles/blob/main/train.py) and `model.offload()` in Megatron actors
- **`--offload-rollout`** moves inference tensors (KV-cache, weights) off-GPU between rollouts, controlled by `--offload-rollout-level` and executed through `InferenceController`
- Both flags are **required** in `--colocate` mode to fit large models on single GPUs, enforced by assertions at `train.py:35`
- Target selection (`cpu` vs `disk`) and chunk sizing are validated in [`miles/utils/arguments.py`](https://github.com/radixark/miles/blob/main/miles/utils/arguments.py) lines **342-348** and **3434-3435**
- FSDP backend automatically disables native CPU offload to avoid conflicts, as shown in [`miles/backends/fsdp_utils/actor.py`](https://github.com/radixark/miles/blob/main/miles/backends/fsdp_utils/actor.py) lines **98-100**

---

## Frequently Asked Questions

### What happens if I enable `--offload-train-target=disk` without specifying `--offload-train`?

The argument parser raises an error. According to [`miles/utils/arguments.py`](https://github.com/radixark/miles/blob/main/miles/utils/arguments.py) lines **3434-3435**, disk targeting requires the base offloading flag to be enabled, preventing misconfiguration where NVMe paths would be ignored.

### Can I use `--offload-train` with FSDP's native CPU offloading?

No—Miles detects this conflict and automatically disables FSDP's native CPU offload. As implemented in [`miles/backends/fsdp_utils/actor.py`](https://github.com/radixark/miles/blob/main/miles/backends/fsdp_utils/actor.py) lines **98-100**, enabling both would cause double data movement and undefined memory states.

### What is the difference between `--offload-rollout-level kv_cache` and `kv_cache weight`?

The `kv_cache` option offloads only the attention cache (typically 10-20% of model memory), while `kv_cache weight` also offloads model parameters (50-70% additional savings). The latter maximizes GPU availability for training but increases transfer overhead between steps.

### How does `--stream-optimizer-state-to-disk` improve performance?

Rather than blocking the training loop while writing optimizer tensors to NVMe, streaming writes chunks asynchronously using the chunk size from `--offload-train-disk-chunk-mb`. This overlaps IO with computation, reducing step-time overhead as configured in [`miles/ray/specs/train.py`](https://github.com/radixark/miles/blob/main/miles/ray/specs/train.py) lines **128-144**.