# Best Practices for Managing Memory Peaks with Miles' Offloading Strategies

> Master memory peaks with Miles' offloading strategies. Learn to serialize GPU memory usage and prevent OOM crashes using --offload-train, --offload-rollout, and --colocate-memory-peak-device=gpu.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: best-practices
- Published: 2026-09-06

---

**Enable both `--offload-train` and `--offload-rollout` plus `--colocate-memory-peak-device=gpu` to serialize GPU memory usage and prevent overlapping spikes that cause OOM crashes.**

Miles provides a sophisticated **offloading system** that moves model parameters, optimizer state, and KV-cache between GPU, CPU, and disk during reinforcement learning training. This article covers proven practices for **managing memory peaks** when using these offloading strategies, based on the implementation in `radixark/miles`.

## Enable Dual Offloading for Training and Rollout Phases

The foundation of **memory peak management** is ensuring model weights never reside on GPU when not actively needed.

In [`main/train.py`](https://github.com/radixark/miles/blob/main/main/train.py) at lines [32-36](https://github.com/radixark/miles/blob/main/main/train.py#L32-L36), Miles checks both offloading flags before entering the training loop:

```python

# From main/train.py (simplified)

if args.offload_train:
    await actor_model.offload()      # moves weights/optimizer to CPU/disk

else:
    await actor_model.clear_memory() # frees remaining GPU tensors

```

**Best practice:** Always enable both flags together:

| Flag | Purpose |
|------|---------|
| `--offload-train` | Moves actor model weights off GPU after each training step |
| `--offload-rollout` | Clears or offloads tensors before/after rollout generation |

Without both enabled, you risk overlapping memory peaks when training and rollout tensors coexist on GPU.

## Co-locate Memory Peaks on GPU with Serialized Scheduling

The `--colocate-memory-peak-device=gpu` flag is critical for **preventing simultaneous memory spikes**. When enabled, Miles forces sequential scheduling of all GPU-resident tensors.

This logic appears in [`main/train.py`](https://github.com/radixark/miles/blob/main/main/train.py) lines [32-38](https://github.com/radixark/miles/blob/main/main/train.py#L32-L38):

```python
if args.colocate_memory_peak_device == "gpu":
    await inference_controller.offload_kv()
    await actor_model.onload()          # bring model back for rollout

    await inference_controller.offload_weights()
else:
    # Non-co-located path: selective offloading by tag

    offload_tags = [GPU_MEMORY_TYPE_CUDA_GRAPH]
    if "kv_cache" in args.offload_rollout_level:
        offload_tags.append(GPU_MEMORY_TYPE_KV_CACHE)
    if "weight" in args.offload_rollout_level:
        offload_tags.append(GPU_MEMORY_TYPE_WEIGHTS)
    await inference_controller.offload(tags=offload_tags)

```

**Key insight:** Co-location guarantees that weights, KV-cache, and CUDA graphs are never resident simultaneously. The GPU sees only one peak at a time, serialized across phases.

## Clear Stale Tensors After Each Rollout

Temporary tensors from generation can linger in GPU memory. Miles addresses this with explicit cleanup at line [45](https://github.com/radixark/miles/blob/main/main/train.py#L45) and lines [45-47](https://github.com/radixark/miles/blob/main/main/train.py#L45-L47):

```python

# From main/train.py

await clear_memory()  # explicitly frees rollout temporaries

```

This call is especially important when `--offload-rollout` is disabled or when using CUDA graphs that allocate persistent workspace memory.

## Enable Granular Offloading Levels for KV-Cache and Weights

Different components consume distinct memory pools. Miles supports independent control via `--offload-rollout-level`, implemented in [`main/train.py`](https://github.com/radixark/miles/blob/main/main/train.py) lines [71-78](https://github.com/radixark/miles/blob/main/main/train.py#L71-L78) and [118-124](https://github.com/radixark/miles/blob/main/main/train.py#L118-L124):

```bash

# Offload only KV-cache (keeps weights on GPU for faster onload)

--offload-rollout-level=kv_cache

# Offload both (maximum memory savings)

--offload-rollout-level=kv_cache,weight

```

**Recommendation:** Start with `kv_cache,weight` for maximum safety, then relax to `kv_cache` only if generation throughput becomes bottlenecked.

## Persist to Disk for Multi-Terabyte Models

When CPU RAM is insufficient, Miles supports **disk-backed offloading**. This is demonstrated in [`run_inkling.py`](https://github.com/radixark/miles/blob/main/run_inkling.py) lines [89-90](https://github.com/radixark/miles/blob/main/main/tests/snapshots/launch_scripts/py/scripts/run_inkling.py#L89-L90):

```bash
python -m miles.main.train \
    --offload-train-target=disk \
    --offload-train-disk-dir=/tmp/train_offload

```

Disk offloading removes the CPU memory constraint entirely. The tradeoff is slower checkpoint I/O—plan `/tmp` or NVMe storage for acceptable performance.

## Offload Optimizer State with `--optimizer-cpu-offload`

Large optimizers (especially 8-bit Adam or fused variants) can exceed parameter memory. Enable CPU offloading as shown in launch scripts like [`run_gpt_oss_20b.py`](https://github.com/radixark/miles/blob/main/run_gpt_oss_20b.py) at line [67](https://github.com/radixark/miles/blob/main/main/tests/snapshots/launch_scripts/py/scripts/run_gpt_oss_20b.py#L67):

```bash
--optimizer-cpu-offload

```

This moves optimizer momentum buffers and state to CPU during backpropagation, preventing GPU memory inflation without affecting training semantics.

## Synchronize Weight Updates to Prevent Stale Resident Copies

After each rollout, explicitly synchronize weights to ensure the rollout engine sees updates without keeping duplicate copies. In [`main/train.py`](https://github.com/radixark/miles/blob/main/main/train.py) line [150](https://github.com/radixark/miles/blob/main/main/train.py#L150):

```python
await update_weights(actor_model, rollout_executor, rollout_id=rollout_id)

```

This atomic update prevents race conditions where both old and new weight versions might temporarily coexist in GPU memory.

## Validate Offloading with CPU Memory Profiler

Miles includes a **dedicated profiling tool** to visualize memory peaks and verify offloading behavior.

Run the profiler:

```bash
python -m miles.main.tools.cpu_memory_profiler \
    --log-dir=/tmp/memory_logs \
    --phase=offload,train,rollout

```

Then generate visualizations:

```bash
python -m miles.main.tools.visualize.cpu_memory_profiler_visualize.py \
    /tmp/memory_logs

```

The output shows tagged phases (`generate`, `offload`, `train`) with precise memory deltas. Use this to confirm that:
- Offload phases show expected memory drops
- No unexpected spikes occur during transitions
- Disk offloading achieves lower CPU RAM usage

## Avoid Critic Path with GPU Co-Location

**Important limitation:** GPU co-location is incompatible with the critic path. The assertion at line [36](https://github.com/radixark/miles/blob/main/main/train.py#L36) in [`main/train.py`](https://github.com/radixark/miles/blob/main/main/train.py) prevents this configuration:

```python
assert not (args.use_critic and args.colocate_memory_peak_device == "gpu"), \
    "Critic path does not support GPU co-location"

```

If using PPO with a critic model, either:
- Disable co-location and rely on selective tag-based offloading
- Accept higher peak memory usage

## Complete Recommended Configuration

For most large-model training scenarios:

```bash
python -m miles.main.train \
    --offload-train \
    --offload-rollout \
    --colocate-memory-peak-device=gpu \
    --offload-rollout-level=kv_cache,weight \
    --optimizer-cpu-offload

```

For models exceeding CPU RAM, add disk persistence:

```bash
    --offload-train-target=disk \
    --offload-train-disk-dir=/nvme/train_offload

```

## Summary

- **Enable both `--offload-train` and `--offload-rollout`** to ensure weights leave GPU when not needed
- **Use `--colocate-memory-peak-device=gpu`** to serialize memory peaks and prevent overlapping spikes
- **Call `clear_memory()` after rollouts** to free temporary generation tensors
- **Configure `--offload-rollout-level`** for granular control over KV-cache and weights
- **Add disk offloading** when CPU RAM is insufficient for multi-terabyte models
- **Enable `--optimizer-cpu-offload`** to prevent optimizer state from inflating GPU usage
- **Profile with [`cpu_memory_profiler.py`](https://github.com/radixark/miles/blob/main/cpu_memory_profiler.py)** to validate expected memory patterns
- **Avoid critic path with GPU co-location** due to known incompatibility

## Frequently Asked Questions

### What causes OOM errors even with offloading enabled?

OOM errors typically occur when memory peaks overlap—when training tensors arrive before rollout tensors fully depart. **Enable `--colocate-memory-peak-device=gpu`** to force serialized scheduling, or verify that `--offload-rollout-level` includes both `kv_cache` and `weight` tags.

### How do I know if disk offloading is working correctly?

Run the **CPU memory profiler** and check that CPU RAM peaks remain flat during offload phases. The visualizer in [`tools/visualize/cpu_memory_profiler_visualize.py`](https://github.com/radixark/miles/blob/main/tools/visualize/cpu_memory_profiler_visualize.py) will show `offload` phases with reduced or flat CPU usage when disk persistence is active.

### Can I use GPU co-location with PPO or other critic-based algorithms?

**No.** The current Miles implementation explicitly blocks this combination with an assertion in [`main/train.py`](https://github.com/radixark/miles/blob/main/main/train.py) line [36](https://github.com/radixark/miles/blob/main/main/train.py#L36). When using a critic, disable co-location and rely on tag-based selective offloading instead, accepting moderately higher peak memory usage.