Best Practices for Managing Memory Peaks with Miles' Offloading Strategies
Enable both --offload-train and --offload-rollout plus --colocate-memory-peak-device=gpu to serialize GPU memory usage and prevent overlapping spikes that cause OOM crashes.
Miles provides a sophisticated offloading system that moves model parameters, optimizer state, and KV-cache between GPU, CPU, and disk during reinforcement learning training. This article covers proven practices for managing memory peaks when using these offloading strategies, based on the implementation in radixark/miles.
Enable Dual Offloading for Training and Rollout Phases
The foundation of memory peak management is ensuring model weights never reside on GPU when not actively needed.
In main/train.py at lines 32-36, Miles checks both offloading flags before entering the training loop:
# From main/train.py (simplified)
if args.offload_train:
await actor_model.offload() # moves weights/optimizer to CPU/disk
else:
await actor_model.clear_memory() # frees remaining GPU tensors
Best practice: Always enable both flags together:
| Flag | Purpose |
|---|---|
--offload-train |
Moves actor model weights off GPU after each training step |
--offload-rollout |
Clears or offloads tensors before/after rollout generation |
Without both enabled, you risk overlapping memory peaks when training and rollout tensors coexist on GPU.
Co-locate Memory Peaks on GPU with Serialized Scheduling
The --colocate-memory-peak-device=gpu flag is critical for preventing simultaneous memory spikes. When enabled, Miles forces sequential scheduling of all GPU-resident tensors.
This logic appears in main/train.py lines 32-38:
if args.colocate_memory_peak_device == "gpu":
await inference_controller.offload_kv()
await actor_model.onload() # bring model back for rollout
await inference_controller.offload_weights()
else:
# Non-co-located path: selective offloading by tag
offload_tags = [GPU_MEMORY_TYPE_CUDA_GRAPH]
if "kv_cache" in args.offload_rollout_level:
offload_tags.append(GPU_MEMORY_TYPE_KV_CACHE)
if "weight" in args.offload_rollout_level:
offload_tags.append(GPU_MEMORY_TYPE_WEIGHTS)
await inference_controller.offload(tags=offload_tags)
Key insight: Co-location guarantees that weights, KV-cache, and CUDA graphs are never resident simultaneously. The GPU sees only one peak at a time, serialized across phases.
Clear Stale Tensors After Each Rollout
Temporary tensors from generation can linger in GPU memory. Miles addresses this with explicit cleanup at line 45 and lines 45-47:
# From main/train.py
await clear_memory() # explicitly frees rollout temporaries
This call is especially important when --offload-rollout is disabled or when using CUDA graphs that allocate persistent workspace memory.
Enable Granular Offloading Levels for KV-Cache and Weights
Different components consume distinct memory pools. Miles supports independent control via --offload-rollout-level, implemented in main/train.py lines 71-78 and 118-124:
# Offload only KV-cache (keeps weights on GPU for faster onload)
--offload-rollout-level=kv_cache
# Offload both (maximum memory savings)
--offload-rollout-level=kv_cache,weight
Recommendation: Start with kv_cache,weight for maximum safety, then relax to kv_cache only if generation throughput becomes bottlenecked.
Persist to Disk for Multi-Terabyte Models
When CPU RAM is insufficient, Miles supports disk-backed offloading. This is demonstrated in run_inkling.py lines 89-90:
python -m miles.main.train \
--offload-train-target=disk \
--offload-train-disk-dir=/tmp/train_offload
Disk offloading removes the CPU memory constraint entirely. The tradeoff is slower checkpoint I/O—plan /tmp or NVMe storage for acceptable performance.
Offload Optimizer State with --optimizer-cpu-offload
Large optimizers (especially 8-bit Adam or fused variants) can exceed parameter memory. Enable CPU offloading as shown in launch scripts like run_gpt_oss_20b.py at line 67:
--optimizer-cpu-offload
This moves optimizer momentum buffers and state to CPU during backpropagation, preventing GPU memory inflation without affecting training semantics.
Synchronize Weight Updates to Prevent Stale Resident Copies
After each rollout, explicitly synchronize weights to ensure the rollout engine sees updates without keeping duplicate copies. In main/train.py line 150:
await update_weights(actor_model, rollout_executor, rollout_id=rollout_id)
This atomic update prevents race conditions where both old and new weight versions might temporarily coexist in GPU memory.
Validate Offloading with CPU Memory Profiler
Miles includes a dedicated profiling tool to visualize memory peaks and verify offloading behavior.
Run the profiler:
python -m miles.main.tools.cpu_memory_profiler \
--log-dir=/tmp/memory_logs \
--phase=offload,train,rollout
Then generate visualizations:
python -m miles.main.tools.visualize.cpu_memory_profiler_visualize.py \
/tmp/memory_logs
The output shows tagged phases (generate, offload, train) with precise memory deltas. Use this to confirm that:
- Offload phases show expected memory drops
- No unexpected spikes occur during transitions
- Disk offloading achieves lower CPU RAM usage
Avoid Critic Path with GPU Co-Location
Important limitation: GPU co-location is incompatible with the critic path. The assertion at line 36 in main/train.py prevents this configuration:
assert not (args.use_critic and args.colocate_memory_peak_device == "gpu"), \
"Critic path does not support GPU co-location"
If using PPO with a critic model, either:
- Disable co-location and rely on selective tag-based offloading
- Accept higher peak memory usage
Complete Recommended Configuration
For most large-model training scenarios:
python -m miles.main.train \
--offload-train \
--offload-rollout \
--colocate-memory-peak-device=gpu \
--offload-rollout-level=kv_cache,weight \
--optimizer-cpu-offload
For models exceeding CPU RAM, add disk persistence:
--offload-train-target=disk \
--offload-train-disk-dir=/nvme/train_offload
Summary
- Enable both
--offload-trainand--offload-rolloutto ensure weights leave GPU when not needed - Use
--colocate-memory-peak-device=gputo serialize memory peaks and prevent overlapping spikes - Call
clear_memory()after rollouts to free temporary generation tensors - Configure
--offload-rollout-levelfor granular control over KV-cache and weights - Add disk offloading when CPU RAM is insufficient for multi-terabyte models
- Enable
--optimizer-cpu-offloadto prevent optimizer state from inflating GPU usage - Profile with
cpu_memory_profiler.pyto validate expected memory patterns - Avoid critic path with GPU co-location due to known incompatibility
Frequently Asked Questions
What causes OOM errors even with offloading enabled?
OOM errors typically occur when memory peaks overlap—when training tensors arrive before rollout tensors fully depart. Enable --colocate-memory-peak-device=gpu to force serialized scheduling, or verify that --offload-rollout-level includes both kv_cache and weight tags.
How do I know if disk offloading is working correctly?
Run the CPU memory profiler and check that CPU RAM peaks remain flat during offload phases. The visualizer in tools/visualize/cpu_memory_profiler_visualize.py will show offload phases with reduced or flat CPU usage when disk persistence is active.
Can I use GPU co-location with PPO or other critic-based algorithms?
No. The current Miles implementation explicitly blocks this combination with an assertion in main/train.py line 36. When using a critic, disable co-location and rely on tag-based selective offloading instead, accepting moderately higher peak memory usage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →