# miles | RadixArk | Knowledge Base | Instagit

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

GitHub Stars: 2.6k

Repository: https://github.com/radixark/miles

---

## Articles

### [How to Integrate Custom Agentic Environments with Miles: A Complete Guide](/radixark/miles/miles-integrate-custom-agentic-environments)

Integrate custom agentic environments with Miles by implementing a generate function and registering it via command-line or dataset configuration. Unlock Miles's full potential today.

- Tags: how-to-guide
- Published: 2026-09-06

### [Best Practices for Managing Memory Peaks with Miles' Offloading Strategies](/radixark/miles/miles-memory-peak-management-offloading-best-practices)

Master memory peaks with Miles' offloading strategies. Learn to serialize GPU memory usage and prevent OOM crashes using --offload-train, --offload-rollout, and --colocate-memory-peak-device=gpu.

- Tags: best-practices
- Published: 2026-09-06

### [How to Configure SGLang for Efficient Rollout in the Miles Framework](/radixark/miles/configure-sglang-efficient-rollout-miles)

Configure SGLang for efficient rollout in Miles by passing YAML config and tuning GPU allocation, memory offloading, and LoRA via per-group overrides for optimal performance.

- Tags: how-to-guide
- Published: 2026-09-06

### [P2P RDMA Weight Transfer Performance in Miles: Benchmarks, Trade‑offs, and Implementation](/radixark/miles/miles-p2p-rdma-weight-transfer-performance)

Discover P2P RDMA weight transfer performance in Miles. Slash transfer times up to 86% for MoE models and understand benchmarks trade-offs. Optimize your deep learning.

- Tags: performance
- Published: 2026-09-06

### [How the Multi-LoRA Async Trainer Differs from the Standard Async Trainer in Miles](/radixark/miles/miles-multi-lora-async-vs-standard-trainer)

Discover how the Miles multi-LoRA async trainer enhances standard async training with dynamic adapter management for seamless on-the-fly registration, retirement, and weight propagation.

- Tags: deep-dive
- Published: 2026-09-06

### [How the Mini Fine-Tuning Controller Complements the Main Training Loop in Miles](/radixark/miles/miles-mini-fine-tuning-controller-training-loop)

Discover how the mini fine-tuning controller in Miles complements the main training loop by monitoring rollouts and triggering adapter updates without interrupting data-parallel training.

- Tags: internals
- Published: 2026-09-06

### [Asynchronous Rollout and Evaluation Modes in Miles: Beyond the Default Pipeline](/radixark/miles/miles-async-modes-rollout-evaluation)

Discover Miles' async rollout strategies: async rollout, fully-async rollout, and multi-LoRA async. Decouple generation from training to cut wall-time and enable parallel evaluation.

- Tags: deep-dive
- Published: 2026-09-06

### [How Ray Placement Groups Manage Distributed Rollout and Training in Miles](/radixark/miles/miles-ray-placement-group-distributed-management)

Learn how Miles leverages Ray placement groups to manage distributed rollout and training. Achieve deterministic execution of RL-HF pipelines with partitioned GPU resources.

- Tags: architecture
- Published: 2026-09-06

### [Weight Update Validation Options in Miles: `--check-weight-update-equal` and `--check-weight-update-selector` Explained](/radixark/miles/miles-weight-update-validation-options)

Explore Miles weight update validation options --check-weight-update-equal and --check-weight-update-selector. Ensure precise model weight synchronization with these powerful tools.

- Tags: how-to-guide
- Published: 2026-09-06

### [How the Rollout Executor (Worker Manager) Connects to SGLang Engines in Miles](/radixark/miles/miles-rollout-executor-sglang-engine-connection)

Discover how the RolloutExecutor connects to SGLang engines in Miles. Learn how this Ray actor manages generation and evaluation requests via HTTP to the SGLang router.

- Tags: internals
- Published: 2026-09-06

### [On-Policy Distillation with SGLang in Miles: A Complete Technical Guide](/radixark/miles/on-policy-distillation-sglang-miles)

Explore on-policy distillation with SGLang in Miles. Learn how token-level knowledge transfer from teacher to student models enhances performance. A complete technical guide.

- Tags: deep-dive
- Published: 2026-09-06

### [Agentic Environments in Miles: Harbor, HUD, NeMo Gym & OpenEnv Integration Guide](/radixark/miles/miles-supported-agentic-environments-integration)

Discover agentic environments supported by Miles including Harbor, HUD, NeMo Gym, and OpenEnv. Learn how lightweight HTTP adaptors integrate these environments seamlessly.

- Tags: how-to-guide
- Published: 2026-09-06

### [How `--colocate-memory-peak-device gpu` Optimizes Memory Allocation in Miles](/radixark/miles/miles-colocate-memory-peak-device-gpu-optimization)

Discover how the --colocate-memory-peak-device gpu option in Miles optimizes GPU memory by offloading KV-cache and weights to CPU, significantly reducing peak consumption for larger models.

- Tags: performance
- Published: 2026-09-06

### [How GRPO, GSPO, PPO, and REINFORCE++ Differ in Implementation: A Deep Dive into the Miles Codebase](/radixark/miles/grpo-gspo-ppo-reinforce-implementation-miles)

Explore GRPO, GSPO, PPO, and REINFORCE++ implementation differences in Miles. Understand their core mechanics for advanced reinforcement learning.

- Tags: deep-dive
- Published: 2026-09-06

### [How the InferenceController Coordinates Rollout and Training in Miles](/radixark/miles/miles-inference-controller-rollout-training-coordination)

Discover how the InferenceController orchestrates continuous training-while-serving in Miles. Learn about its role in exclusive weight updates, checkpoint synchronization, and RolloutServer coordination.

- Tags: internals
- Published: 2026-09-06

### [How `--offload-train` and `--offload-rollout` Memory Offloading Works in Miles](/radixark/miles/miles-memory-offloading-offload-train-rollout)

Learn how Miles memory offloading with --offload-train and --offload-rollout trains larger models by moving optimizer state, parameters, and inference tensors off-GPU to CPU or NVMe.

- Tags: internals
- Published: 2026-09-06

### [How Miles Handles Fault Tolerance for SGLang Engine Failures and Run Recovery](/radixark/miles/miles-fault-tolerance-sglang-engine-recovery)

Learn how Miles ensures fault tolerance for SGLang engine failures. Discover automatic crash recovery, actor restarts, weight reloads, and request replays for uninterrupted training and inference.

- Tags: how-to-guide
- Published: 2026-09-06

### [Token-in-Token-Out (TITO) in Miles: How Incremental Tokenization Works Across Model Architectures](/radixark/miles/token-in-token-out-tito-miles-architectures)

Discover Token-in-Token-Out (TITO) in Miles. Learn how this incremental tokenization unifies chat histories across model architectures by merging prefix tokens with new messages.

- Tags: deep-dive
- Published: 2026-09-06

### [How LoRA and Multi-LoRA Training Work with SGLang Rollout Integration in Miles](/radixark/miles/lora-multi-lora-training-sglang-integration)

Discover how Miles integrates LoRA and Multi-LoRA training with SGLang. Learn about dual-mode architecture, async producer-consumer patterns, and dynamic staleness filtering for efficient rollouts.

- Tags: deep-dive
- Published: 2026-09-06

### [Miles Low‑Precision Formats: MXFP8, NVFP4, FP8, and INT4 QAT Stability Guide](/radixark/miles/miles-supported-low-precision-formats-stability)

Explore Miles low-precision formats MXFP8 NVFP4 FP8 INT4. Understand stability guarantees from production-ready to experimental for efficient AI.

- Tags: guide
- Published: 2026-09-06

### [How P2P RDMA Weight Transfer Enables Fast In-Loop Model Updates in Miles](/radixark/miles/p2p-rdma-weight-transfer-fast-model-updates)

Discover how P2P RDMA weight transfer in Miles accelerates in-loop model updates by reducing per-step latency from hundreds to 30-50ms for faster training.

- Tags: deep-dive
- Published: 2026-09-06

### [Rollout Routing Replay (R3) in Miles: Preventing MoE Routing Mismatch for Stable RL Training](/radixark/miles/rollout-routing-replay-r3-mechanism-miles)

Discover Miles Rollout Routing Replay (R3) to prevent MoE routing mismatch in RL training. R3 ensures stable training by replaying expert decisions, eliminating stochastic mismatches caused by quantization and non-deterministic...

- Tags: deep-dive
- Published: 2026-09-06

### [How the Fully Async RL Pipeline in Miles Decouples Rollout and Training Workers](/radixark/miles/fully-async-rl-pipeline-miles-decoupling)

Discover how the Miles RL pipeline decouples rollout and training workers using a persistent background Ray actor and queue based buffer for efficient asynchronous data fetching and training.

- Tags: internals
- Published: 2026-09-06

