miles

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

23 articles 2.6k View on GitHub ↗
23 articles
How to Integrate Custom Agentic Environments with Miles: A Complete Guide

Integrate custom agentic environments with Miles by implementing a generate function and registering it via command-line or dataset configuration. Unlock Miles's full potential today.

how-to-guide
Sep 6, 2026
Best Practices for Managing Memory Peaks with Miles' Offloading Strategies

Master memory peaks with Miles' offloading strategies. Learn to serialize GPU memory usage and prevent OOM crashes using --offload-train, --offload-rollout, and --colocate-memory-peak-device=gpu.

best-practices
Sep 6, 2026
How to Configure SGLang for Efficient Rollout in the Miles Framework

Configure SGLang for efficient rollout in Miles by passing YAML config and tuning GPU allocation, memory offloading, and LoRA via per-group overrides for optimal performance.

how-to-guide
Sep 6, 2026
P2P RDMA Weight Transfer Performance in Miles: Benchmarks, Trade‑offs, and Implementation

Discover P2P RDMA weight transfer performance in Miles. Slash transfer times up to 86% for MoE models and understand benchmarks trade-offs. Optimize your deep learning.

performance
Sep 6, 2026
How the Multi-LoRA Async Trainer Differs from the Standard Async Trainer in Miles

Discover how the Miles multi-LoRA async trainer enhances standard async training with dynamic adapter management for seamless on-the-fly registration, retirement, and weight propagation.

deep-dive
Sep 6, 2026
How the Mini Fine-Tuning Controller Complements the Main Training Loop in Miles

Discover how the mini fine-tuning controller in Miles complements the main training loop by monitoring rollouts and triggering adapter updates without interrupting data-parallel training.

internals
Sep 6, 2026
Asynchronous Rollout and Evaluation Modes in Miles: Beyond the Default Pipeline

Discover Miles' async rollout strategies: async rollout, fully-async rollout, and multi-LoRA async. Decouple generation from training to cut wall-time and enable parallel evaluation.

deep-dive
Sep 6, 2026
How Ray Placement Groups Manage Distributed Rollout and Training in Miles

Learn how Miles leverages Ray placement groups to manage distributed rollout and training. Achieve deterministic execution of RL-HF pipelines with partitioned GPU resources.

architecture
Sep 6, 2026
Weight Update Validation Options in Miles: `--check-weight-update-equal` and `--check-weight-update-selector` Explained

Explore Miles weight update validation options --check-weight-update-equal and --check-weight-update-selector. Ensure precise model weight synchronization with these powerful tools.

how-to-guide
Sep 6, 2026
How the Rollout Executor (Worker Manager) Connects to SGLang Engines in Miles

Discover how the RolloutExecutor connects to SGLang engines in Miles. Learn how this Ray actor manages generation and evaluation requests via HTTP to the SGLang router.

internals
Sep 6, 2026
On-Policy Distillation with SGLang in Miles: A Complete Technical Guide

Explore on-policy distillation with SGLang in Miles. Learn how token-level knowledge transfer from teacher to student models enhances performance. A complete technical guide.

deep-dive
Sep 6, 2026
Agentic Environments in Miles: Harbor, HUD, NeMo Gym & OpenEnv Integration Guide

Discover agentic environments supported by Miles including Harbor, HUD, NeMo Gym, and OpenEnv. Learn how lightweight HTTP adaptors integrate these environments seamlessly.

how-to-guide
Sep 6, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →