miles
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
Integrate custom agentic environments with Miles by implementing a generate function and registering it via command-line or dataset configuration. Unlock Miles's full potential today.
Best Practices for Managing Memory Peaks with Miles' Offloading StrategiesMaster memory peaks with Miles' offloading strategies. Learn to serialize GPU memory usage and prevent OOM crashes using --offload-train, --offload-rollout, and --colocate-memory-peak-device=gpu.
How to Configure SGLang for Efficient Rollout in the Miles FrameworkConfigure SGLang for efficient rollout in Miles by passing YAML config and tuning GPU allocation, memory offloading, and LoRA via per-group overrides for optimal performance.
P2P RDMA Weight Transfer Performance in Miles: Benchmarks, Trade‑offs, and ImplementationDiscover P2P RDMA weight transfer performance in Miles. Slash transfer times up to 86% for MoE models and understand benchmarks trade-offs. Optimize your deep learning.
How the Multi-LoRA Async Trainer Differs from the Standard Async Trainer in MilesDiscover how the Miles multi-LoRA async trainer enhances standard async training with dynamic adapter management for seamless on-the-fly registration, retirement, and weight propagation.
How the Mini Fine-Tuning Controller Complements the Main Training Loop in MilesDiscover how the mini fine-tuning controller in Miles complements the main training loop by monitoring rollouts and triggering adapter updates without interrupting data-parallel training.
Asynchronous Rollout and Evaluation Modes in Miles: Beyond the Default PipelineDiscover Miles' async rollout strategies: async rollout, fully-async rollout, and multi-LoRA async. Decouple generation from training to cut wall-time and enable parallel evaluation.
How Ray Placement Groups Manage Distributed Rollout and Training in MilesLearn how Miles leverages Ray placement groups to manage distributed rollout and training. Achieve deterministic execution of RL-HF pipelines with partitioned GPU resources.
Weight Update Validation Options in Miles: `--check-weight-update-equal` and `--check-weight-update-selector` ExplainedExplore Miles weight update validation options --check-weight-update-equal and --check-weight-update-selector. Ensure precise model weight synchronization with these powerful tools.
How the Rollout Executor (Worker Manager) Connects to SGLang Engines in MilesDiscover how the RolloutExecutor connects to SGLang engines in Miles. Learn how this Ray actor manages generation and evaluation requests via HTTP to the SGLang router.
On-Policy Distillation with SGLang in Miles: A Complete Technical GuideExplore on-policy distillation with SGLang in Miles. Learn how token-level knowledge transfer from teacher to student models enhances performance. A complete technical guide.
Agentic Environments in Miles: Harbor, HUD, NeMo Gym & OpenEnv Integration GuideDiscover agentic environments supported by Miles including Harbor, HUD, NeMo Gym, and OpenEnv. Learn how lightweight HTTP adaptors integrate these environments seamlessly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →