How the Mini Fine-Tuning Controller Complements the Main Training Loop in Miles

The Mini Fine-Tuning Controller runs as a daemon thread alongside the main training loop, monitoring rollout health signals and triggering lightweight adapter updates without interrupting data-parallel training.

In distributed LLM training, fault tolerance and rapid adaptation are critical. The Miles repository implements a Mini Fine-Tuning (FT) Controller that integrates directly with the primary training pipeline to enable on-demand adapter fine-tuning. This article examines how maybe_start_mini_ft_controller() bridges the main training loop in train.py and train_async.py with responsive, low-overhead model updates.


What the Mini FT Controller Does

The Mini FT Controller serves two complementary functions that keep long-running training jobs resilient:

Function Implementation Impact on Training
Health signal polling A daemon thread queries the local API server at mini_ft_controller_poll_interval for "heal" signals from the rollout subsystem Zero interruption to gradient computation; asynchronous detection of rollout failures
Reactive adapter fine-tuning Upon signal detection, the controller launches short-lived LoRA/IA³ updates via _MiniFTControllerRunner Model reference remains stable; only adapter weights change, preserving training throughput

This architecture separates concerns: the main loop handles data parallelism and gradient aggregation, while the controller manages fine-grained model health without process restarts.


Integration Point: maybe_start_mini_ft_controller()

The controller hooks into training through a single entry point. In miles/main/train.py and miles/main/train_async.py, the startup sequence follows this pattern:

from miles.utils.ft_utils.mini_ft_controller import maybe_start_mini_ft_controller

def main():
    args = parse_args()
    
    # Initialize controller before heavy training setup

    mini_ft_handle = maybe_start_mini_ft_controller(args)
    
    # Proceed with standard training pipeline

    trainer = Trainer(args)
    trainer.train()

The maybe_start_mini_ft_controller() function checks activation flags and either returns a handle to the running controller or None if disabled. This design keeps the integration non-invasive—training logic requires no conditional branches to accommodate the controller.


Activation and Configuration

The controller respects explicit CLI flags defined in the argument parser of train.py:

Flag Default Purpose
--mini-ft-controller-enable False (enabled implicitly via --use-fault-tolerance) Explicit activation
--no-mini-ft-controller-enable — Force disable
--mini-ft-controller-poll-interval 1.0 Seconds between health checks
--mini-ft-controller-resume-delay 0.0 Cooldown after fine-tuning before resuming rollouts
python -m miles.main.train \
    --model qwen3-5-35b \
    --num-rollout 4 \
    --use-fault-tolerance

The --use-fault-tolerance flag implicitly enables the Mini FT Controller with sensible defaults for production deployments.

Explicit Tuning for Async Training

python -m miles.main.train_async \
    --model qwen3-6-35b-a3b \
    --mini-ft-controller-enable \
    --mini-ft-controller-poll-interval 2.0 \
    --mini-ft-controller-resume-delay 5.0

This configuration polls every 2 seconds and imposes a 5-second stabilization period post-healing—useful when rollout variance is high.

Complete Disable

python -m miles.main.train \
    --model gemma-4-31b \
    --no-mini-ft-controller-enable

The maybe_start_mini_ft_controller() call returns immediately, and training proceeds without sidecar monitoring.


Internal Architecture: mini_ft_controller.py

The implementation in miles/utils/ft_utils/mini_ft_controller.py organizes functionality into discrete classes:

  • _MiniFTControllerRunner — Encapsulates the daemon thread lifecycle and signal processing loop
  • Health check logic — HTTP requests to localhost:{args.api_server_port}/health or equivalent endpoint
  • Adapter update dispatch — Invokes the fine-tuning pipeline with frozen base model weights

Key design decisions visible in the source:

  1. Thread-based, not process-based — Shares Python interpreter state with the trainer; no serialization overhead
  2. Signal-driven, not timer-driven — Polling interval only bounds maximum latency; actual response is event-triggered
  3. Graceful degradation — Controller exceptions are caught and logged; training continues even if healing fails

Testing and Verification

The repository includes targeted tests in tests/fast/utils/test_mini_ft_controller.py that validate:

  • Correct startup/shutdown behavior under various flag combinations
  • Thread safety of shared model references
  • Proper handling of malformed heal signals

Integration with the API server is exercised in tests/fast/api_server/test_handles.py, which defines the signal schema that production controllers consume.


Summary

  • The Mini FT Controller runs as a daemon thread spawned by maybe_start_mini_ft_controller() in train.py and train_async.py
  • It polls for health signals at configurable intervals without blocking gradient computation
  • Upon detection, it executes lightweight adapter fine-tuning (LoRA/IA³) while preserving the base model reference
  • Activation is controlled via --use-fault-tolerance or explicit --mini-ft-controller-enable flags
  • The implementation in miles/utils/ft_utils/mini_ft_controller.py prioritizes zero-overhead integration through thread-sharing and defensive error handling

Frequently Asked Questions

What happens if the Mini FT Controller crashes during training?

The controller wraps its main loop in exception handlers that log errors without propagating to the parent thread. Training continues uninterrupted; the model simply won't receive automatic healing updates until restart. Monitor logs for MiniFTController errors to detect degradation.

Can I use the Mini FT Controller with custom adapter types beyond LoRA?

The controller delegates actual fine-tuning to Miles' internal adapter system. Any adapter registered in the model config—including IA³, DoRA, or custom PEFT implementations—can be updated. The controller itself is adapter-agnostic; it only triggers the existing fine-tuning pipeline with appropriate freeze/unfreeze settings.

How does the controller interact with distributed data parallelism?

Because the controller runs in the main process of each rank, only rank 0 typically activates the controller in practice (enforced by args.local_rank == 0 checks). Health signals and adapter state are then broadcast through standard distributed communication primitives. The implementation avoids MPI or NCAR directly, relying on the trainer's existing synchronization.

What's the performance overhead of health polling?

At default 1-second intervals, the controller adds negligible overhead—each poll is a lightweight HTTP GET to the local API server. For sub-second latency requirements, mini_ft_controller_poll_interval can be reduced to 0.1, though this increases CPU utilization marginally.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →