# How the Mini Fine-Tuning Controller Complements the Main Training Loop in Miles

> Discover how the mini fine-tuning controller in Miles complements the main training loop by monitoring rollouts and triggering adapter updates without interrupting data-parallel training.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: internals
- Published: 2026-09-06

---

**The Mini Fine-Tuning Controller runs as a daemon thread alongside the main training loop, monitoring rollout health signals and triggering lightweight adapter updates without interrupting data-parallel training.**

In distributed LLM training, fault tolerance and rapid adaptation are critical. The Miles repository implements a **Mini Fine-Tuning (FT) Controller** that integrates directly with the primary training pipeline to enable on-demand adapter fine-tuning. This article examines how `maybe_start_mini_ft_controller()` bridges the main training loop in [`train.py`](https://github.com/radixark/miles/blob/main/train.py) and [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py) with responsive, low-overhead model updates.

---

## What the Mini FT Controller Does

The Mini FT Controller serves two complementary functions that keep long-running training jobs resilient:

| Function | Implementation | Impact on Training |
|----------|---------------|------------------|
| **Health signal polling** | A daemon thread queries the local API server at `mini_ft_controller_poll_interval` for "heal" signals from the rollout subsystem | Zero interruption to gradient computation; asynchronous detection of rollout failures |
| **Reactive adapter fine-tuning** | Upon signal detection, the controller launches short-lived LoRA/IA³ updates via `_MiniFTControllerRunner` | Model reference remains stable; only adapter weights change, preserving training throughput |

This architecture separates concerns: the main loop handles data parallelism and gradient aggregation, while the controller manages fine-grained model health without process restarts.

---

## Integration Point: `maybe_start_mini_ft_controller()`

The controller hooks into training through a single entry point. In [`miles/main/train.py`](https://github.com/radixark/miles/blob/main/miles/main/train.py) and [`miles/main/train_async.py`](https://github.com/radixark/miles/blob/main/miles/main/train_async.py), the startup sequence follows this pattern:

```python
from miles.utils.ft_utils.mini_ft_controller import maybe_start_mini_ft_controller

def main():
    args = parse_args()
    
    # Initialize controller before heavy training setup

    mini_ft_handle = maybe_start_mini_ft_controller(args)
    
    # Proceed with standard training pipeline

    trainer = Trainer(args)
    trainer.train()

```

The `maybe_start_mini_ft_controller()` function checks activation flags and either returns a handle to the running controller or `None` if disabled. This design keeps the integration non-invasive—training logic requires no conditional branches to accommodate the controller.

---

## Activation and Configuration

The controller respects explicit CLI flags defined in the argument parser of [`train.py`](https://github.com/radixark/miles/blob/main/train.py):

| Flag | Default | Purpose |
|------|---------|---------|
| `--mini-ft-controller-enable` | `False` (enabled implicitly via `--use-fault-tolerance`) | Explicit activation |
| `--no-mini-ft-controller-enable` | — | Force disable |
| `--mini-ft-controller-poll-interval` | `1.0` | Seconds between health checks |
| `--mini-ft-controller-resume-delay` | `0.0` | Cooldown after fine-tuning before resuming rollouts |

### Enable with Fault Tolerance (Recommended)

```bash
python -m miles.main.train \
    --model qwen3-5-35b \
    --num-rollout 4 \
    --use-fault-tolerance

```

The `--use-fault-tolerance` flag implicitly enables the Mini FT Controller with sensible defaults for production deployments.

### Explicit Tuning for Async Training

```bash
python -m miles.main.train_async \
    --model qwen3-6-35b-a3b \
    --mini-ft-controller-enable \
    --mini-ft-controller-poll-interval 2.0 \
    --mini-ft-controller-resume-delay 5.0

```

This configuration polls every 2 seconds and imposes a 5-second stabilization period post-healing—useful when rollout variance is high.

### Complete Disable

```bash
python -m miles.main.train \
    --model gemma-4-31b \
    --no-mini-ft-controller-enable

```

The `maybe_start_mini_ft_controller()` call returns immediately, and training proceeds without sidecar monitoring.

---

## Internal Architecture: [`mini_ft_controller.py`](https://github.com/radixark/miles/blob/main/mini_ft_controller.py)

The implementation in [`miles/utils/ft_utils/mini_ft_controller.py`](https://github.com/radixark/miles/blob/main/miles/utils/ft_utils/mini_ft_controller.py) organizes functionality into discrete classes:

- **`_MiniFTControllerRunner`** — Encapsulates the daemon thread lifecycle and signal processing loop
- **Health check logic** — HTTP requests to `localhost:{args.api_server_port}/health` or equivalent endpoint
- **Adapter update dispatch** — Invokes the fine-tuning pipeline with frozen base model weights

Key design decisions visible in the source:

1. **Thread-based, not process-based** — Shares Python interpreter state with the trainer; no serialization overhead
2. **Signal-driven, not timer-driven** — Polling interval only bounds maximum latency; actual response is event-triggered
3. **Graceful degradation** — Controller exceptions are caught and logged; training continues even if healing fails

---

## Testing and Verification

The repository includes targeted tests in [`tests/fast/utils/test_mini_ft_controller.py`](https://github.com/radixark/miles/blob/main/tests/fast/utils/test_mini_ft_controller.py) that validate:

- Correct startup/shutdown behavior under various flag combinations
- Thread safety of shared model references
- Proper handling of malformed heal signals

Integration with the API server is exercised in [`tests/fast/api_server/test_handles.py`](https://github.com/radixark/miles/blob/main/tests/fast/api_server/test_handles.py), which defines the signal schema that production controllers consume.

---

## Summary

- The **Mini FT Controller** runs as a **daemon thread** spawned by `maybe_start_mini_ft_controller()` in [`train.py`](https://github.com/radixark/miles/blob/main/train.py) and [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py)
- It **polls for health signals** at configurable intervals without blocking gradient computation
- Upon detection, it executes **lightweight adapter fine-tuning** (LoRA/IA³) while preserving the base model reference
- Activation is controlled via `--use-fault-tolerance` or explicit `--mini-ft-controller-enable` flags
- The implementation in [`miles/utils/ft_utils/mini_ft_controller.py`](https://github.com/radixark/miles/blob/main/miles/utils/ft_utils/mini_ft_controller.py) prioritizes **zero-overhead integration** through thread-sharing and defensive error handling

---

## Frequently Asked Questions

### What happens if the Mini FT Controller crashes during training?

The controller wraps its main loop in exception handlers that log errors without propagating to the parent thread. Training continues uninterrupted; the model simply won't receive automatic healing updates until restart. Monitor logs for `MiniFTController` errors to detect degradation.

### Can I use the Mini FT Controller with custom adapter types beyond LoRA?

The controller delegates actual fine-tuning to Miles' internal adapter system. Any adapter registered in the model config—including IA³, DoRA, or custom PEFT implementations—can be updated. The controller itself is adapter-agnostic; it only triggers the existing fine-tuning pipeline with appropriate freeze/unfreeze settings.

### How does the controller interact with distributed data parallelism?

Because the controller runs in the main process of each rank, **only rank 0 typically activates the controller** in practice (enforced by `args.local_rank == 0` checks). Health signals and adapter state are then broadcast through standard distributed communication primitives. The implementation avoids MPI or NCAR directly, relying on the trainer's existing synchronization.

### What's the performance overhead of health polling?

At default 1-second intervals, the controller adds negligible overhead—each poll is a lightweight HTTP GET to the local API server. For sub-second latency requirements, `mini_ft_controller_poll_interval` can be reduced to 0.1, though this increases CPU utilization marginally.