# How Miles Handles Fault Tolerance for SGLang Engine Failures and Run Recovery

> Learn how Miles ensures fault tolerance for SGLang engine failures. Discover automatic crash recovery, actor restarts, weight reloads, and request replays for uninterrupted training and inference.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Miles detects SGLang engine failures through health‑checking, automatically restarts the crashed Ray actor, reloads weights, and replays pending requests to recover the training or inference run without manual intervention.**

Fault tolerance for SGLang engine failures is a core capability in **Miles** (Massively‑parallel In‑Context Learning Engine Service), a distributed inference system built on Ray. When an SGLang engine crashes due to out‑of‑memory errors, segmentation faults, or router failures, Miles provides automatic detection, restart, and request replay to ensure runs continue uninterrupted. This article explains the three‑layer architecture—launch supervision, health‑checking, and recovery replay—using actual source code paths from the `radixark/miles` repository.

## Layer 1: Launch and Supervision with Fault Tolerance Flag

The fault‑tolerance flow begins at launch time. Users enable the feature via a command‑line flag that propagates through the entire stack.

The argument parser in [`miles/utils/arguments.py`](https://github.com/radixark/miles/blob/main/miles/utils/arguments.py) (line 967) defines the toggle:

```python
parser.add_argument(
    "--use-fault-tolerance",
    action="store_true",
    help="Enable fault tolerance for the SGLang engine and router.",
)

```

When [`train_async.py`](https://github.com/radixark/miles/blob/main/train_async.py) or [`train_multi_lora_async.py`](https://github.com/radixark/miles/blob/main/train_multi_lora_async.py) invokes `launch_server_process`, this flag is forwarded to both `RouterArgs` and `ServerArgs`. The launch helper creates a Ray‑remote SGLang engine and attaches a watchdog process that monitors engine health throughout the run lifetime.

## Layer 2: Health‑Checking and Failure Detection

Miles detects engine death through continuous heart‑beating implemented in `SGLangApiClient`.

The client ([`miles/backends/sglang_utils/sglang_api_client.py`](https://github.com/radixark/miles/blob/main/miles/backends/sglang_utils/sglang_api_client.py)) runs a background coroutine that pings the router's `/healthz` endpoint. When any of these conditions occur, the client raises `SGLangEngineUnavailable`:

- **RPC timeout** or connection refused
- **Non‑200 HTTP response** from `/healthz`
- **Explicit engine kill** signal

The rollout manager `InferenceController` in [`miles/ray/rollout/inference_controller.py`](https://github.com/radixark/miles/blob/main/miles/ray/rollout/inference_controller.py) catches this exception and initiates the recovery protocol. This separation of detection (client) and orchestration (controller) allows for flexible retry policies and centralized failure handling across distributed workers.

## Layer 3: Recovery, Weight Reload, and Request Replay

Once a failure is detected, Miles performs three coordinated actions to restore service:

**1. Weight Reload**

Before restarting the engine, the controller calls `SGLangApiClient.update_weights_from_disk` to ensure the new process loads the latest checkpoint. This prevents stale model state from corrupting the resumed run.

**2. Engine Restart**

The controller terminates the broken Ray actor via `engine.shutdown()` or `engine.kill_subprocess()`, then spawns a fresh actor with identical configuration. The new engine performs its own health check before accepting traffic.

**3. Request Replay**

During normal operation, `InferenceController` buffers each request's JSON payload in `self.pending_requests`. After restart, the controller iterates this FIFO queue and reissues calls via `SGLangApiClient.generate`. This provides **exact‑once semantics** with preserved ordering—critical for deterministic training loops.

```python

# Simplified recovery flow inside InferenceController

try:
    response = await self.client.generate(request_payload)
except SGLangEngineUnavailable:
    # Trigger recovery

    await self.client.update_weights_from_disk(checkpoint_path)
    await self.engine.shutdown()
    self.engine = await launch_server_process(
        router_ip=self.router_ip,
        router_port=self.router_port,
        use_fault_tolerance=True,
    )
    # Replay buffered requests

    for buffered_request in self.pending_requests:
        response = await self.client.generate(buffered_request)

```

## Mock Engine Testing for Fault Injection

The `MockSGLangEngine` class in [`miles/utils/test_utils/mock_sglang_engine.py`](https://github.com/radixark/miles/blob/main/miles/utils/test_utils/mock_sglang_engine.py) validates the fault‑tolerance path without requiring real GPU resources. It implements three key methods:

- **`set_fault(method, exception)`**: Registers an exception to inject on the next call to `method`
- **`_maybe_fault(method)`**: Raises the stored exception if one is registered
- **`inject_fault()`**: Forces immediate failure for testing recovery timing

The mock records all calls in `self.calls` and mimics real SGLang behavior through `MockSGLangHttpServer`. Tests in [`tests/fast/router/test_sessions_v2.py`](https://github.com/radixark/miles/blob/main/tests/fast/router/test_sessions_v2.py) and [`test_sessions_v1_pins.py`](https://github.com/radixark/miles/blob/main/test_sessions_v1_pins.py) verify that injecting a fault on the `run` method triggers actor restart and successful subsequent `generate` calls.

## Router Health Check Validation

The router's availability is verified independently in [`tests/fast/router/test_router.py`](https://github.com/radixark/miles/blob/main/tests/fast/router/test_router.py). The test `test_check_worker_health_success` confirms that `/healthz` returns success when workers are healthy; a failing health check provokes the same restart flow used in production, ensuring consistency between test and runtime behavior.

## Practical Usage Example

Enable fault tolerance when launching a training run:

```bash
python -m miles.main.train \
    --model qwen3-8b \
    --sglang-router-ip 127.0.0.1 \
    --sglang-router-port 31000 \
    --use-fault-tolerance

```

Programmatic engine launch with fault tolerance:

```python
from miles.backends.sglang_utils.sglang_engine import launch_server_process
from miles.utils.arguments import parse_args

args = parse_args()
engine = launch_server_process(
    router_ip=args.sglang_router_ip,
    router_port=args.sglang_router_port,
    use_fault_tolerance=args.use_fault_tolerance,
)

# engine is a Ray actor with built-in health monitoring

```

When failure occurs, the sequence runs automatically: health‑check detection → weight reload → actor restart → request replay.

## Summary

- **Enable with `--use-fault-tolerance`** — propagates through `RouterArgs` and `ServerArgs` from [`miles/utils/arguments.py`](https://github.com/radixark/miles/blob/main/miles/utils/arguments.py)
- **Detect via `SGLangApiClient`** — continuous `/healthz` polling raises `SGLangEngineUnavailable` on failure
- **Recover through `InferenceController`** — orchestrates weight reload, Ray actor restart, and buffered request replay
- **Test with `MockSGLangEngine`** — `set_fault` and `_maybe_fault` inject failures for validation without hardware

## Frequently Asked Questions

### What types of SGLang engine failures does Miles handle?

Miles handles process crashes (segmentation faults, OOM kills), RPC timeouts, router unavailability, and explicit shutdown signals. The health‑check mechanism in `SGLangApiClient` treats any non‑responsive `/healthz` endpoint as a failure trigger, regardless of root cause.

### How does Miles ensure no training data is lost during recovery?

The `InferenceController` maintains `self.pending_requests` as a FIFO buffer of JSON payloads. When restart completes, it replays these buffered requests in original order. Successful responses are cleared from the buffer; only unacknowledged requests persist for retry.

### Can fault tolerance be used during inference-only deployments?

Yes. The same `--use-fault-tolerance` flag and `launch_server_process` API work for inference workloads. The weight reload step becomes optional if the model weights haven't changed, but the request replay mechanism remains active to recover in‑flight generation calls.

### Where is the fault tolerance tested in the Miles repository?

Key test files include [`tests/fast/router/test_sessions_v2.py`](https://github.com/radixark/miles/blob/main/tests/fast/router/test_sessions_v2.py) and [`test_sessions_v1_pins.py`](https://github.com/radixark/miles/blob/main/test_sessions_v1_pins.py) for end‑to‑end recovery scenarios, [`tests/fast/router/test_router.py`](https://github.com/radixark/miles/blob/main/tests/fast/router/test_router.py) for health‑check validation, and [`miles/utils/test_utils/mock_sglang_engine.py`](https://github.com/radixark/miles/blob/main/miles/utils/test_utils/mock_sglang_engine.py) which provides the `MockSGLangEngine` infrastructure for fault injection without GPU dependencies.