How Miles Handles Fault Tolerance for SGLang Engine Failures and Run Recovery
Miles detects SGLang engine failures through health‑checking, automatically restarts the crashed Ray actor, reloads weights, and replays pending requests to recover the training or inference run without manual intervention.
Fault tolerance for SGLang engine failures is a core capability in Miles (Massively‑parallel In‑Context Learning Engine Service), a distributed inference system built on Ray. When an SGLang engine crashes due to out‑of‑memory errors, segmentation faults, or router failures, Miles provides automatic detection, restart, and request replay to ensure runs continue uninterrupted. This article explains the three‑layer architecture—launch supervision, health‑checking, and recovery replay—using actual source code paths from the radixark/miles repository.
Layer 1: Launch and Supervision with Fault Tolerance Flag
The fault‑tolerance flow begins at launch time. Users enable the feature via a command‑line flag that propagates through the entire stack.
The argument parser in miles/utils/arguments.py (line 967) defines the toggle:
parser.add_argument(
"--use-fault-tolerance",
action="store_true",
help="Enable fault tolerance for the SGLang engine and router.",
)
When train_async.py or train_multi_lora_async.py invokes launch_server_process, this flag is forwarded to both RouterArgs and ServerArgs. The launch helper creates a Ray‑remote SGLang engine and attaches a watchdog process that monitors engine health throughout the run lifetime.
Layer 2: Health‑Checking and Failure Detection
Miles detects engine death through continuous heart‑beating implemented in SGLangApiClient.
The client (miles/backends/sglang_utils/sglang_api_client.py) runs a background coroutine that pings the router's /healthz endpoint. When any of these conditions occur, the client raises SGLangEngineUnavailable:
- RPC timeout or connection refused
- Non‑200 HTTP response from
/healthz - Explicit engine kill signal
The rollout manager InferenceController in miles/ray/rollout/inference_controller.py catches this exception and initiates the recovery protocol. This separation of detection (client) and orchestration (controller) allows for flexible retry policies and centralized failure handling across distributed workers.
Layer 3: Recovery, Weight Reload, and Request Replay
Once a failure is detected, Miles performs three coordinated actions to restore service:
1. Weight Reload
Before restarting the engine, the controller calls SGLangApiClient.update_weights_from_disk to ensure the new process loads the latest checkpoint. This prevents stale model state from corrupting the resumed run.
2. Engine Restart
The controller terminates the broken Ray actor via engine.shutdown() or engine.kill_subprocess(), then spawns a fresh actor with identical configuration. The new engine performs its own health check before accepting traffic.
3. Request Replay
During normal operation, InferenceController buffers each request's JSON payload in self.pending_requests. After restart, the controller iterates this FIFO queue and reissues calls via SGLangApiClient.generate. This provides exact‑once semantics with preserved ordering—critical for deterministic training loops.
# Simplified recovery flow inside InferenceController
try:
response = await self.client.generate(request_payload)
except SGLangEngineUnavailable:
# Trigger recovery
await self.client.update_weights_from_disk(checkpoint_path)
await self.engine.shutdown()
self.engine = await launch_server_process(
router_ip=self.router_ip,
router_port=self.router_port,
use_fault_tolerance=True,
)
# Replay buffered requests
for buffered_request in self.pending_requests:
response = await self.client.generate(buffered_request)
Mock Engine Testing for Fault Injection
The MockSGLangEngine class in miles/utils/test_utils/mock_sglang_engine.py validates the fault‑tolerance path without requiring real GPU resources. It implements three key methods:
set_fault(method, exception): Registers an exception to inject on the next call tomethod_maybe_fault(method): Raises the stored exception if one is registeredinject_fault(): Forces immediate failure for testing recovery timing
The mock records all calls in self.calls and mimics real SGLang behavior through MockSGLangHttpServer. Tests in tests/fast/router/test_sessions_v2.py and test_sessions_v1_pins.py verify that injecting a fault on the run method triggers actor restart and successful subsequent generate calls.
Router Health Check Validation
The router's availability is verified independently in tests/fast/router/test_router.py. The test test_check_worker_health_success confirms that /healthz returns success when workers are healthy; a failing health check provokes the same restart flow used in production, ensuring consistency between test and runtime behavior.
Practical Usage Example
Enable fault tolerance when launching a training run:
python -m miles.main.train \
--model qwen3-8b \
--sglang-router-ip 127.0.0.1 \
--sglang-router-port 31000 \
--use-fault-tolerance
Programmatic engine launch with fault tolerance:
from miles.backends.sglang_utils.sglang_engine import launch_server_process
from miles.utils.arguments import parse_args
args = parse_args()
engine = launch_server_process(
router_ip=args.sglang_router_ip,
router_port=args.sglang_router_port,
use_fault_tolerance=args.use_fault_tolerance,
)
# engine is a Ray actor with built-in health monitoring
When failure occurs, the sequence runs automatically: health‑check detection → weight reload → actor restart → request replay.
Summary
- Enable with
--use-fault-tolerance— propagates throughRouterArgsandServerArgsfrommiles/utils/arguments.py - Detect via
SGLangApiClient— continuous/healthzpolling raisesSGLangEngineUnavailableon failure - Recover through
InferenceController— orchestrates weight reload, Ray actor restart, and buffered request replay - Test with
MockSGLangEngine—set_faultand_maybe_faultinject failures for validation without hardware
Frequently Asked Questions
What types of SGLang engine failures does Miles handle?
Miles handles process crashes (segmentation faults, OOM kills), RPC timeouts, router unavailability, and explicit shutdown signals. The health‑check mechanism in SGLangApiClient treats any non‑responsive /healthz endpoint as a failure trigger, regardless of root cause.
How does Miles ensure no training data is lost during recovery?
The InferenceController maintains self.pending_requests as a FIFO buffer of JSON payloads. When restart completes, it replays these buffered requests in original order. Successful responses are cleared from the buffer; only unacknowledged requests persist for retry.
Can fault tolerance be used during inference-only deployments?
Yes. The same --use-fault-tolerance flag and launch_server_process API work for inference workloads. The weight reload step becomes optional if the model weights haven't changed, but the request replay mechanism remains active to recover in‑flight generation calls.
Where is the fault tolerance tested in the Miles repository?
Key test files include tests/fast/router/test_sessions_v2.py and test_sessions_v1_pins.py for end‑to‑end recovery scenarios, tests/fast/router/test_router.py for health‑check validation, and miles/utils/test_utils/mock_sglang_engine.py which provides the MockSGLangEngine infrastructure for fault injection without GPU dependencies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →