How to Perform a Health Check on the MTPLX Server: Complete Guide
MTPLX exposes a lightweight /health HTTP endpoint that returns runtime status, model information, resource usage, and configuration flags without triggering expensive model loading or inference.
The MTPLX inference server includes a built-in health monitoring system accessible through a simple HTTP GET request. This endpoint aggregates real-time data from multiple subsystems including the model scheduler, thermal management, and feature flags—making it ideal for automated monitoring, load balancer health probes, and operational diagnostics.
Understanding the MTPLX Health Check Endpoint
The health endpoint is implemented in the FastAPI server at mtplx/server/openai.py and registered with the decorator @app.get("/health") lines 28868. When invoked, the handler constructs a comprehensive JSON payload from several in-memory subsystems without performing any model inference.
Core Design Principles
The endpoint follows two critical design constraints:
- Read-only operation: All metrics are pre-computed counters stored in memory
- Zero inference cost: No model loading or token generation occurs
This architecture allows health checks to run every few seconds in production environments without degrading serving performance.
Health Response Structure
The JSON payload returned by /health contains seven primary sections:
| Section | Description | Key Fields |
|---|---|---|
| Model & Runtime | Current model identifier and operational mode | model, runtime_mode, fan_boost_active |
| Scheduler | Request queuing and scheduling statistics | scheduler, active_requests |
| Thermal & Fan | Hardware thermal management status | thermal, smart_fan_* |
| Profile & Sampler | Active configuration profile and sampling parameters | profile, sampler |
| Resource & Queue | Request counts across foreground, dashboard, and scheduler queues | foreground_active, dashboard_active_requests, idle_seconds |
| Feature Flags | Runtime configuration toggles | api_key_required, verify_core, enable_thinking |
| Diagnostics | Kernel self-check and GPU state | kernel_selfcheck, gpu_keepalive, qwen4_install_reports |
These fields are populated from lines 28905-28999 in mtplx/server/openai.py, with subsystem-specific data sourced from mtplx/thermal.py (thermal payload) and mtplx/model_scheduler.py (scheduler state).
Security note: The response intentionally excludes sensitive values. Only boolean flags like
api_key_requiredare exposed—actual keys remain server-side.
Methods to Check MTPLX Server Health
Method 1: curl (Command-Line Verification)
The fastest way to verify server liveness:
curl -s http://localhost:8000/health | jq .
For production automation without jq:
curl -sf http://localhost:8000/health > /dev/null && echo "healthy" || echo "unhealthy"
Method 2: Python requests Client
Programmatic access matching the built-in CLI behavior:
import requests
import json
def check_mtplx_health(base_url: str = "http://localhost:8000", timeout: float = 2.0) -> dict:
"""Perform health check on MTPLX server."""
resp = requests.get(f"{base_url}/health", timeout=timeout)
resp.raise_for_status()
return resp.json()
health = check_mtplx_health()
print(json.dumps(health, indent=2))
Method 3: MTPLX CLI Helper
Native command-line integration with formatted output:
mtplx health http://localhost:8000
This command is implemented in mtplx/commands/public.py lines 1971-2099 and uses an internal _http_json helper with a 1.5-second default timeout.
Method 4: Prometheus Monitoring Configuration
For production observability stacks:
scrape_configs:
- job_name: 'mtplx'
static_configs:
- targets: ['localhost:8000']
metrics_path: '/health'
scheme: http
scrape_interval: 5s
Key Source Files for Health Implementation
| File | Purpose |
|---|---|
mtplx/server/openai.py |
FastAPI route registration and payload assembly |
mtplx/commands/public.py |
CLI wrapper (mtplx health) |
mtplx/thermal.py |
Thermal subsystem data provider (_thermal_health_payload) |
mtplx/model_scheduler.py |
Scheduler state export (_mtplx_scheduler_state) |
mtplx/daemon_client.py |
Daemon discovery via health probe (url = f"http://{host}:{port}/health") |
mtplx/dashboard/__init__.py |
Web UI health verification on startup |
Interpreting Health Response Fields
Critical Fields for Load Balancers
runtime_mode: Expected values includeloaded,loading, orunavailableactive_requests: Current inference backlog—use for capacity-based routingidle_seconds: Time since last request; useful for scaling decisions
Diagnostic Fields for Troubleshooting
kernel_selfcheck: Boolean indicating kernel module integritygpu_keepalive: GPU communication channel statussmart_fan_status: Thermal management operational state
Configuration Verification
api_key_required: Whether authentication is enforcedverify_core: Core file validation toggleenable_thinking: Extended reasoning mode availability
Production Health Check Strategies
Recommended Polling Intervals
- Load balancer health probes: 5-10 seconds with 2-second timeout
- Kubernetes liveness checks: 10-second period, 3-failure threshold
- Detailed monitoring dashboards: 1-5 seconds for real-time visibility
Timeout Configuration Guidelines
| Client Type | Recommended Timeout | Rationale |
|---|---|---|
| CLI interactive use | 5.0 seconds | User tolerance for slow responses |
| Automated scripts | 2.0 seconds | Balance reliability vs. speed |
| Service mesh / mesh proxies | 1.5 seconds | Match MTPLX CLI default |
| Aggressive load balancers | 1.0 second | Fast failover requirements |
Summary
- The MTPLX
/healthendpoint provides comprehensive runtime status through a single HTTP GET request tohttp://<host>:<port>/health - Response payload includes model state, scheduler statistics, thermal metrics, and feature flags sourced from
mtplx/server/openai.pylines 28905-28999 - Zero inference overhead design permits frequent polling without performance impact
- Multiple access methods: direct
curl, Pythonrequests, nativemtplx healthCLI, or Prometheus scrape configuration - Security-hardened response excludes actual secrets while exposing configuration boolean flags
Frequently Asked Questions
What port does the MTPLX health endpoint use?
The default port is 8000, matching the standard FastAPI server binding in mtplx/server/openai.py. This can be configured at server startup through environment variables or command-line arguments.
Does the health check endpoint require authentication?
Authentication depends on runtime configuration. If api_key_required is true in the health response, the server enforces API key validation on inference endpoints—the /health endpoint itself remains accessible for load balancer probes. Check the api_key_required field in your health response to determine current policy.
Why does my health check show runtime_mode: "loading"?
This indicates the server has started but the model weights are not yet fully loaded into GPU memory. Depending on model size and storage speed, this state may persist from several seconds to minutes. Inference requests will queue during this period; scale your timeout thresholds accordingly or implement a readiness gate that waits for runtime_mode: "loaded".
Can I use the health endpoint for automatic failover?
Yes—the lightweight design supports sub-second polling intervals ideal for load balancer health checks. Key fields for failover logic include runtime_mode (fail if not loaded), kernel_selfcheck (fail if false), and gpu_keepalive (fail if disconnected). Combine these with response time thresholds to build robust availability detection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →