# How to Perform a Health Check on the MTPLX Server: Complete Guide

> Learn how to perform a health check on the MTPLX server using its /health endpoint. Get runtime status, model info, resource usage, and config flags quickly.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-06

---

**MTPLX exposes a lightweight `/health` HTTP endpoint that returns runtime status, model information, resource usage, and configuration flags without triggering expensive model loading or inference.**

The MTPLX inference server includes a built-in health monitoring system accessible through a simple HTTP GET request. This endpoint aggregates real-time data from multiple subsystems including the model scheduler, thermal management, and feature flags—making it ideal for automated monitoring, load balancer health probes, and operational diagnostics.

## Understanding the MTPLX Health Check Endpoint

The health endpoint is implemented in the FastAPI server at [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) and registered with the decorator `@app.get("/health")` [lines 28868](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L28868). When invoked, the handler constructs a comprehensive JSON payload from several in-memory subsystems without performing any model inference.

### Core Design Principles

The endpoint follows two critical design constraints:

- **Read-only operation**: All metrics are pre-computed counters stored in memory
- **Zero inference cost**: No model loading or token generation occurs

This architecture allows health checks to run every few seconds in production environments without degrading serving performance.

## Health Response Structure

The JSON payload returned by `/health` contains seven primary sections:

| Section | Description | Key Fields |
|:---|:---|:---|
| **Model & Runtime** | Current model identifier and operational mode | `model`, `runtime_mode`, `fan_boost_active` |
| **Scheduler** | Request queuing and scheduling statistics | `scheduler`, `active_requests` |
| **Thermal & Fan** | Hardware thermal management status | `thermal`, `smart_fan_*` |
| **Profile & Sampler** | Active configuration profile and sampling parameters | `profile`, `sampler` |
| **Resource & Queue** | Request counts across foreground, dashboard, and scheduler queues | `foreground_active`, `dashboard_active_requests`, `idle_seconds` |
| **Feature Flags** | Runtime configuration toggles | `api_key_required`, `verify_core`, `enable_thinking` |
| **Diagnostics** | Kernel self-check and GPU state | `kernel_selfcheck`, `gpu_keepalive`, `qwen4_install_reports` |

These fields are populated from lines [28905-28999](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L28905-L28999) in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), with subsystem-specific data sourced from [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) (thermal payload) and [`mtplx/model_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_scheduler.py) (scheduler state).

> **Security note**: The response intentionally excludes sensitive values. Only boolean flags like `api_key_required` are exposed—actual keys remain server-side.

## Methods to Check MTPLX Server Health

### Method 1: curl (Command-Line Verification)

The fastest way to verify server liveness:

```bash
curl -s http://localhost:8000/health | jq .

```

For production automation without `jq`:

```bash
curl -sf http://localhost:8000/health > /dev/null && echo "healthy" || echo "unhealthy"

```

### Method 2: Python requests Client

Programmatic access matching the built-in CLI behavior:

```python
import requests
import json

def check_mtplx_health(base_url: str = "http://localhost:8000", timeout: float = 2.0) -> dict:
    """Perform health check on MTPLX server."""
    resp = requests.get(f"{base_url}/health", timeout=timeout)
    resp.raise_for_status()
    return resp.json()

health = check_mtplx_health()
print(json.dumps(health, indent=2))

```

### Method 3: MTPLX CLI Helper

Native command-line integration with formatted output:

```bash
mtplx health http://localhost:8000

```

This command is implemented in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) [lines 1971-2099](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py#L1971-L2099) and uses an internal `_http_json` helper with a **1.5-second default timeout**.

### Method 4: Prometheus Monitoring Configuration

For production observability stacks:

```yaml
scrape_configs:
  - job_name: 'mtplx'
    static_configs:
      - targets: ['localhost:8000']
    metrics_path: '/health'
    scheme: http
    scrape_interval: 5s

```

## Key Source Files for Health Implementation

| File | Purpose |
|:---|:---|
| [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) | FastAPI route registration and payload assembly |
| [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) | CLI wrapper (`mtplx health`) |
| [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) | Thermal subsystem data provider (`_thermal_health_payload`) |
| [`mtplx/model_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_scheduler.py) | Scheduler state export (`_mtplx_scheduler_state`) |
| [`mtplx/daemon_client.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/daemon_client.py) | Daemon discovery via health probe (`url = f"http://{host}:{port}/health"`) |
| [`mtplx/dashboard/__init__.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/dashboard/__init__.py) | Web UI health verification on startup |

## Interpreting Health Response Fields

### Critical Fields for Load Balancers

- `runtime_mode`: Expected values include `loaded`, `loading`, or `unavailable`
- `active_requests`: Current inference backlog—use for capacity-based routing
- `idle_seconds`: Time since last request; useful for scaling decisions

### Diagnostic Fields for Troubleshooting

- `kernel_selfcheck`: Boolean indicating kernel module integrity
- `gpu_keepalive`: GPU communication channel status
- `smart_fan_status`: Thermal management operational state

### Configuration Verification

- `api_key_required`: Whether authentication is enforced
- `verify_core`: Core file validation toggle
- `enable_thinking`: Extended reasoning mode availability

## Production Health Check Strategies

### Recommended Polling Intervals

- **Load balancer health probes**: 5-10 seconds with 2-second timeout
- **Kubernetes liveness checks**: 10-second period, 3-failure threshold
- **Detailed monitoring dashboards**: 1-5 seconds for real-time visibility

### Timeout Configuration Guidelines

| Client Type | Recommended Timeout | Rationale |
|:---|:---|:---|
| CLI interactive use | 5.0 seconds | User tolerance for slow responses |
| Automated scripts | 2.0 seconds | Balance reliability vs. speed |
| Service mesh / mesh proxies | 1.5 seconds | Match MTPLX CLI default |
| Aggressive load balancers | 1.0 second | Fast failover requirements |

## Summary

- The MTPLX `/health` endpoint provides comprehensive runtime status through a **single HTTP GET request** to `http://<host>:<port>/health`
- Response payload includes model state, scheduler statistics, thermal metrics, and feature flags sourced from [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) [lines 28905-28999](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L28905-L28999)
- **Zero inference overhead** design permits frequent polling without performance impact
- Multiple access methods: direct `curl`, Python `requests`, native `mtplx health` CLI, or Prometheus scrape configuration
- Security-hardened response excludes actual secrets while exposing configuration boolean flags

## Frequently Asked Questions

### What port does the MTPLX health endpoint use?

The default port is **8000**, matching the standard FastAPI server binding in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py). This can be configured at server startup through environment variables or command-line arguments.

### Does the health check endpoint require authentication?

Authentication depends on runtime configuration. If `api_key_required` is `true` in the health response, the server enforces API key validation on inference endpoints—the `/health` endpoint itself remains accessible for load balancer probes. Check the `api_key_required` field in your health response to determine current policy.

### Why does my health check show `runtime_mode: "loading"`?

This indicates the server has started but the model weights are not yet fully loaded into GPU memory. Depending on model size and storage speed, this state may persist from several seconds to minutes. Inference requests will queue during this period; scale your timeout thresholds accordingly or implement a readiness gate that waits for `runtime_mode: "loaded"`.

### Can I use the health endpoint for automatic failover?

Yes—the lightweight design supports sub-second polling intervals ideal for load balancer health checks. Key fields for failover logic include `runtime_mode` (fail if not `loaded`), `kernel_selfcheck` (fail if `false`), and `gpu_keepalive` (fail if disconnected). Combine these with response time thresholds to build robust availability detection.