How to Perform a Health Check on the MTPLX Server: Complete Guide

MTPLX exposes a lightweight /health HTTP endpoint that returns runtime status, model information, resource usage, and configuration flags without triggering expensive model loading or inference.

The MTPLX inference server includes a built-in health monitoring system accessible through a simple HTTP GET request. This endpoint aggregates real-time data from multiple subsystems including the model scheduler, thermal management, and feature flags—making it ideal for automated monitoring, load balancer health probes, and operational diagnostics.

Understanding the MTPLX Health Check Endpoint

The health endpoint is implemented in the FastAPI server at mtplx/server/openai.py and registered with the decorator @app.get("/health") lines 28868. When invoked, the handler constructs a comprehensive JSON payload from several in-memory subsystems without performing any model inference.

Core Design Principles

The endpoint follows two critical design constraints:

  • Read-only operation: All metrics are pre-computed counters stored in memory
  • Zero inference cost: No model loading or token generation occurs

This architecture allows health checks to run every few seconds in production environments without degrading serving performance.

Health Response Structure

The JSON payload returned by /health contains seven primary sections:

Section Description Key Fields
Model & Runtime Current model identifier and operational mode model, runtime_mode, fan_boost_active
Scheduler Request queuing and scheduling statistics scheduler, active_requests
Thermal & Fan Hardware thermal management status thermal, smart_fan_*
Profile & Sampler Active configuration profile and sampling parameters profile, sampler
Resource & Queue Request counts across foreground, dashboard, and scheduler queues foreground_active, dashboard_active_requests, idle_seconds
Feature Flags Runtime configuration toggles api_key_required, verify_core, enable_thinking
Diagnostics Kernel self-check and GPU state kernel_selfcheck, gpu_keepalive, qwen4_install_reports

These fields are populated from lines 28905-28999 in mtplx/server/openai.py, with subsystem-specific data sourced from mtplx/thermal.py (thermal payload) and mtplx/model_scheduler.py (scheduler state).

Security note: The response intentionally excludes sensitive values. Only boolean flags like api_key_required are exposed—actual keys remain server-side.

Methods to Check MTPLX Server Health

Method 1: curl (Command-Line Verification)

The fastest way to verify server liveness:

curl -s http://localhost:8000/health | jq .

For production automation without jq:

curl -sf http://localhost:8000/health > /dev/null && echo "healthy" || echo "unhealthy"

Method 2: Python requests Client

Programmatic access matching the built-in CLI behavior:

import requests
import json

def check_mtplx_health(base_url: str = "http://localhost:8000", timeout: float = 2.0) -> dict:
    """Perform health check on MTPLX server."""
    resp = requests.get(f"{base_url}/health", timeout=timeout)
    resp.raise_for_status()
    return resp.json()

health = check_mtplx_health()
print(json.dumps(health, indent=2))

Method 3: MTPLX CLI Helper

Native command-line integration with formatted output:

mtplx health http://localhost:8000

This command is implemented in mtplx/commands/public.py lines 1971-2099 and uses an internal _http_json helper with a 1.5-second default timeout.

Method 4: Prometheus Monitoring Configuration

For production observability stacks:

scrape_configs:
  - job_name: 'mtplx'
    static_configs:
      - targets: ['localhost:8000']
    metrics_path: '/health'
    scheme: http
    scrape_interval: 5s

Key Source Files for Health Implementation

File Purpose
mtplx/server/openai.py FastAPI route registration and payload assembly
mtplx/commands/public.py CLI wrapper (mtplx health)
mtplx/thermal.py Thermal subsystem data provider (_thermal_health_payload)
mtplx/model_scheduler.py Scheduler state export (_mtplx_scheduler_state)
mtplx/daemon_client.py Daemon discovery via health probe (url = f"http://{host}:{port}/health")
mtplx/dashboard/__init__.py Web UI health verification on startup

Interpreting Health Response Fields

Critical Fields for Load Balancers

  • runtime_mode: Expected values include loaded, loading, or unavailable
  • active_requests: Current inference backlog—use for capacity-based routing
  • idle_seconds: Time since last request; useful for scaling decisions

Diagnostic Fields for Troubleshooting

  • kernel_selfcheck: Boolean indicating kernel module integrity
  • gpu_keepalive: GPU communication channel status
  • smart_fan_status: Thermal management operational state

Configuration Verification

  • api_key_required: Whether authentication is enforced
  • verify_core: Core file validation toggle
  • enable_thinking: Extended reasoning mode availability

Production Health Check Strategies

  • Load balancer health probes: 5-10 seconds with 2-second timeout
  • Kubernetes liveness checks: 10-second period, 3-failure threshold
  • Detailed monitoring dashboards: 1-5 seconds for real-time visibility

Timeout Configuration Guidelines

Client Type Recommended Timeout Rationale
CLI interactive use 5.0 seconds User tolerance for slow responses
Automated scripts 2.0 seconds Balance reliability vs. speed
Service mesh / mesh proxies 1.5 seconds Match MTPLX CLI default
Aggressive load balancers 1.0 second Fast failover requirements

Summary

  • The MTPLX /health endpoint provides comprehensive runtime status through a single HTTP GET request to http://<host>:<port>/health
  • Response payload includes model state, scheduler statistics, thermal metrics, and feature flags sourced from mtplx/server/openai.py lines 28905-28999
  • Zero inference overhead design permits frequent polling without performance impact
  • Multiple access methods: direct curl, Python requests, native mtplx health CLI, or Prometheus scrape configuration
  • Security-hardened response excludes actual secrets while exposing configuration boolean flags

Frequently Asked Questions

What port does the MTPLX health endpoint use?

The default port is 8000, matching the standard FastAPI server binding in mtplx/server/openai.py. This can be configured at server startup through environment variables or command-line arguments.

Does the health check endpoint require authentication?

Authentication depends on runtime configuration. If api_key_required is true in the health response, the server enforces API key validation on inference endpoints—the /health endpoint itself remains accessible for load balancer probes. Check the api_key_required field in your health response to determine current policy.

Why does my health check show runtime_mode: "loading"?

This indicates the server has started but the model weights are not yet fully loaded into GPU memory. Depending on model size and storage speed, this state may persist from several seconds to minutes. Inference requests will queue during this period; scale your timeout thresholds accordingly or implement a readiness gate that waits for runtime_mode: "loaded".

Can I use the health endpoint for automatic failover?

Yes—the lightweight design supports sub-second polling intervals ideal for load balancer health checks. Key fields for failover logic include runtime_mode (fail if not loaded), kernel_selfcheck (fail if false), and gpu_keepalive (fail if disconnected). Combine these with response time thresholds to build robust availability detection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →