MTPLX GET /health Endpoint: Complete Runtime State Reference

The GET /health endpoint returns a comprehensive JSON snapshot of the MTPLX API server's runtime state, including the loaded model identifier, active sampling parameters, API key requirements, KV cache quantization mode, and real-time concurrency metrics—containing no sensitive secrets and designed for monitoring dashboards.

The GET /health endpoint in the youssofal/MTPLX repository serves as the primary diagnostic interface for the FastAPI server. Located in mtplx/server/openai.py, this route aggregates telemetry from the model profile, backend configuration, and request scheduler into a single public-safe payload that health monitors, load balancers, and benchmarking tools consume to assess server readiness.

Health Check Response Structure

When you query GET /health, MTPLX assembles a JSON object containing nine distinct top-level sections. Each field provides visibility into a specific subsystem without exposing secrets like raw API keys or model weights.

Model Identity and Sampling Configuration

The model field returns the string identifier of the currently loaded model (e.g., "mtplx-test-model"). Nested within the profile object, you find model_id and profile_default_model_id distinguishing the active model from the default profile. The profile.sampler sub-object exposes the generation hyperparameters currently in effect, including temperature, top_p, and top_k values.

Authentication and Security Flags

Two boolean and string fields govern access control:

  • api_key_required: Boolean indicating if requests must include a bearer token.
  • api_key_source: String describing the key origin (e.g., "none" when authentication is disabled).

According to the test suite in tests/test_server_openai.py, these fields are validated to ensure the endpoint never leaks credential material while accurately reflecting security posture.

Memory Management and Quantization

The paged_kv_quantization field reports the KV cache compression mode as "off", "q4", or "q8", revealing how the server manages attention memory. Backend-specific implementations in mtplx/backends/qwen3_next.py (lines 39-45) and mtplx/backends/step3p5_mtp.py contribute these values based on the active model family’s capabilities.

Runtime Concurrency and Thermal State

Real-time operational metrics include:

  • active_requests: Integer count of in-flight generation jobs (normally 0 when idle).
  • foreground_active: Integer count of foreground-priority tasks.
  • last_request_started_at: Unix timestamp (seconds) of the most recent request start; returns 0.0 when the server has processed no requests since startup.
  • thermal.max_requested: Boolean flag indicating whether "max-thermal" performance mode has been requested.

Startup Metadata and Special Modes

The nested startup object contains process metadata populated during initialization:

  • pid: The server process ID.
  • warmup.ran: Boolean indicating whether the model warm-up phase completed.
  • model_controls: Object detailing model_family (e.g., "qwen3_8") and draft_control.maximum limits for speculative decoding.
  • tool_prompt_mode and tool_contract_active: Strings and booleans describing the current tool-use configuration.

For deployments using the "opencode" short-context optimization, the payload includes opencode_short_context_depth2_tokens (nullable integer) and opencode_short_context_depth_policy (object with type and depth fields), as verified in tests/test_server_openai.py around lines 2516-2550.

How to Call the GET /health Endpoint

Basic curl Request

Query the endpoint from any HTTP client:

curl http://127.0.0.1:8000/health

Sample JSON Response:

{
  "model": "mtplx-test-model",
  "profile": {
    "model_id": "mtplx-test-model",
    "profile_default_model_id": "mtplx-default",
    "sampler": { "temperature": 0.6, "top_p": 0.95, "top_k": 40 }
  },
  "api_key_required": false,
  "api_key_source": "none",
  "paged_kv_quantization": "q4",
  "warmup": { "ran": false },
  "startup": {
    "pid": 12345,
    "warmup": { "ran": false },
    "api_key_source": "none",
    "model_controls": {
      "model_family": "qwen3_8",
      "draft_control": { "maximum": 3 }
    },
    "tool_prompt_mode": "hybrid",
    "tool_contract_active": true
  },
  "thermal": { "max_requested": false },
  "foreground_active": 0,
  "active_requests": 0,
  "last_request_started_at": 0.0,
  "opencode_short_context_depth2_tokens": null,
  "opencode_short_context_depth_policy": { "type": "fixed", "depth": 2 }
}

Python Client Integration

Use requests to programmatically check server status:

import requests

resp = requests.get("http://127.0.0.1:8000/health")
health = resp.json()

print(f"Loaded model: {health['model']}")
print(f"Active generations: {health['active_requests']}")
print(f"Sampler temp: {health['profile']['sampler']['temperature']}")

# Check authentication requirements before sending sensitive prompts

if health["api_key_required"]:
    print("Warning: API key required for this endpoint")
else:
    print("No authentication required")

Source Code Implementation

The health payload assembly occurs in mtplx/server/openai.py between lines 28630-28645, where the FastAPI route constructs the dictionary from runtime state objects. Backend modules such as mtplx/backends/qwen3_next.py inject family-specific values like draft control limits into the startup.model_controls section.

The response schema is regression-tested in tests/test_server_openai.py (lines 2505-2525), which asserts that health.json()["model"] matches the expected identifier and validates the sampler configuration structure at lines 2510-2515. These tests ensure that changes to the health payload remain backward compatible for monitoring integrations.

Summary

  • The MTPLX GET /health endpoint exposes a JSON object containing model identity, sampling parameters, authentication requirements, and runtime metrics.
  • Security-safe design means the endpoint reveals whether keys are required (api_key_required) but never exposes the keys themselves.
  • Performance visibility includes KV cache quantization mode (paged_kv_quantization), active request counts, and thermal throttling status.
  • Implementation resides in mtplx/server/openai.py with backend contributions from files like mtplx/backends/qwen3_next.py, verified by tests/test_server_openai.py.

Frequently Asked Questions

Does the MTPLX GET /health endpoint expose sensitive API keys?

No. While the response includes api_key_required (boolean) and api_key_source (string indicating origin like "none"), it deliberately omits the actual secret material. This design allows public exposure for load balancer health checks without security risks.

How can I determine if the model has finished warming up?

Check the warmup.ran field inside both the top-level and startup objects. A value of true indicates the initialization and warm-up sequence completed successfully. The test suite at tests/test_server_openai.py specifically validates this flag to ensure accurate readiness reporting.

What does the paged_kv_quantization field indicate?

This field reports the compression level applied to the paged key-value cache, returning "off", "q4", or "q8". This value originates from backend implementations such as mtplx/backends/qwen3_next.py and helps operators verify memory optimization settings for the loaded model family.

Why are active_requests and foreground_active sometimes zero when jobs are running?

These counters reflect instantaneous states sampled at request time. If you query /health between generation tokens or during scheduler idle cycles, transient counts may show zero. For precise concurrency tracking, poll the endpoint or correlate with last_request_started_at timestamps to detect recent activity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →