MTPLX GET /health Endpoint: Complete Runtime State Reference
The GET /health endpoint returns a comprehensive JSON snapshot of the MTPLX API server's runtime state, including the loaded model identifier, active sampling parameters, API key requirements, KV cache quantization mode, and real-time concurrency metrics—containing no sensitive secrets and designed for monitoring dashboards.
The GET /health endpoint in the youssofal/MTPLX repository serves as the primary diagnostic interface for the FastAPI server. Located in mtplx/server/openai.py, this route aggregates telemetry from the model profile, backend configuration, and request scheduler into a single public-safe payload that health monitors, load balancers, and benchmarking tools consume to assess server readiness.
Health Check Response Structure
When you query GET /health, MTPLX assembles a JSON object containing nine distinct top-level sections. Each field provides visibility into a specific subsystem without exposing secrets like raw API keys or model weights.
Model Identity and Sampling Configuration
The model field returns the string identifier of the currently loaded model (e.g., "mtplx-test-model"). Nested within the profile object, you find model_id and profile_default_model_id distinguishing the active model from the default profile. The profile.sampler sub-object exposes the generation hyperparameters currently in effect, including temperature, top_p, and top_k values.
Authentication and Security Flags
Two boolean and string fields govern access control:
api_key_required: Boolean indicating if requests must include a bearer token.api_key_source: String describing the key origin (e.g.,"none"when authentication is disabled).
According to the test suite in tests/test_server_openai.py, these fields are validated to ensure the endpoint never leaks credential material while accurately reflecting security posture.
Memory Management and Quantization
The paged_kv_quantization field reports the KV cache compression mode as "off", "q4", or "q8", revealing how the server manages attention memory. Backend-specific implementations in mtplx/backends/qwen3_next.py (lines 39-45) and mtplx/backends/step3p5_mtp.py contribute these values based on the active model family’s capabilities.
Runtime Concurrency and Thermal State
Real-time operational metrics include:
active_requests: Integer count of in-flight generation jobs (normally0when idle).foreground_active: Integer count of foreground-priority tasks.last_request_started_at: Unix timestamp (seconds) of the most recent request start; returns0.0when the server has processed no requests since startup.thermal.max_requested: Boolean flag indicating whether "max-thermal" performance mode has been requested.
Startup Metadata and Special Modes
The nested startup object contains process metadata populated during initialization:
pid: The server process ID.warmup.ran: Boolean indicating whether the model warm-up phase completed.model_controls: Object detailingmodel_family(e.g.,"qwen3_8") anddraft_control.maximumlimits for speculative decoding.tool_prompt_modeandtool_contract_active: Strings and booleans describing the current tool-use configuration.
For deployments using the "opencode" short-context optimization, the payload includes opencode_short_context_depth2_tokens (nullable integer) and opencode_short_context_depth_policy (object with type and depth fields), as verified in tests/test_server_openai.py around lines 2516-2550.
How to Call the GET /health Endpoint
Basic curl Request
Query the endpoint from any HTTP client:
curl http://127.0.0.1:8000/health
Sample JSON Response:
{
"model": "mtplx-test-model",
"profile": {
"model_id": "mtplx-test-model",
"profile_default_model_id": "mtplx-default",
"sampler": { "temperature": 0.6, "top_p": 0.95, "top_k": 40 }
},
"api_key_required": false,
"api_key_source": "none",
"paged_kv_quantization": "q4",
"warmup": { "ran": false },
"startup": {
"pid": 12345,
"warmup": { "ran": false },
"api_key_source": "none",
"model_controls": {
"model_family": "qwen3_8",
"draft_control": { "maximum": 3 }
},
"tool_prompt_mode": "hybrid",
"tool_contract_active": true
},
"thermal": { "max_requested": false },
"foreground_active": 0,
"active_requests": 0,
"last_request_started_at": 0.0,
"opencode_short_context_depth2_tokens": null,
"opencode_short_context_depth_policy": { "type": "fixed", "depth": 2 }
}
Python Client Integration
Use requests to programmatically check server status:
import requests
resp = requests.get("http://127.0.0.1:8000/health")
health = resp.json()
print(f"Loaded model: {health['model']}")
print(f"Active generations: {health['active_requests']}")
print(f"Sampler temp: {health['profile']['sampler']['temperature']}")
# Check authentication requirements before sending sensitive prompts
if health["api_key_required"]:
print("Warning: API key required for this endpoint")
else:
print("No authentication required")
Source Code Implementation
The health payload assembly occurs in mtplx/server/openai.py between lines 28630-28645, where the FastAPI route constructs the dictionary from runtime state objects. Backend modules such as mtplx/backends/qwen3_next.py inject family-specific values like draft control limits into the startup.model_controls section.
The response schema is regression-tested in tests/test_server_openai.py (lines 2505-2525), which asserts that health.json()["model"] matches the expected identifier and validates the sampler configuration structure at lines 2510-2515. These tests ensure that changes to the health payload remain backward compatible for monitoring integrations.
Summary
- The MTPLX
GET /healthendpoint exposes a JSON object containing model identity, sampling parameters, authentication requirements, and runtime metrics. - Security-safe design means the endpoint reveals whether keys are required (
api_key_required) but never exposes the keys themselves. - Performance visibility includes KV cache quantization mode (
paged_kv_quantization), active request counts, and thermal throttling status. - Implementation resides in
mtplx/server/openai.pywith backend contributions from files likemtplx/backends/qwen3_next.py, verified bytests/test_server_openai.py.
Frequently Asked Questions
Does the MTPLX GET /health endpoint expose sensitive API keys?
No. While the response includes api_key_required (boolean) and api_key_source (string indicating origin like "none"), it deliberately omits the actual secret material. This design allows public exposure for load balancer health checks without security risks.
How can I determine if the model has finished warming up?
Check the warmup.ran field inside both the top-level and startup objects. A value of true indicates the initialization and warm-up sequence completed successfully. The test suite at tests/test_server_openai.py specifically validates this flag to ensure accurate readiness reporting.
What does the paged_kv_quantization field indicate?
This field reports the compression level applied to the paged key-value cache, returning "off", "q4", or "q8". This value originates from backend implementations such as mtplx/backends/qwen3_next.py and helps operators verify memory optimization settings for the loaded model family.
Why are active_requests and foreground_active sometimes zero when jobs are running?
These counters reflect instantaneous states sampled at request time. If you query /health between generation tokens or during scheduler idle cycles, transient counts may show zero. For precise concurrency tracking, poll the endpoint or correlate with last_request_started_at timestamps to detect recent activity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →