What Metrics Are Available Through the GET /metrics Endpoint of the MTPLX API Server
The GET /metrics endpoint returns a JSON snapshot containing three top-level objects—latest for current turn metrics, recent for historical request data, and tool_parse_counters for tool-parsing diagnostics—enabling lightweight, real-time observability of the MTPLX inference server without exposing private model data.
The MTPLX API server provides an OpenAI-compatible inference layer with built-in observability hooks for monitoring runtime performance. According to the source code in mtplx/server/openai.py, the GET /metrics route assembles a concise KPI payload that aggregates per-turn statistics from the in-memory dashboard state, offering developers instant visibility into latency, token throughput, and tool-use reliability.
The Three Core Metric Categories
The JSON response from GET /metrics is structured into three distinct top-level fields, each serving a specific observability purpose. These fields are assembled in mtplx/server/openai.py and populated by mtplx/server/dashboard_state.py.
Latest Turn Metrics (latest)
The latest object captures real-time statistics for the most recent request/turn, providing immediate insight into current system performance. This field includes latency measurements in milliseconds, token input/output counts, and CPU/GPU utilization metrics gathered during the active inference pass.
According to mtplx/server/dashboard_state.py, this data is updated continuously as the server processes each request, making it ideal for monitoring the immediate health of individual inference calls.
Historical Performance Buffer (recent)
The recent array maintains a ring buffer of the last 32 turn metrics, where each entry mirrors the structure of the latest object. This circular buffer allows clients to detect performance trends, identify latency spikes, or correlate slowdowns with specific request patterns over time.
The implementation in mtplx/server/dashboard_state.py manages this buffer automatically, discarding the oldest entry once the 32-item limit is reached. This provides a lightweight alternative to full logging infrastructure while retaining enough history for meaningful trend analysis.
Tool Parsing Diagnostics (tool_parse_counters)
The tool_parse_counters object tracks tool-use reliability metrics critical for debugging agentic workflows. This field contains counters for:
successful: Properly parsed tool invocationsmalformed: Tool calls with syntactic errorsunrecognized: Calls referencing undefined toolsskipped: Tool markup that was bypassed
These counters are incremented in mtplx/server/openai.py whenever a request containing tool markup is processed. Developers use these metrics to determine how often the server falls back to plain text generation versus successfully executing tool calls.
Accessing the Metrics Endpoint
The endpoint is accessible via standard HTTP clients and supports both snapshot polling and real-time streaming.
Basic HTTP Request
Fetch the current metrics snapshot using curl:
curl http://127.0.0.1:8000/metrics
Or using Python's requests library:
import requests
resp = requests.get("http://127.0.0.1:8000/metrics")
metrics = resp.json()
print("Latest latency (ms):", metrics["latest"]["latency_ms"])
print("Recent turn count:", len(metrics["recent"]))
print("Tool parse successes:", metrics["tool_parse_counters"]["successful"])
Server-Sent Events Streaming
For dashboard applications requiring live updates, the server exposes an SSE stream endpoint:
curl http://127.0.0.1:8000/v1/mtplx/metrics/stream?snapshot_interval_ms=500
This streaming interface, documented in docs/dashboard.md, pushes metric snapshots at configurable intervals (defaulting to 500ms), enabling real-time visualization without repeated polling overhead.
Implementation Architecture
Understanding the source code locations helps developers extend or debug the metrics pipeline:
mtplx/server/openai.py: Registers bothGET /metricsandGET /v1/mtplx/metrics/stream, assembles the JSON payload, and updates tool parsing counters during request processing.mtplx/server/dashboard_state.py: Maintains the in-memory structures forlatestturn statistics and the circular buffer forrecenthistory.mtplx/server/flight_recorder.py: Underlies the telemetry infrastructure by managing optional JSONL flight-recorder files, providing persistent storage complementary to the ephemeral/metricsendpoint.docs/api.md: Contains the canonical API reference documentation for the three top-level response fields.
Summary
- The
GET /metricsendpoint provides a JSON snapshot with three fields:latestfor current turn data,recentfor the last 32 historical entries, andtool_parse_countersfor tool-use reliability. - Latency and token counts are tracked per-turn and aggregated in the dashboard state maintained by
mtplx/server/dashboard_state.py. - Tool parsing metrics help identify malformed tool calls and fallback patterns during agentic workflows.
- The endpoint supports both polling and SSE streaming for flexible integration with monitoring systems.
- All metrics are exposed without revealing private model weights or training data, ensuring secure observability.
Frequently Asked Questions
How many historical entries does the recent array store?
The recent field maintains a fixed-size ring buffer containing a maximum of 32 historical turn entries. Once this limit is reached, the server discards the oldest entry to make room for new metrics, ensuring constant memory usage regardless of uptime.
What purpose do the tool_parse_counters serve?
The tool_parse_counters object provides diagnostic visibility into tool-use reliability by tracking successful, malformed, unrecognized, and skipped tool calls. Developers use these counters to fine-tune tool schemas, identify prompt engineering issues, or measure how frequently the system falls back to plain text generation instead of executing tools.
Is the metrics endpoint compatible with Prometheus?
While the endpoint returns JSON rather than Prometheus exposition format, the structured data in mtplx/server/openai.py can be easily adapted for Prometheus scraping. The latest and recent fields expose standard KPIs like latency milliseconds and token counts that map directly to Prometheus gauge and histogram metrics.
What is the difference between /metrics and /v1/mtplx/metrics/stream?
The standard GET /metrics endpoint returns a point-in-time JSON snapshot suitable for periodic polling, while GET /v1/mtplx/metrics/stream establishes a Server-Sent Events connection that pushes updated snapshots at configurable intervals (controlled via the snapshot_interval_ms parameter). The streaming endpoint is optimized for live dashboards requiring real-time updates without connection overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →