What Metrics Are Available Through the GET /metrics Endpoint of the MTPLX API Server

The GET /metrics endpoint returns a JSON snapshot containing three top-level objects—latest for current turn metrics, recent for historical request data, and tool_parse_counters for tool-parsing diagnostics—enabling lightweight, real-time observability of the MTPLX inference server without exposing private model data.

The MTPLX API server provides an OpenAI-compatible inference layer with built-in observability hooks for monitoring runtime performance. According to the source code in mtplx/server/openai.py, the GET /metrics route assembles a concise KPI payload that aggregates per-turn statistics from the in-memory dashboard state, offering developers instant visibility into latency, token throughput, and tool-use reliability.

The Three Core Metric Categories

The JSON response from GET /metrics is structured into three distinct top-level fields, each serving a specific observability purpose. These fields are assembled in mtplx/server/openai.py and populated by mtplx/server/dashboard_state.py.

Latest Turn Metrics (latest)

The latest object captures real-time statistics for the most recent request/turn, providing immediate insight into current system performance. This field includes latency measurements in milliseconds, token input/output counts, and CPU/GPU utilization metrics gathered during the active inference pass.

According to mtplx/server/dashboard_state.py, this data is updated continuously as the server processes each request, making it ideal for monitoring the immediate health of individual inference calls.

Historical Performance Buffer (recent)

The recent array maintains a ring buffer of the last 32 turn metrics, where each entry mirrors the structure of the latest object. This circular buffer allows clients to detect performance trends, identify latency spikes, or correlate slowdowns with specific request patterns over time.

The implementation in mtplx/server/dashboard_state.py manages this buffer automatically, discarding the oldest entry once the 32-item limit is reached. This provides a lightweight alternative to full logging infrastructure while retaining enough history for meaningful trend analysis.

Tool Parsing Diagnostics (tool_parse_counters)

The tool_parse_counters object tracks tool-use reliability metrics critical for debugging agentic workflows. This field contains counters for:

  • successful: Properly parsed tool invocations
  • malformed: Tool calls with syntactic errors
  • unrecognized: Calls referencing undefined tools
  • skipped: Tool markup that was bypassed

These counters are incremented in mtplx/server/openai.py whenever a request containing tool markup is processed. Developers use these metrics to determine how often the server falls back to plain text generation versus successfully executing tool calls.

Accessing the Metrics Endpoint

The endpoint is accessible via standard HTTP clients and supports both snapshot polling and real-time streaming.

Basic HTTP Request

Fetch the current metrics snapshot using curl:

curl http://127.0.0.1:8000/metrics

Or using Python's requests library:

import requests

resp = requests.get("http://127.0.0.1:8000/metrics")
metrics = resp.json()

print("Latest latency (ms):", metrics["latest"]["latency_ms"])
print("Recent turn count:", len(metrics["recent"]))
print("Tool parse successes:", metrics["tool_parse_counters"]["successful"])

Server-Sent Events Streaming

For dashboard applications requiring live updates, the server exposes an SSE stream endpoint:

curl http://127.0.0.1:8000/v1/mtplx/metrics/stream?snapshot_interval_ms=500

This streaming interface, documented in docs/dashboard.md, pushes metric snapshots at configurable intervals (defaulting to 500ms), enabling real-time visualization without repeated polling overhead.

Implementation Architecture

Understanding the source code locations helps developers extend or debug the metrics pipeline:

  • mtplx/server/openai.py: Registers both GET /metrics and GET /v1/mtplx/metrics/stream, assembles the JSON payload, and updates tool parsing counters during request processing.
  • mtplx/server/dashboard_state.py: Maintains the in-memory structures for latest turn statistics and the circular buffer for recent history.
  • mtplx/server/flight_recorder.py: Underlies the telemetry infrastructure by managing optional JSONL flight-recorder files, providing persistent storage complementary to the ephemeral /metrics endpoint.
  • docs/api.md: Contains the canonical API reference documentation for the three top-level response fields.

Summary

  • The GET /metrics endpoint provides a JSON snapshot with three fields: latest for current turn data, recent for the last 32 historical entries, and tool_parse_counters for tool-use reliability.
  • Latency and token counts are tracked per-turn and aggregated in the dashboard state maintained by mtplx/server/dashboard_state.py.
  • Tool parsing metrics help identify malformed tool calls and fallback patterns during agentic workflows.
  • The endpoint supports both polling and SSE streaming for flexible integration with monitoring systems.
  • All metrics are exposed without revealing private model weights or training data, ensuring secure observability.

Frequently Asked Questions

How many historical entries does the recent array store?

The recent field maintains a fixed-size ring buffer containing a maximum of 32 historical turn entries. Once this limit is reached, the server discards the oldest entry to make room for new metrics, ensuring constant memory usage regardless of uptime.

What purpose do the tool_parse_counters serve?

The tool_parse_counters object provides diagnostic visibility into tool-use reliability by tracking successful, malformed, unrecognized, and skipped tool calls. Developers use these counters to fine-tune tool schemas, identify prompt engineering issues, or measure how frequently the system falls back to plain text generation instead of executing tools.

Is the metrics endpoint compatible with Prometheus?

While the endpoint returns JSON rather than Prometheus exposition format, the structured data in mtplx/server/openai.py can be easily adapted for Prometheus scraping. The latest and recent fields expose standard KPIs like latency milliseconds and token counts that map directly to Prometheus gauge and histogram metrics.

What is the difference between /metrics and /v1/mtplx/metrics/stream?

The standard GET /metrics endpoint returns a point-in-time JSON snapshot suitable for periodic polling, while GET /v1/mtplx/metrics/stream establishes a Server-Sent Events connection that pushes updated snapshots at configurable intervals (controlled via the snapshot_interval_ms parameter). The streaming endpoint is optimized for live dashboards requiring real-time updates without connection overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →