# What Metrics Are Available Through the GET /metrics Endpoint of the MTPLX API Server

> Discover the metrics available via the MTPLX API GET /metrics endpoint. Access current turn data, historical requests, and tool-parsing diagnostics for real-time server observability.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: api-reference
- Published: 2026-09-05

---

**The GET /metrics endpoint returns a JSON snapshot containing three top-level objects—`latest` for current turn metrics, `recent` for historical request data, and `tool_parse_counters` for tool-parsing diagnostics—enabling lightweight, real-time observability of the MTPLX inference server without exposing private model data.**

The **MTPLX API server** provides an OpenAI-compatible inference layer with built-in observability hooks for monitoring runtime performance. According to the source code in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), the `GET /metrics` route assembles a concise KPI payload that aggregates per-turn statistics from the in-memory dashboard state, offering developers instant visibility into latency, token throughput, and tool-use reliability.

## The Three Core Metric Categories

The JSON response from `GET /metrics` is structured into three distinct top-level fields, each serving a specific observability purpose. These fields are assembled in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) and populated by [`mtplx/server/dashboard_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/dashboard_state.py).

### Latest Turn Metrics (`latest`)

The `latest` object captures **real-time statistics for the most recent request/turn**, providing immediate insight into current system performance. This field includes latency measurements in milliseconds, token input/output counts, and CPU/GPU utilization metrics gathered during the active inference pass.

According to [`mtplx/server/dashboard_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/dashboard_state.py), this data is updated continuously as the server processes each request, making it ideal for monitoring the immediate health of individual inference calls.

### Historical Performance Buffer (`recent`)

The `recent` array maintains a **ring buffer of the last 32 turn metrics**, where each entry mirrors the structure of the `latest` object. This circular buffer allows clients to detect performance trends, identify latency spikes, or correlate slowdowns with specific request patterns over time.

The implementation in [`mtplx/server/dashboard_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/dashboard_state.py) manages this buffer automatically, discarding the oldest entry once the 32-item limit is reached. This provides a lightweight alternative to full logging infrastructure while retaining enough history for meaningful trend analysis.

### Tool Parsing Diagnostics (`tool_parse_counters`)

The `tool_parse_counters` object tracks **tool-use reliability metrics** critical for debugging agentic workflows. This field contains counters for:
- `successful`: Properly parsed tool invocations
- `malformed`: Tool calls with syntactic errors
- `unrecognized`: Calls referencing undefined tools
- `skipped`: Tool markup that was bypassed

These counters are incremented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) whenever a request containing tool markup is processed. Developers use these metrics to determine how often the server falls back to plain text generation versus successfully executing tool calls.

## Accessing the Metrics Endpoint

The endpoint is accessible via standard HTTP clients and supports both snapshot polling and real-time streaming.

### Basic HTTP Request

Fetch the current metrics snapshot using `curl`:

```bash
curl http://127.0.0.1:8000/metrics

```

Or using Python's `requests` library:

```python
import requests

resp = requests.get("http://127.0.0.1:8000/metrics")
metrics = resp.json()

print("Latest latency (ms):", metrics["latest"]["latency_ms"])
print("Recent turn count:", len(metrics["recent"]))
print("Tool parse successes:", metrics["tool_parse_counters"]["successful"])

```

### Server-Sent Events Streaming

For dashboard applications requiring live updates, the server exposes an SSE stream endpoint:

```bash
curl http://127.0.0.1:8000/v1/mtplx/metrics/stream?snapshot_interval_ms=500

```

This streaming interface, documented in [`docs/dashboard.md`](https://github.com/youssofal/MTPLX/blob/main/docs/dashboard.md), pushes metric snapshots at configurable intervals (defaulting to 500ms), enabling real-time visualization without repeated polling overhead.

## Implementation Architecture

Understanding the source code locations helps developers extend or debug the metrics pipeline:

- **[`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)**: Registers both `GET /metrics` and `GET /v1/mtplx/metrics/stream`, assembles the JSON payload, and updates tool parsing counters during request processing.
- **[`mtplx/server/dashboard_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/dashboard_state.py)**: Maintains the in-memory structures for `latest` turn statistics and the circular buffer for `recent` history.
- **[`mtplx/server/flight_recorder.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/flight_recorder.py)**: Underlies the telemetry infrastructure by managing optional JSONL flight-recorder files, providing persistent storage complementary to the ephemeral `/metrics` endpoint.
- **[`docs/api.md`](https://github.com/youssofal/MTPLX/blob/main/docs/api.md)**: Contains the canonical API reference documentation for the three top-level response fields.

## Summary

- The `GET /metrics` endpoint provides a JSON snapshot with three fields: `latest` for current turn data, `recent` for the last 32 historical entries, and `tool_parse_counters` for tool-use reliability.
- **Latency and token counts** are tracked per-turn and aggregated in the dashboard state maintained by [`mtplx/server/dashboard_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/dashboard_state.py).
- **Tool parsing metrics** help identify malformed tool calls and fallback patterns during agentic workflows.
- The endpoint supports both **polling and SSE streaming** for flexible integration with monitoring systems.
- All metrics are exposed without revealing private model weights or training data, ensuring secure observability.

## Frequently Asked Questions

### How many historical entries does the `recent` array store?

The `recent` field maintains a fixed-size ring buffer containing a maximum of **32 historical turn entries**. Once this limit is reached, the server discards the oldest entry to make room for new metrics, ensuring constant memory usage regardless of uptime.

### What purpose do the tool_parse_counters serve?

The `tool_parse_counters` object provides **diagnostic visibility into tool-use reliability** by tracking successful, malformed, unrecognized, and skipped tool calls. Developers use these counters to fine-tune tool schemas, identify prompt engineering issues, or measure how frequently the system falls back to plain text generation instead of executing tools.

### Is the metrics endpoint compatible with Prometheus?

While the endpoint returns JSON rather than Prometheus exposition format, the structured data in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) can be easily adapted for Prometheus scraping. The `latest` and `recent` fields expose standard KPIs like latency milliseconds and token counts that map directly to Prometheus gauge and histogram metrics.

### What is the difference between `/metrics` and `/v1/mtplx/metrics/stream`?

The standard `GET /metrics` endpoint returns a **point-in-time JSON snapshot** suitable for periodic polling, while `GET /v1/mtplx/metrics/stream` establishes a **Server-Sent Events connection** that pushes updated snapshots at configurable intervals (controlled via the `snapshot_interval_ms` parameter). The streaming endpoint is optimized for live dashboards requiring real-time updates without connection overhead.