# How to Get Server Metrics from MTPLX: Complete API Guide with Code Examples

> Learn to get server metrics from MTPLX using its API. Explore code examples for snapshot and real-time streaming endpoints, accessing token throughput, acceptance rates, and request statistics.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: api-reference
- Published: 2026-09-06

---

**MTPLX exposes server metrics through two HTTP endpoints—`/metrics` for snapshots and `/v1/mtplx/metrics/stream` for real-time streaming—both returning JSON data about token throughput, acceptance rates, and request statistics.**

The MTPLX inference server, implemented in the youssofal/MTPLX repository, provides built-in observability through a FastAPI-based metrics API. Whether you need a one-time health check or a live dashboard feed, the server exposes performance data without requiring external monitoring tools. This guide covers both endpoints, their underlying implementation, and practical code examples for accessing **server metrics from MTPLX**.

## Available Metrics Endpoints

MTPLX registers two routes in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) that serve different monitoring needs:

| Route | Method | Use Case |
|-------|--------|----------|
| `/metrics` | GET | Retrieve the latest snapshot of server-wide metrics |
| `/v1/mtplx/metrics/stream` | GET | Stream metric updates in real time for live dashboards |

Both endpoints return JSON payloads constructed by the internal `_metrics_envelope` helper function.

### What Metrics Are Collected

When processing requests, the server records granular performance data through `_record_request_metrics` and `_merge_final_bridge_stats_into_latest_metrics`. These populate `state.last_metrics`, a ring buffer that powers both endpoints.

Key fields in each metrics entry include:

- **`tok_s`** — tokens generated per second for the request
- **`accept_rate`** — proportion of tokens accepted by the speculative decoding pipeline
- **`request_id`** — unique identifier for request correlation
- **`session_prompt_prefix_commit`** — snapshot of prefix caching state
- **`session_postcommit_snapshot`** — post-processing execution pipeline state

## Fetching a Metrics Snapshot

The `/metrics` endpoint returns the most recent entry from the internal ring buffer, wrapped in a structured envelope.

### Using curl

```bash

# Get latest metrics as formatted JSON

curl -s http://localhost:8000/metrics | jq .

```

Typical response structure:

```json
{
  "latest": {
    "tok_s": 142.7,
    "accept_rate": 0.89,
    "request_id": "req_7a3f9d2e1b",
    "timestamp": "2024-01-15T09:23:47.123456"
  },
  "history": []
}

```

### Using Python requests

```python
import requests

BASE_URL = "http://localhost:8000"

resp = requests.get(f"{BASE_URL}/metrics")
data = resp.json()

print(f"Tokens/sec: {data['latest']['tok_s']}")
print(f"Accept rate: {data['latest']['accept_rate']:.1%}")

```

## Streaming Real-Time Metrics

For monitoring dashboards or continuous health tracking, the streaming endpoint yields each new metric entry as `state.last_metrics` is updated.

### Using curl with streaming

```bash

# Stream metrics until interrupted (Ctrl-C)

curl -N http://localhost:8000/v1/mtplx/metrics/stream | jq .

```

### Using Python with Server-Sent Events

```python
import requests
import json

BASE_URL = "http://localhost:8000"

with requests.get(
    f"{BASE_URL}/v1/mtplx/metrics/stream",
    stream=True
) as r:
    for line in r.iter_lines():
        if line and line.startswith(b"data: "):
            payload = json.loads(line[6:])  # Strip "data: " prefix

            print(f"tok_s={payload['tok_s']:.1f}, "
                  f"accept_rate={payload['accept_rate']:.2f}")

```

## Implementation Details in the Source Code

Understanding the internal architecture helps when building custom monitoring integrations.

### Endpoint Registration

In [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), the routes attach to the FastAPI `app` instance:

```python
@app.get("/metrics")
@app.get("/v1/mtplx/metrics/stream")

```

The non-streaming handler calls `_metrics_envelope(state)` and returns its result directly. The streaming handler uses an async generator to yield newline-delimited JSON as new metrics arrive.

### Metric Collection Pipeline

After each request completes, `_record_request_metrics` pushes a dictionary onto `state.last_metrics`. The later call to `_merge_final_bridge_stats_into_latest_metrics` enriches this with final execution statistics before the endpoint can serve it.

### The Envelope Function

`_metrics_envelope(state)` constructs the response format you receive, placing the newest entry under the `latest` key and optionally including historical data. This abstraction allows the same payload structure for both snapshot and streaming consumers.

## Using the Bundled Example Client

The repository includes a reference implementation in [`examples/openai-python-client.py`](https://github.com/youssofal/MTPLX/blob/main/examples/openai-python-client.py) that demonstrates idiomatic client usage:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-dummy")

# Access raw metrics via the underlying HTTP client

metrics = client.get("/metrics")
print(metrics.json())

```

This pattern is verified in [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py), which asserts correct payload structure and HTTP status codes.

## Key Files Reference

| File | Purpose |
|------|---------|
| [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) | FastAPI server with `/metrics` and streaming route implementations; contains `_metrics_envelope`, `_record_request_metrics`, `_merge_final_bridge_stats_into_latest_metrics` |
| [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) | Unit tests validating endpoint behavior and response schemas |
| [`docs/server.md`](https://github.com/youssofal/MTPLX/blob/main/docs/server.md) | High-level API documentation covering all server endpoints |
| [`examples/openai-python-client.py`](https://github.com/youssofal/MTPLX/blob/main/examples/openai-python-client.py) | Reference Python client showing metrics access patterns |

## Summary

- **Two endpoints serve MTPLX metrics**: `/metrics` for snapshots, `/v1/mtplx/metrics/stream` for real-time feeds
- **Core metrics include** `tok_s`, `accept_rate`, and request pipeline state captured per completion
- **Implementation lives in** [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) through `_record_request_metrics` and `_metrics_envelope`
- **Both curl and Python clients work**; the streaming endpoint uses standard HTTP streaming for dashboard integration
- **Tests and examples** in [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) and [`examples/openai-python-client.py`](https://github.com/youssofal/MTPLX/blob/main/examples/openai-python-client.py) provide verified reference code

## Frequently Asked Questions

### What port does the MTPLX metrics API use?

By default, the MTPLX server binds to port 8000. The metrics endpoints are accessible at `http://localhost:8000/metrics` and `http://localhost:8000/v1/mtplx/metrics/stream`. Check your server startup logs or configuration if running on a non-default port.

### Do I need authentication to access MTPLX metrics?

The open-source implementation in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) does not enforce authentication on the metrics endpoints by default. For production deployments, place the server behind a reverse proxy with appropriate access controls, as the metrics may expose internal performance characteristics.

### Can I access historical metrics beyond the latest snapshot?

The `/metrics` endpoint includes a `history` array in its envelope, though population depends on the `state.last_metrics` ring buffer configuration. For persistent historical analysis, poll the endpoint periodically and store results externally, or consume the streaming endpoint and archive the feed.

### How accurate is the `tok_s` measurement?

The `tok_s` value is computed per-request in `_record_request_metrics` based on actual token generation duration. It reflects end-to-end throughput including speculative decoding overhead, making it suitable for capacity planning and performance regression detection.