How to Get Server Metrics from MTPLX: Complete API Guide with Code Examples

MTPLX exposes server metrics through two HTTP endpoints—/metrics for snapshots and /v1/mtplx/metrics/stream for real-time streaming—both returning JSON data about token throughput, acceptance rates, and request statistics.

The MTPLX inference server, implemented in the youssofal/MTPLX repository, provides built-in observability through a FastAPI-based metrics API. Whether you need a one-time health check or a live dashboard feed, the server exposes performance data without requiring external monitoring tools. This guide covers both endpoints, their underlying implementation, and practical code examples for accessing server metrics from MTPLX.

Available Metrics Endpoints

MTPLX registers two routes in mtplx/server/openai.py that serve different monitoring needs:

Route Method Use Case
/metrics GET Retrieve the latest snapshot of server-wide metrics
/v1/mtplx/metrics/stream GET Stream metric updates in real time for live dashboards

Both endpoints return JSON payloads constructed by the internal _metrics_envelope helper function.

What Metrics Are Collected

When processing requests, the server records granular performance data through _record_request_metrics and _merge_final_bridge_stats_into_latest_metrics. These populate state.last_metrics, a ring buffer that powers both endpoints.

Key fields in each metrics entry include:

  • tok_s — tokens generated per second for the request
  • accept_rate — proportion of tokens accepted by the speculative decoding pipeline
  • request_id — unique identifier for request correlation
  • session_prompt_prefix_commit — snapshot of prefix caching state
  • session_postcommit_snapshot — post-processing execution pipeline state

Fetching a Metrics Snapshot

The /metrics endpoint returns the most recent entry from the internal ring buffer, wrapped in a structured envelope.

Using curl


# Get latest metrics as formatted JSON

curl -s http://localhost:8000/metrics | jq .

Typical response structure:

{
  "latest": {
    "tok_s": 142.7,
    "accept_rate": 0.89,
    "request_id": "req_7a3f9d2e1b",
    "timestamp": "2024-01-15T09:23:47.123456"
  },
  "history": []
}

Using Python requests

import requests

BASE_URL = "http://localhost:8000"

resp = requests.get(f"{BASE_URL}/metrics")
data = resp.json()

print(f"Tokens/sec: {data['latest']['tok_s']}")
print(f"Accept rate: {data['latest']['accept_rate']:.1%}")

Streaming Real-Time Metrics

For monitoring dashboards or continuous health tracking, the streaming endpoint yields each new metric entry as state.last_metrics is updated.

Using curl with streaming


# Stream metrics until interrupted (Ctrl-C)

curl -N http://localhost:8000/v1/mtplx/metrics/stream | jq .

Using Python with Server-Sent Events

import requests
import json

BASE_URL = "http://localhost:8000"

with requests.get(
    f"{BASE_URL}/v1/mtplx/metrics/stream",
    stream=True
) as r:
    for line in r.iter_lines():
        if line and line.startswith(b"data: "):
            payload = json.loads(line[6:])  # Strip "data: " prefix

            print(f"tok_s={payload['tok_s']:.1f}, "
                  f"accept_rate={payload['accept_rate']:.2f}")

Implementation Details in the Source Code

Understanding the internal architecture helps when building custom monitoring integrations.

Endpoint Registration

In mtplx/server/openai.py, the routes attach to the FastAPI app instance:

@app.get("/metrics")
@app.get("/v1/mtplx/metrics/stream")

The non-streaming handler calls _metrics_envelope(state) and returns its result directly. The streaming handler uses an async generator to yield newline-delimited JSON as new metrics arrive.

Metric Collection Pipeline

After each request completes, _record_request_metrics pushes a dictionary onto state.last_metrics. The later call to _merge_final_bridge_stats_into_latest_metrics enriches this with final execution statistics before the endpoint can serve it.

The Envelope Function

_metrics_envelope(state) constructs the response format you receive, placing the newest entry under the latest key and optionally including historical data. This abstraction allows the same payload structure for both snapshot and streaming consumers.

Using the Bundled Example Client

The repository includes a reference implementation in examples/openai-python-client.py that demonstrates idiomatic client usage:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-dummy")

# Access raw metrics via the underlying HTTP client

metrics = client.get("/metrics")
print(metrics.json())

This pattern is verified in tests/test_server_openai.py, which asserts correct payload structure and HTTP status codes.

Key Files Reference

File Purpose
mtplx/server/openai.py FastAPI server with /metrics and streaming route implementations; contains _metrics_envelope, _record_request_metrics, _merge_final_bridge_stats_into_latest_metrics
tests/test_server_openai.py Unit tests validating endpoint behavior and response schemas
docs/server.md High-level API documentation covering all server endpoints
examples/openai-python-client.py Reference Python client showing metrics access patterns

Summary

  • Two endpoints serve MTPLX metrics: /metrics for snapshots, /v1/mtplx/metrics/stream for real-time feeds
  • Core metrics include tok_s, accept_rate, and request pipeline state captured per completion
  • Implementation lives in mtplx/server/openai.py through _record_request_metrics and _metrics_envelope
  • Both curl and Python clients work; the streaming endpoint uses standard HTTP streaming for dashboard integration
  • Tests and examples in tests/test_server_openai.py and examples/openai-python-client.py provide verified reference code

Frequently Asked Questions

What port does the MTPLX metrics API use?

By default, the MTPLX server binds to port 8000. The metrics endpoints are accessible at http://localhost:8000/metrics and http://localhost:8000/v1/mtplx/metrics/stream. Check your server startup logs or configuration if running on a non-default port.

Do I need authentication to access MTPLX metrics?

The open-source implementation in mtplx/server/openai.py does not enforce authentication on the metrics endpoints by default. For production deployments, place the server behind a reverse proxy with appropriate access controls, as the metrics may expose internal performance characteristics.

Can I access historical metrics beyond the latest snapshot?

The /metrics endpoint includes a history array in its envelope, though population depends on the state.last_metrics ring buffer configuration. For persistent historical analysis, poll the endpoint periodically and store results externally, or consume the streaming endpoint and archive the feed.

How accurate is the tok_s measurement?

The tok_s value is computed per-request in _record_request_metrics based on actual token generation duration. It reflects end-to-end throughput including speculative decoding overhead, making it suitable for capacity planning and performance regression detection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →