How to Get Server Metrics from MTPLX: Complete API Guide with Code Examples
MTPLX exposes server metrics through two HTTP endpoints—/metrics for snapshots and /v1/mtplx/metrics/stream for real-time streaming—both returning JSON data about token throughput, acceptance rates, and request statistics.
The MTPLX inference server, implemented in the youssofal/MTPLX repository, provides built-in observability through a FastAPI-based metrics API. Whether you need a one-time health check or a live dashboard feed, the server exposes performance data without requiring external monitoring tools. This guide covers both endpoints, their underlying implementation, and practical code examples for accessing server metrics from MTPLX.
Available Metrics Endpoints
MTPLX registers two routes in mtplx/server/openai.py that serve different monitoring needs:
| Route | Method | Use Case |
|---|---|---|
/metrics |
GET | Retrieve the latest snapshot of server-wide metrics |
/v1/mtplx/metrics/stream |
GET | Stream metric updates in real time for live dashboards |
Both endpoints return JSON payloads constructed by the internal _metrics_envelope helper function.
What Metrics Are Collected
When processing requests, the server records granular performance data through _record_request_metrics and _merge_final_bridge_stats_into_latest_metrics. These populate state.last_metrics, a ring buffer that powers both endpoints.
Key fields in each metrics entry include:
tok_s— tokens generated per second for the requestaccept_rate— proportion of tokens accepted by the speculative decoding pipelinerequest_id— unique identifier for request correlationsession_prompt_prefix_commit— snapshot of prefix caching statesession_postcommit_snapshot— post-processing execution pipeline state
Fetching a Metrics Snapshot
The /metrics endpoint returns the most recent entry from the internal ring buffer, wrapped in a structured envelope.
Using curl
# Get latest metrics as formatted JSON
curl -s http://localhost:8000/metrics | jq .
Typical response structure:
{
"latest": {
"tok_s": 142.7,
"accept_rate": 0.89,
"request_id": "req_7a3f9d2e1b",
"timestamp": "2024-01-15T09:23:47.123456"
},
"history": []
}
Using Python requests
import requests
BASE_URL = "http://localhost:8000"
resp = requests.get(f"{BASE_URL}/metrics")
data = resp.json()
print(f"Tokens/sec: {data['latest']['tok_s']}")
print(f"Accept rate: {data['latest']['accept_rate']:.1%}")
Streaming Real-Time Metrics
For monitoring dashboards or continuous health tracking, the streaming endpoint yields each new metric entry as state.last_metrics is updated.
Using curl with streaming
# Stream metrics until interrupted (Ctrl-C)
curl -N http://localhost:8000/v1/mtplx/metrics/stream | jq .
Using Python with Server-Sent Events
import requests
import json
BASE_URL = "http://localhost:8000"
with requests.get(
f"{BASE_URL}/v1/mtplx/metrics/stream",
stream=True
) as r:
for line in r.iter_lines():
if line and line.startswith(b"data: "):
payload = json.loads(line[6:]) # Strip "data: " prefix
print(f"tok_s={payload['tok_s']:.1f}, "
f"accept_rate={payload['accept_rate']:.2f}")
Implementation Details in the Source Code
Understanding the internal architecture helps when building custom monitoring integrations.
Endpoint Registration
In mtplx/server/openai.py, the routes attach to the FastAPI app instance:
@app.get("/metrics")
@app.get("/v1/mtplx/metrics/stream")
The non-streaming handler calls _metrics_envelope(state) and returns its result directly. The streaming handler uses an async generator to yield newline-delimited JSON as new metrics arrive.
Metric Collection Pipeline
After each request completes, _record_request_metrics pushes a dictionary onto state.last_metrics. The later call to _merge_final_bridge_stats_into_latest_metrics enriches this with final execution statistics before the endpoint can serve it.
The Envelope Function
_metrics_envelope(state) constructs the response format you receive, placing the newest entry under the latest key and optionally including historical data. This abstraction allows the same payload structure for both snapshot and streaming consumers.
Using the Bundled Example Client
The repository includes a reference implementation in examples/openai-python-client.py that demonstrates idiomatic client usage:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-dummy")
# Access raw metrics via the underlying HTTP client
metrics = client.get("/metrics")
print(metrics.json())
This pattern is verified in tests/test_server_openai.py, which asserts correct payload structure and HTTP status codes.
Key Files Reference
| File | Purpose |
|---|---|
mtplx/server/openai.py |
FastAPI server with /metrics and streaming route implementations; contains _metrics_envelope, _record_request_metrics, _merge_final_bridge_stats_into_latest_metrics |
tests/test_server_openai.py |
Unit tests validating endpoint behavior and response schemas |
docs/server.md |
High-level API documentation covering all server endpoints |
examples/openai-python-client.py |
Reference Python client showing metrics access patterns |
Summary
- Two endpoints serve MTPLX metrics:
/metricsfor snapshots,/v1/mtplx/metrics/streamfor real-time feeds - Core metrics include
tok_s,accept_rate, and request pipeline state captured per completion - Implementation lives in
mtplx/server/openai.pythrough_record_request_metricsand_metrics_envelope - Both curl and Python clients work; the streaming endpoint uses standard HTTP streaming for dashboard integration
- Tests and examples in
tests/test_server_openai.pyandexamples/openai-python-client.pyprovide verified reference code
Frequently Asked Questions
What port does the MTPLX metrics API use?
By default, the MTPLX server binds to port 8000. The metrics endpoints are accessible at http://localhost:8000/metrics and http://localhost:8000/v1/mtplx/metrics/stream. Check your server startup logs or configuration if running on a non-default port.
Do I need authentication to access MTPLX metrics?
The open-source implementation in mtplx/server/openai.py does not enforce authentication on the metrics endpoints by default. For production deployments, place the server behind a reverse proxy with appropriate access controls, as the metrics may expose internal performance characteristics.
Can I access historical metrics beyond the latest snapshot?
The /metrics endpoint includes a history array in its envelope, though population depends on the state.last_metrics ring buffer configuration. For persistent historical analysis, poll the endpoint periodically and store results externally, or consume the streaming endpoint and archive the feed.
How accurate is the tok_s measurement?
The tok_s value is computed per-request in _record_request_metrics based on actual token generation duration. It reflects end-to-end throughput including speculative decoding overhead, making it suitable for capacity planning and performance regression detection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →