MTPLX API Endpoints: Complete Reference for FastAPI Server Routes
MTPLX exposes 35+ HTTP endpoints through a FastAPI server, grouped into health checks, OpenAI-compatible chat/completions, thermal management, AIME benchmarks, and administrative controls—all defined in mtplx/server/openai.py.
The MTPLX inference engine provides a comprehensive REST API that mirrors OpenAI's interface while adding hardware-aware extensions for thermal control, benchmarking, and real-time monitoring. This guide catalogs every available endpoint with practical examples you can run immediately against a local server.
Health and Base Endpoints
MTPLX provides simple entry points for orchestration tools and API discovery.
-
GET /health– Returns server status for load balancers and health probes. Implemented at line 28868 inmtplx/server/openai.py. -
GET /v1andGET /v1/– Root path returning minimal API information. These routes share handler logic at line 28843.
curl http://localhost:5000/health
Server Settings Endpoints
Runtime configuration inspection and modification.
Retrieve Current Settings
GET /v1/mtplx/settings(canonical)GET /mtplx/settings(alias)
Returns JSON with active parameters: model identifier, batch size, temperature, and max tokens. Route defined at line 29184.
Update Settings
POST /v1/mtplx/settings(canonical)POST /mtplx/settings(alias)
Accepts partial JSON payloads to mutate server behavior without restart. Route defined at line 29189.
curl -X POST http://localhost:5000/v1/mtplx/settings \
-H "Content-Type: application/json" \
-d '{"temperature": 0.7, "max_tokens": 512}'
Cache and Prefill Inspection
Endpoints for debugging model state.
-
GET /v1/mtplx/snapshot– Retrieves the latest KV cache snapshot for analysis. Line 29201. -
GET /v1/mtplx/prefill_history– Returns historical prefill operations with timings. Line 29205.
Request Lifecycle Management
Live Request Tracing
GET /v1/mtplx/flight– Streams real-time trace data for in-flight requests via Server-Sent Events. Line 29212.
Request Cancellation
POST /v1/mtplx/cancel/{request_id}– Terminates a running inference request by UUID. Line 29219.
curl -X POST http://localhost:5000/v1/mtplx/cancel/123e4567-e89b-12d3-a456-426614174000
Application and Thermal Control
MTPLX uniquely exposes hardware thermal state for edge deployments.
Capability Discovery
GET /v1/mtplx/app/capabilities– Lists UI-available features including thermal control support. Line 29239.
Fan Mode Control
POST /v1/mtplx/thermal/fan_mode(canonical)POST /mtplx/thermal/fan_mode(alias)
Body accepts {"mode": "auto"} or {"mode": "max"}. Line 29243.
Thermal Status Query
GET /v1/mtplx/thermal/status(canonical)GET /mtplx/thermal/status(alias)
Returns temperature sensors and current fan RPM. Line 29316.
AIME Benchmark Endpoints
Full lifecycle management for AIME benchmark runs.
| Endpoint | Method | Purpose |
|---|---|---|
/v1/mtplx/benchmarks/aime/start |
POST |
Launch new benchmark. Line 29775. |
/v1/mtplx/benchmarks/aime/active |
GET |
Get running benchmark ID. Line 29832. |
/v1/mtplx/benchmarks/aime/history |
GET |
List completed runs. Line 29838. |
/v1/mtplx/benchmarks/aime/{run_id} |
GET |
Specific run details. Line 29879. |
/v1/mtplx/benchmarks/aime/{run_id}/pause |
POST |
Pause execution. |
/v1/mtplx/benchmarks/aime/{run_id}/resume |
POST |
Resume from pause. |
/v1/mtplx/benchmarks/aime/{run_id}/skip |
POST |
Skip current step. |
/v1/mtplx/benchmarks/aime/{run_id}/cancel |
POST |
Abort run. |
/v1/mtplx/benchmarks/aime/{run_id}/stream |
GET |
SSE stream of logs. |
curl -X POST http://localhost:5000/v1/mtplx/benchmarks/aime/start \
-H "Content-Type: application/json" \
-d '{"model": "qwen2", "prompt_length": 2048}'
Real-Time Metrics Streaming
GET /v1/mtplx/metrics/stream– SSE endpoint pushing throughput, latency, and memory statistics every second.
import sseclient, requests
response = requests.get(
"http://localhost:5000/v1/mtplx/metrics/stream",
stream=True,
)
for event in sseclient.SSEClient(response).events():
print(event.data) # JSON: {"tokens_per_sec": 142.3, "latency_ms": 23.5, ...}
Administrative Endpoints
Cache and session management without authentication (local deployment assumed).
| Endpoint | Method | Function |
|---|---|---|
/admin/sessions |
GET |
List active inference sessions. |
/admin/sessions/{session_id}/clear |
POST |
Drop session state. |
/admin/cache/clear |
POST |
Purge all on-disk cache. |
/admin/cache/ssd |
GET |
Inspect SSD cache statistics. |
/admin/cache/ssd/archive |
POST |
Snapshot cache to archive. |
OpenAI-Compatible API Endpoints
Drop-in replacements for OpenAI client libraries.
| Endpoint | Method | OpenAI Equivalent |
|---|---|---|
GET /v1/models |
GET |
List models |
POST /v1/embeddings |
POST |
Create embeddings |
POST /v1/rerank |
POST |
Rerank candidates |
POST /v1/chat/completions |
POST |
Chat completions |
POST /v1/messages |
POST |
UI message endpoint |
POST /v1/messages/count_tokens |
POST |
Token counting |
POST /v1/completions |
POST |
Legacy completions |
import requests, json
resp = requests.post(
"http://localhost:5000/v1/chat/completions",
headers={"Content-Type": "application/json"},
data=json.dumps({
"model": "qwen2",
"messages": [{"role": "user", "content": "Explain quantum tunneling"}],
"max_tokens": 256,
})
)
print(resp.json()["choices"][0]["message"]["content"])
Dashboard Endpoint
GET /dashboard– Serves the built-in React-based web UI for monitoring and control.
Implementation Architecture
All routes attach to a single FastAPI application created by create_app() in mtplx/server/openai.py. Key architectural decisions:
- Dual path schemes –
/v1/mtplx/*for MTPLX-native features,/v1/*for OpenAI parity - SSE for streaming – Metrics and benchmark logs use Server-Sent Events rather than WebSockets for simpler client integration
- No authentication – Designed for local execution on secured edge devices
Supporting files include mtplx/commands/public.py for CLI wrappers, mtplx/benchmarks/runners/aime.py for benchmark implementation, and mtplx/server/dashboard_state.py for SSE stream management.
Summary
- 35+ endpoints across 8 functional categories in
mtplx/server/openai.py - OpenAI-compatible core at
/v1/chat/completions,/v1/embeddings,/v1/models - Hardware extensions for thermal control (
/v1/mtplx/thermal/*) and AIME benchmarking - Streaming first – SSE used throughout for metrics, logs, and live traces
- Admin routes at
/admin/*for session and cache management
Frequently Asked Questions
How do I list all available MTPLX API endpoints programmatically?
Start the server and fetch the auto-generated OpenAPI schema at http://localhost:5000/openapi.json. FastAPI produces this automatically from the route definitions in mtplx/server/openai.py, including all 35+ endpoints with their request and response models.
What's the difference between /v1/mtplx/settings and /v1/messages?
/v1/mtplx/settings (GET/POST) controls server configuration like temperature and batch size. /v1/messages (POST) is the chat inference endpoint used by the web UI—similar to /v1/chat/completions but with MTPLX-specific payload extensions for internal state tracking.
Does MTPLX require authentication for API access?
No. According to the source code in mtplx/server/openai.py, no authentication middleware is applied. The server assumes local deployment on a physically secured device. For remote access, place MTPLX behind a reverse proxy with TLS and authentication.
How can I cancel a long-running inference request?
Capture the request_id from the response headers or logs, then POST to /v1/mtplx/cancel/{request_id}. The cancellation handler at line 29219 signals the inference loop to abort at the next token boundary.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →