Complete Guide to MTPLX Server API Endpoints: OpenAI-Compatible and Admin Routes
The MTPLX server exposes a FastAPI-based HTTP API with OpenAI-compatible endpoints for chat completions and embeddings, alongside proprietary administrative routes for cache management, thermal control, and benchmarking, all defined in mtplx/server/openai.py.
The MTPLX repository implements a high-performance inference server that mirrors the OpenAI API specification while adding specialized endpoints for hardware optimization and observability. This article catalogs every public MTPLX server API endpoint, tracing their implementation to specific line numbers in the main server file to help developers integrate and manage deployments effectively.
OpenAI-Compatible Core Endpoints
The MTPLX server maintains full compatibility with the OpenAI REST API specification, implementing the standard routes required by modern LLM clients and SDKs.
Chat and Completions
POST /v1/chat/completions– The primary streaming endpoint for multi-turn conversations. Supports tool calling, stop sequences, and Server-Sent Events (SSE) for token-by-token delivery. Implemented at line 30280 inmtplx/server/openai.py.POST /v1/completions– Legacy non-chat completion endpoint for single-prompt generation. Located at line 34915.POST /v1/messages– Creates a message object following the OpenAI messages API pattern. Found at line 34845.POST /v1/messages/count_tokens– Returns token counts for a message payload without performing generation. Defined at line 34877.
Models and Embeddings
GET /v1/models– Returns the list of available model cards, required by the OpenAI spec for client discovery. See line 30102.POST /v1/embeddings– Generates vector embeddings for input text with OpenAI-compatible request/response formats. Implemented at line 30158.POST /v1/rerank– Reranks documents against a query using the OpenAI-style reranking protocol. Located at line 30239.
MTPLX Operational and Diagnostic Endpoints
Beyond the OpenAI specification, MTPLX exposes proprietary routes for runtime configuration, hardware monitoring, and request lifecycle management.
Runtime Configuration and State
GET /v1/mtplx/settingsandGET /mtplx/settings– Retrieve current server parameters including depth, fan mode, and profiling status. Defined at lines 29184–29185.POST /v1/mtplx/settingsandPOST /mtplx/settings– Update mutable settings such as inference depth or fan mode toggles. See lines 29189–29190.GET /v1/mtplx/snapshot– Dumps the internal cache state for debugging complex inference scenarios. Located at line 29201.GET /v1/mtplx/prefill_history– Exposes recent pre-fill operation logs for observability pipelines. Found at line 29205.
Request Management and Monitoring
- `POST /v1/mtplx/cancel/{request_id}`` – Aborts an active generation by its unique request identifier. Implemented at line 29219.
GET /v1/mtplx/flight– Opens an SSE stream delivering real-time flight-recorder logs for progress visualization in WebUI clients. See line 29212.GET /v1/mtplx/app/capabilities– Advertises optional server features such as supported scheduler modes. Located at line 29239.
Hardware Thermal Control
POST /v1/mtplx/thermal/fan_modeandPOST /mtplx/thermal/fan_mode– Adjusts GPU fan modes betweensmart,max, anddefault. Defined at line 29243.GET /v1/mtplx/thermal/statusandGET /mtplx/thermal/status– Retrieves current fan speeds and temperature metrics. See lines 29316–29317.
Administrative and Cache Management Endpoints
The /admin namespace provides low-level control over sessions and the SSD cache, useful for debugging and performance tuning.
Session Control
GET /admin/sessions– Lists all active engine sessions with metadata. Located at line 30042.POST /admin/sessions/{session_id}/clear– Forces a specific session’s cache eviction. Found at line 30046.
SSD Cache Operations
POST /admin/cache/clear– Flushes the entire SSD cache immediately. See line 30050.GET /admin/cache/ssd– Queries SSD-cache utilization statistics. Defined at line 30089.POST /admin/cache/ssd/archive– Creates a checkpoint archive of the current SSD cache. Located at line 30096.
Health, Metrics, and Benchmarking
MTPLX includes dedicated endpoints for orchestration health-checks, Prometheus scraping, and AIME benchmark orchestration.
System Health and Observability
GET /– Serves the minimal HTML dashboard used by Open WebUI and other front-ends. See line 28811.GET /v1andGET /v1/– Lightweight health checks returning a JSON meta-object. Located at lines 28843–28844.GET /health– Bare-bones liveness probe for Kubernetes and systemd orchestration. Found at line 28868.GET /metrics– Prometheus-compatible metrics export for external monitoring stacks. See line 30032.GET /v1/mtplx/metrics/stream– SSE stream broadcasting real-time metrics including tokens-per-second and latency. Defined at line 29977.
AIME Benchmark Suite
The experimental AIME benchmark endpoints provide lifecycle management for performance testing runs:
POST /v1/mtplx/benchmarks/aime/start– Initiates a new benchmark run. See line 29775.GET /v1/mtplx/benchmarks/aime/active– Queries the currently active benchmark, if any. Located at line 29832.GET /v1/mtplx/benchmarks/aime/history– Lists completed benchmark runs with summary statistics. Found at line 29838.- `GET /v1/mtplx/benchmarks/aime/{run_id}`` – Retrieves detailed metadata for a specific run. See line 29879.
POST /v1/mtplx/benchmarks/aime/{run_id}/pause– Pauses an active benchmark. Defined at line 29888.POST /v1/mtplx/benchmarks/aime/{run_id}/resume– Resumes a paused benchmark. Located at line 29898.POST /v1/mtplx/benchmarks/aime/{run_id}/skip– Skips the current benchmark step. Found at line 29908.POST /v1/mtplx/benchmarks/aime/{run_id}/cancel– Terminates an entire benchmark run. See line 29918.GET /v1/mtplx/benchmarks/aime/{run_id}/stream– SSE stream delivering live benchmark metrics during execution. Defined at line 29928.
Client Integration Examples
The following Python snippets demonstrate interaction with key MTPLX server API endpoints using standard HTTP requests:
import requests
BASE = "http://localhost:8080/v1"
# Retrieve available models
resp = requests.get(f"{BASE}/models")
print(resp.json())
# Streaming chat completion
payload = {
"model": "meta-llama/Meta-Llama-3-8B-Instruct",
"messages": [{"role": "user", "content": "Explain quantum entanglement"}],
"stream": True,
}
with requests.post(f"{BASE}/chat/completions", json=payload, stream=True) as r:
for line in r.iter_lines():
if line:
print(line.decode())
# Update server settings dynamically
settings = {"depth": 4}
resp = requests.post(f"{BASE}/mtplx/settings", json=settings)
print(resp.json())
# Debug cache state
resp = requests.get(f"{BASE}/mtplx/snapshot")
print(resp.json())
# Cancel a running request
request_id = "abc123"
resp = requests.post(f"{BASE}/mtplx/cancel/{request_id}")
print(resp.status_code)
Summary
- OpenAI Compatibility: The MTPLX server implements standard
/v1/models,/v1/chat/completions,/v1/embeddings, and/v1/completionsendpoints required for drop-in SDK replacement. - Operational Control: Proprietary
/v1/mtplx/*routes expose runtime settings, thermal management, and request cancellation capabilities unique to the MTPLX inference engine. - Administrative Access: The
/admin/*namespace provides cache clearing, session management, and SSD archiving for maintenance and debugging workflows. - Observability: SSE streams at
/v1/mtplx/flightand/v1/mtplx/metrics/streamenable real-time monitoring, while/metricsserves Prometheus scrapes. - Benchmarking: The AIME benchmark suite offers full lifecycle control via
/v1/mtplx/benchmarks/aime/*endpoints with live metric streaming.
Frequently Asked Questions
What file contains all the MTPLX API endpoint definitions?
All route handlers are registered in mtplx/server/openai.py, where the FastAPI application instance is created in create_app() at line 28641 and routes are attached through line 35540.
How does MTPLX handle OpenAI API compatibility?
According to the source code in mtplx/server/openai.py, MTPLX mirrors the OpenAI request/response schemas exactly for chat completions, embeddings, and completions while adding custom headers and MTPLX-specific metadata endpoints under the /v1/mtplx prefix.
Can I monitor GPU temperature and fan status via the API?
Yes. The GET /v1/mtplx/thermal/status endpoint returns current GPU temperatures and fan speeds, while POST /v1/mtplx/thermal/fan_mode allows dynamic switching between smart, max, and default cooling profiles.
Is there a way to stream real-time server metrics?
Yes. The GET /v1/mtplx/metrics/stream endpoint provides a Server-Sent Events (SSE) stream that pushes live tokens-per-second and latency metrics, suitable for building custom dashboards without polling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →