Complete Guide to MTPLX Server API Endpoints: OpenAI-Compatible and Admin Routes

The MTPLX server exposes a FastAPI-based HTTP API with OpenAI-compatible endpoints for chat completions and embeddings, alongside proprietary administrative routes for cache management, thermal control, and benchmarking, all defined in mtplx/server/openai.py.

The MTPLX repository implements a high-performance inference server that mirrors the OpenAI API specification while adding specialized endpoints for hardware optimization and observability. This article catalogs every public MTPLX server API endpoint, tracing their implementation to specific line numbers in the main server file to help developers integrate and manage deployments effectively.

OpenAI-Compatible Core Endpoints

The MTPLX server maintains full compatibility with the OpenAI REST API specification, implementing the standard routes required by modern LLM clients and SDKs.

Chat and Completions

  • POST /v1/chat/completions – The primary streaming endpoint for multi-turn conversations. Supports tool calling, stop sequences, and Server-Sent Events (SSE) for token-by-token delivery. Implemented at line 30280 in mtplx/server/openai.py.
  • POST /v1/completions – Legacy non-chat completion endpoint for single-prompt generation. Located at line 34915.
  • POST /v1/messages – Creates a message object following the OpenAI messages API pattern. Found at line 34845.
  • POST /v1/messages/count_tokens – Returns token counts for a message payload without performing generation. Defined at line 34877.

Models and Embeddings

  • GET /v1/models – Returns the list of available model cards, required by the OpenAI spec for client discovery. See line 30102.
  • POST /v1/embeddings – Generates vector embeddings for input text with OpenAI-compatible request/response formats. Implemented at line 30158.
  • POST /v1/rerank – Reranks documents against a query using the OpenAI-style reranking protocol. Located at line 30239.

MTPLX Operational and Diagnostic Endpoints

Beyond the OpenAI specification, MTPLX exposes proprietary routes for runtime configuration, hardware monitoring, and request lifecycle management.

Runtime Configuration and State

  • GET /v1/mtplx/settings and GET /mtplx/settings – Retrieve current server parameters including depth, fan mode, and profiling status. Defined at lines 29184–29185.
  • POST /v1/mtplx/settings and POST /mtplx/settings – Update mutable settings such as inference depth or fan mode toggles. See lines 29189–29190.
  • GET /v1/mtplx/snapshot – Dumps the internal cache state for debugging complex inference scenarios. Located at line 29201.
  • GET /v1/mtplx/prefill_history – Exposes recent pre-fill operation logs for observability pipelines. Found at line 29205.

Request Management and Monitoring

  • `POST /v1/mtplx/cancel/{request_id}`` – Aborts an active generation by its unique request identifier. Implemented at line 29219.
  • GET /v1/mtplx/flight – Opens an SSE stream delivering real-time flight-recorder logs for progress visualization in WebUI clients. See line 29212.
  • GET /v1/mtplx/app/capabilities – Advertises optional server features such as supported scheduler modes. Located at line 29239.

Hardware Thermal Control

  • POST /v1/mtplx/thermal/fan_mode and POST /mtplx/thermal/fan_mode – Adjusts GPU fan modes between smart, max, and default. Defined at line 29243.
  • GET /v1/mtplx/thermal/status and GET /mtplx/thermal/status – Retrieves current fan speeds and temperature metrics. See lines 29316–29317.

Administrative and Cache Management Endpoints

The /admin namespace provides low-level control over sessions and the SSD cache, useful for debugging and performance tuning.

Session Control

  • GET /admin/sessions – Lists all active engine sessions with metadata. Located at line 30042.
  • POST /admin/sessions/{session_id}/clear – Forces a specific session’s cache eviction. Found at line 30046.

SSD Cache Operations

  • POST /admin/cache/clear – Flushes the entire SSD cache immediately. See line 30050.
  • GET /admin/cache/ssd – Queries SSD-cache utilization statistics. Defined at line 30089.
  • POST /admin/cache/ssd/archive – Creates a checkpoint archive of the current SSD cache. Located at line 30096.

Health, Metrics, and Benchmarking

MTPLX includes dedicated endpoints for orchestration health-checks, Prometheus scraping, and AIME benchmark orchestration.

System Health and Observability

  • GET / – Serves the minimal HTML dashboard used by Open WebUI and other front-ends. See line 28811.
  • GET /v1 and GET /v1/ – Lightweight health checks returning a JSON meta-object. Located at lines 28843–28844.
  • GET /health – Bare-bones liveness probe for Kubernetes and systemd orchestration. Found at line 28868.
  • GET /metrics – Prometheus-compatible metrics export for external monitoring stacks. See line 30032.
  • GET /v1/mtplx/metrics/stream – SSE stream broadcasting real-time metrics including tokens-per-second and latency. Defined at line 29977.

AIME Benchmark Suite

The experimental AIME benchmark endpoints provide lifecycle management for performance testing runs:

  • POST /v1/mtplx/benchmarks/aime/start – Initiates a new benchmark run. See line 29775.
  • GET /v1/mtplx/benchmarks/aime/active – Queries the currently active benchmark, if any. Located at line 29832.
  • GET /v1/mtplx/benchmarks/aime/history – Lists completed benchmark runs with summary statistics. Found at line 29838.
  • `GET /v1/mtplx/benchmarks/aime/{run_id}`` – Retrieves detailed metadata for a specific run. See line 29879.
  • POST /v1/mtplx/benchmarks/aime/{run_id}/pause – Pauses an active benchmark. Defined at line 29888.
  • POST /v1/mtplx/benchmarks/aime/{run_id}/resume – Resumes a paused benchmark. Located at line 29898.
  • POST /v1/mtplx/benchmarks/aime/{run_id}/skip – Skips the current benchmark step. Found at line 29908.
  • POST /v1/mtplx/benchmarks/aime/{run_id}/cancel – Terminates an entire benchmark run. See line 29918.
  • GET /v1/mtplx/benchmarks/aime/{run_id}/stream – SSE stream delivering live benchmark metrics during execution. Defined at line 29928.

Client Integration Examples

The following Python snippets demonstrate interaction with key MTPLX server API endpoints using standard HTTP requests:

import requests

BASE = "http://localhost:8080/v1"

# Retrieve available models

resp = requests.get(f"{BASE}/models")
print(resp.json())

# Streaming chat completion

payload = {
    "model": "meta-llama/Meta-Llama-3-8B-Instruct",
    "messages": [{"role": "user", "content": "Explain quantum entanglement"}],
    "stream": True,
}
with requests.post(f"{BASE}/chat/completions", json=payload, stream=True) as r:
    for line in r.iter_lines():
        if line:
            print(line.decode())

# Update server settings dynamically

settings = {"depth": 4}
resp = requests.post(f"{BASE}/mtplx/settings", json=settings)
print(resp.json())

# Debug cache state

resp = requests.get(f"{BASE}/mtplx/snapshot")
print(resp.json())

# Cancel a running request

request_id = "abc123"
resp = requests.post(f"{BASE}/mtplx/cancel/{request_id}")
print(resp.status_code)

Summary

  • OpenAI Compatibility: The MTPLX server implements standard /v1/models, /v1/chat/completions, /v1/embeddings, and /v1/completions endpoints required for drop-in SDK replacement.
  • Operational Control: Proprietary /v1/mtplx/* routes expose runtime settings, thermal management, and request cancellation capabilities unique to the MTPLX inference engine.
  • Administrative Access: The /admin/* namespace provides cache clearing, session management, and SSD archiving for maintenance and debugging workflows.
  • Observability: SSE streams at /v1/mtplx/flight and /v1/mtplx/metrics/stream enable real-time monitoring, while /metrics serves Prometheus scrapes.
  • Benchmarking: The AIME benchmark suite offers full lifecycle control via /v1/mtplx/benchmarks/aime/* endpoints with live metric streaming.

Frequently Asked Questions

What file contains all the MTPLX API endpoint definitions?

All route handlers are registered in mtplx/server/openai.py, where the FastAPI application instance is created in create_app() at line 28641 and routes are attached through line 35540.

How does MTPLX handle OpenAI API compatibility?

According to the source code in mtplx/server/openai.py, MTPLX mirrors the OpenAI request/response schemas exactly for chat completions, embeddings, and completions while adding custom headers and MTPLX-specific metadata endpoints under the /v1/mtplx prefix.

Can I monitor GPU temperature and fan status via the API?

Yes. The GET /v1/mtplx/thermal/status endpoint returns current GPU temperatures and fan speeds, while POST /v1/mtplx/thermal/fan_mode allows dynamic switching between smart, max, and default cooling profiles.

Is there a way to stream real-time server metrics?

Yes. The GET /v1/mtplx/metrics/stream endpoint provides a Server-Sent Events (SSE) stream that pushes live tokens-per-second and latency metrics, suitable for building custom dashboards without polling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →