MTPLX API Endpoints: Complete Reference for FastAPI Server Routes

MTPLX exposes 35+ HTTP endpoints through a FastAPI server, grouped into health checks, OpenAI-compatible chat/completions, thermal management, AIME benchmarks, and administrative controls—all defined in mtplx/server/openai.py.

The MTPLX inference engine provides a comprehensive REST API that mirrors OpenAI's interface while adding hardware-aware extensions for thermal control, benchmarking, and real-time monitoring. This guide catalogs every available endpoint with practical examples you can run immediately against a local server.

Health and Base Endpoints

MTPLX provides simple entry points for orchestration tools and API discovery.

  • GET /health – Returns server status for load balancers and health probes. Implemented at line 28868 in mtplx/server/openai.py.

  • GET /v1 and GET /v1/ – Root path returning minimal API information. These routes share handler logic at line 28843.

curl http://localhost:5000/health

Server Settings Endpoints

Runtime configuration inspection and modification.

Retrieve Current Settings

  • GET /v1/mtplx/settings (canonical)
  • GET /mtplx/settings (alias)

Returns JSON with active parameters: model identifier, batch size, temperature, and max tokens. Route defined at line 29184.

Update Settings

  • POST /v1/mtplx/settings (canonical)
  • POST /mtplx/settings (alias)

Accepts partial JSON payloads to mutate server behavior without restart. Route defined at line 29189.

curl -X POST http://localhost:5000/v1/mtplx/settings \
     -H "Content-Type: application/json" \
     -d '{"temperature": 0.7, "max_tokens": 512}'

Cache and Prefill Inspection

Endpoints for debugging model state.

  • GET /v1/mtplx/snapshot – Retrieves the latest KV cache snapshot for analysis. Line 29201.

  • GET /v1/mtplx/prefill_history – Returns historical prefill operations with timings. Line 29205.

Request Lifecycle Management

Live Request Tracing

  • GET /v1/mtplx/flight – Streams real-time trace data for in-flight requests via Server-Sent Events. Line 29212.

Request Cancellation

  • POST /v1/mtplx/cancel/{request_id} – Terminates a running inference request by UUID. Line 29219.
curl -X POST http://localhost:5000/v1/mtplx/cancel/123e4567-e89b-12d3-a456-426614174000

Application and Thermal Control

MTPLX uniquely exposes hardware thermal state for edge deployments.

Capability Discovery

  • GET /v1/mtplx/app/capabilities – Lists UI-available features including thermal control support. Line 29239.

Fan Mode Control

  • POST /v1/mtplx/thermal/fan_mode (canonical)
  • POST /mtplx/thermal/fan_mode (alias)

Body accepts {"mode": "auto"} or {"mode": "max"}. Line 29243.

Thermal Status Query

  • GET /v1/mtplx/thermal/status (canonical)
  • GET /mtplx/thermal/status (alias)

Returns temperature sensors and current fan RPM. Line 29316.

AIME Benchmark Endpoints

Full lifecycle management for AIME benchmark runs.

Endpoint Method Purpose
/v1/mtplx/benchmarks/aime/start POST Launch new benchmark. Line 29775.
/v1/mtplx/benchmarks/aime/active GET Get running benchmark ID. Line 29832.
/v1/mtplx/benchmarks/aime/history GET List completed runs. Line 29838.
/v1/mtplx/benchmarks/aime/{run_id} GET Specific run details. Line 29879.
/v1/mtplx/benchmarks/aime/{run_id}/pause POST Pause execution.
/v1/mtplx/benchmarks/aime/{run_id}/resume POST Resume from pause.
/v1/mtplx/benchmarks/aime/{run_id}/skip POST Skip current step.
/v1/mtplx/benchmarks/aime/{run_id}/cancel POST Abort run.
/v1/mtplx/benchmarks/aime/{run_id}/stream GET SSE stream of logs.
curl -X POST http://localhost:5000/v1/mtplx/benchmarks/aime/start \
     -H "Content-Type: application/json" \
     -d '{"model": "qwen2", "prompt_length": 2048}'

Real-Time Metrics Streaming

  • GET /v1/mtplx/metrics/stream – SSE endpoint pushing throughput, latency, and memory statistics every second.
import sseclient, requests

response = requests.get(
    "http://localhost:5000/v1/mtplx/metrics/stream",
    stream=True,
)
for event in sseclient.SSEClient(response).events():
    print(event.data)  # JSON: {"tokens_per_sec": 142.3, "latency_ms": 23.5, ...}

Administrative Endpoints

Cache and session management without authentication (local deployment assumed).

Endpoint Method Function
/admin/sessions GET List active inference sessions.
/admin/sessions/{session_id}/clear POST Drop session state.
/admin/cache/clear POST Purge all on-disk cache.
/admin/cache/ssd GET Inspect SSD cache statistics.
/admin/cache/ssd/archive POST Snapshot cache to archive.

OpenAI-Compatible API Endpoints

Drop-in replacements for OpenAI client libraries.

Endpoint Method OpenAI Equivalent
GET /v1/models GET List models
POST /v1/embeddings POST Create embeddings
POST /v1/rerank POST Rerank candidates
POST /v1/chat/completions POST Chat completions
POST /v1/messages POST UI message endpoint
POST /v1/messages/count_tokens POST Token counting
POST /v1/completions POST Legacy completions
import requests, json

resp = requests.post(
    "http://localhost:5000/v1/chat/completions",
    headers={"Content-Type": "application/json"},
    data=json.dumps({
        "model": "qwen2",
        "messages": [{"role": "user", "content": "Explain quantum tunneling"}],
        "max_tokens": 256,
    })
)
print(resp.json()["choices"][0]["message"]["content"])

Dashboard Endpoint

  • GET /dashboard – Serves the built-in React-based web UI for monitoring and control.

Implementation Architecture

All routes attach to a single FastAPI application created by create_app() in mtplx/server/openai.py. Key architectural decisions:

  • Dual path schemes – /v1/mtplx/* for MTPLX-native features, /v1/* for OpenAI parity
  • SSE for streaming – Metrics and benchmark logs use Server-Sent Events rather than WebSockets for simpler client integration
  • No authentication – Designed for local execution on secured edge devices

Supporting files include mtplx/commands/public.py for CLI wrappers, mtplx/benchmarks/runners/aime.py for benchmark implementation, and mtplx/server/dashboard_state.py for SSE stream management.

Summary

  • 35+ endpoints across 8 functional categories in mtplx/server/openai.py
  • OpenAI-compatible core at /v1/chat/completions, /v1/embeddings, /v1/models
  • Hardware extensions for thermal control (/v1/mtplx/thermal/*) and AIME benchmarking
  • Streaming first – SSE used throughout for metrics, logs, and live traces
  • Admin routes at /admin/* for session and cache management

Frequently Asked Questions

How do I list all available MTPLX API endpoints programmatically?

Start the server and fetch the auto-generated OpenAPI schema at http://localhost:5000/openapi.json. FastAPI produces this automatically from the route definitions in mtplx/server/openai.py, including all 35+ endpoints with their request and response models.

What's the difference between /v1/mtplx/settings and /v1/messages?

/v1/mtplx/settings (GET/POST) controls server configuration like temperature and batch size. /v1/messages (POST) is the chat inference endpoint used by the web UI—similar to /v1/chat/completions but with MTPLX-specific payload extensions for internal state tracking.

Does MTPLX require authentication for API access?

No. According to the source code in mtplx/server/openai.py, no authentication middleware is applied. The server assumes local deployment on a physically secured device. For remote access, place MTPLX behind a reverse proxy with TLS and authentication.

How can I cancel a long-running inference request?

Capture the request_id from the response headers or logs, then POST to /v1/mtplx/cancel/{request_id}. The cancellation handler at line 29219 signals the inference loop to abort at the next token boundary.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →