# MTPLX API Endpoints: Complete Reference for FastAPI Server Routes

> Explore MTPLX API endpoints and FastAPI server routes. Discover over 35 endpoints for health checks, OpenAI compatibility, thermal management, benchmarks, and admin controls. Get the full reference.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: api-reference
- Published: 2026-09-06

---

**MTPLX exposes 35+ HTTP endpoints through a FastAPI server, grouped into health checks, OpenAI-compatible chat/completions, thermal management, AIME benchmarks, and administrative controls—all defined in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).**

The MTPLX inference engine provides a comprehensive REST API that mirrors OpenAI's interface while adding hardware-aware extensions for thermal control, benchmarking, and real-time monitoring. This guide catalogs every available endpoint with practical examples you can run immediately against a local server.

## Health and Base Endpoints

MTPLX provides simple entry points for orchestration tools and API discovery.

- **`GET /health`** – Returns server status for load balancers and health probes. Implemented at line 28868 in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).

- **`GET /v1`** and **`GET /v1/`** – Root path returning minimal API information. These routes share handler logic at line 28843.

```bash
curl http://localhost:5000/health

```

## Server Settings Endpoints

Runtime configuration inspection and modification.

### Retrieve Current Settings

- **`GET /v1/mtplx/settings`** (canonical)
- **`GET /mtplx/settings`** (alias)

Returns JSON with active parameters: model identifier, batch size, temperature, and max tokens. Route defined at line 29184.

### Update Settings

- **`POST /v1/mtplx/settings`** (canonical)
- **`POST /mtplx/settings`** (alias)

Accepts partial JSON payloads to mutate server behavior without restart. Route defined at line 29189.

```bash
curl -X POST http://localhost:5000/v1/mtplx/settings \
     -H "Content-Type: application/json" \
     -d '{"temperature": 0.7, "max_tokens": 512}'

```

## Cache and Prefill Inspection

Endpoints for debugging model state.

- **`GET /v1/mtplx/snapshot`** – Retrieves the latest KV cache snapshot for analysis. Line 29201.

- **`GET /v1/mtplx/prefill_history`** – Returns historical prefill operations with timings. Line 29205.

## Request Lifecycle Management

### Live Request Tracing

- **`GET /v1/mtplx/flight`** – Streams real-time trace data for in-flight requests via Server-Sent Events. Line 29212.

### Request Cancellation

- **`POST /v1/mtplx/cancel/{request_id}`** – Terminates a running inference request by UUID. Line 29219.

```bash
curl -X POST http://localhost:5000/v1/mtplx/cancel/123e4567-e89b-12d3-a456-426614174000

```

## Application and Thermal Control

MTPLX uniquely exposes hardware thermal state for edge deployments.

### Capability Discovery

- **`GET /v1/mtplx/app/capabilities`** – Lists UI-available features including thermal control support. Line 29239.

### Fan Mode Control

- **`POST /v1/mtplx/thermal/fan_mode`** (canonical)
- **`POST /mtplx/thermal/fan_mode`** (alias)

Body accepts `{"mode": "auto"}` or `{"mode": "max"}`. Line 29243.

### Thermal Status Query

- **`GET /v1/mtplx/thermal/status`** (canonical)
- **`GET /mtplx/thermal/status`** (alias)

Returns temperature sensors and current fan RPM. Line 29316.

## AIME Benchmark Endpoints

Full lifecycle management for AIME benchmark runs.

| Endpoint | Method | Purpose |
|----------|--------|---------|
| `/v1/mtplx/benchmarks/aime/start` | `POST` | Launch new benchmark. Line 29775. |
| `/v1/mtplx/benchmarks/aime/active` | `GET` | Get running benchmark ID. Line 29832. |
| `/v1/mtplx/benchmarks/aime/history` | `GET` | List completed runs. Line 29838. |
| `/v1/mtplx/benchmarks/aime/{run_id}` | `GET` | Specific run details. Line 29879. |
| `/v1/mtplx/benchmarks/aime/{run_id}/pause` | `POST` | Pause execution. |
| `/v1/mtplx/benchmarks/aime/{run_id}/resume` | `POST` | Resume from pause. |
| `/v1/mtplx/benchmarks/aime/{run_id}/skip` | `POST` | Skip current step. |
| `/v1/mtplx/benchmarks/aime/{run_id}/cancel` | `POST` | Abort run. |
| `/v1/mtplx/benchmarks/aime/{run_id}/stream` | `GET` | SSE stream of logs. |

```bash
curl -X POST http://localhost:5000/v1/mtplx/benchmarks/aime/start \
     -H "Content-Type: application/json" \
     -d '{"model": "qwen2", "prompt_length": 2048}'

```

## Real-Time Metrics Streaming

- **`GET /v1/mtplx/metrics/stream`** – SSE endpoint pushing throughput, latency, and memory statistics every second.

```python
import sseclient, requests

response = requests.get(
    "http://localhost:5000/v1/mtplx/metrics/stream",
    stream=True,
)
for event in sseclient.SSEClient(response).events():
    print(event.data)  # JSON: {"tokens_per_sec": 142.3, "latency_ms": 23.5, ...}

```

## Administrative Endpoints

Cache and session management without authentication (local deployment assumed).

| Endpoint | Method | Function |
|----------|--------|----------|
| `/admin/sessions` | `GET` | List active inference sessions. |
| `/admin/sessions/{session_id}/clear` | `POST` | Drop session state. |
| `/admin/cache/clear` | `POST` | Purge all on-disk cache. |
| `/admin/cache/ssd` | `GET` | Inspect SSD cache statistics. |
| `/admin/cache/ssd/archive` | `POST` | Snapshot cache to archive. |

## OpenAI-Compatible API Endpoints

Drop-in replacements for OpenAI client libraries.

| Endpoint | Method | OpenAI Equivalent |
|----------|--------|-------------------|
| `GET /v1/models` | `GET` | List models |
| `POST /v1/embeddings` | `POST` | Create embeddings |
| `POST /v1/rerank` | `POST` | Rerank candidates |
| `POST /v1/chat/completions` | `POST` | Chat completions |
| `POST /v1/messages` | `POST` | UI message endpoint |
| `POST /v1/messages/count_tokens` | `POST` | Token counting |
| `POST /v1/completions` | `POST` | Legacy completions |

```python
import requests, json

resp = requests.post(
    "http://localhost:5000/v1/chat/completions",
    headers={"Content-Type": "application/json"},
    data=json.dumps({
        "model": "qwen2",
        "messages": [{"role": "user", "content": "Explain quantum tunneling"}],
        "max_tokens": 256,
    })
)
print(resp.json()["choices"][0]["message"]["content"])

```

## Dashboard Endpoint

- **`GET /dashboard`** – Serves the built-in React-based web UI for monitoring and control.

## Implementation Architecture

All routes attach to a single FastAPI application created by `create_app()` in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py). Key architectural decisions:

- **Dual path schemes** – `/v1/mtplx/*` for MTPLX-native features, `/v1/*` for OpenAI parity
- **SSE for streaming** – Metrics and benchmark logs use Server-Sent Events rather than WebSockets for simpler client integration
- **No authentication** – Designed for local execution on secured edge devices

Supporting files include [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) for CLI wrappers, [`mtplx/benchmarks/runners/aime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/runners/aime.py) for benchmark implementation, and [`mtplx/server/dashboard_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/dashboard_state.py) for SSE stream management.

## Summary

- **35+ endpoints** across 8 functional categories in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)
- **OpenAI-compatible core** at `/v1/chat/completions`, `/v1/embeddings`, `/v1/models`
- **Hardware extensions** for thermal control (`/v1/mtplx/thermal/*`) and AIME benchmarking
- **Streaming first** – SSE used throughout for metrics, logs, and live traces
- **Admin routes** at `/admin/*` for session and cache management

## Frequently Asked Questions

### How do I list all available MTPLX API endpoints programmatically?

Start the server and fetch the auto-generated OpenAPI schema at `http://localhost:5000/openapi.json`. FastAPI produces this automatically from the route definitions in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), including all 35+ endpoints with their request and response models.

### What's the difference between `/v1/mtplx/settings` and `/v1/messages`?

`/v1/mtplx/settings` (GET/POST) controls server configuration like temperature and batch size. `/v1/messages` (POST) is the chat inference endpoint used by the web UI—similar to `/v1/chat/completions` but with MTPLX-specific payload extensions for internal state tracking.

### Does MTPLX require authentication for API access?

No. According to the source code in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), no authentication middleware is applied. The server assumes local deployment on a physically secured device. For remote access, place MTPLX behind a reverse proxy with TLS and authentication.

### How can I cancel a long-running inference request?

Capture the `request_id` from the response headers or logs, then POST to `/v1/mtplx/cancel/{request_id}`. The cancellation handler at line 29219 signals the inference loop to abort at the next token boundary.