# MTPLX API Reference: OpenAI-Compatible Server for Local LLMs on Apple Silicon

> Access the MTPLX API reference for an OpenAI-compatible server running local LLMs on Apple Silicon with MLX. Interact with models using standard clients via FastAPI.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: api-reference
- Published: 2026-09-11

---

**MTPLX exposes a lightweight FastAPI-based HTTP server that mirrors the OpenAI and Anthropic API specifications, enabling any standard client to interact with local large language models running on Apple Silicon via MLX.**

The MTPLX API provides drop-in compatibility with existing OpenAI SDKs and tools, eliminating the need for external GPUs or Docker containers. The core implementation resides in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), where a FastAPI application registers routes for chat completions, embeddings, tool calls, and system observability.

## Core API Endpoints

The server registers eight primary endpoints that follow OpenAI’s REST conventions. All routes are defined as asynchronous functions in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) and delegate to MTPLX’s internal runtime components.

### Chat Completions

The `POST /v1/chat/completions` endpoint handles conversational AI requests with support for streaming, tool calling, and optional retrieval-augmented generation. According to the source code at lines 56-69 of [`openai.py`](https://github.com/youssofal/MTPLX/blob/main/openai.py), this route accepts a JSON payload modeled by the `ChatCompletionRequest` Pydantic class. The endpoint supports the `stream` parameter for Server-Sent Events (SSE) and integrates with the `MTPBatchGenerationService` to execute multi-token prediction algorithms.

### Classic Completions

The `POST /v1/completions` endpoint provides legacy single-prompt token generation. This route bypasses thechat message formatting and returns raw token streams, suitable for completion-style tasks.

### Embeddings and Reranking

Two specialized endpoints handle retrieval tasks without invoking the MTP generation pipeline:

- **`POST /v1/embeddings`**: Returns vector embeddings for input text using models like `Qwen3-Embedding-8B-4bit-DWQ`. The server loads embedding models lazily on first request.
- **`POST /v1/rerank`**: Reranks documents against a query using models such as `Qwen3-Reranker-4B-4bit-MLX`. This endpoint shares the same underlying MLX infrastructure but applies a different forward pass.

Both endpoints respect the `--retrieval-max-resident` CLI flag to limit memory usage by capping the number of resident retrieval models.

### Anthropic-Compatible Messages

The `POST /v1/messages` endpoint implements Anthropic’s API specification, supporting tool calls and streaming responses. This allows clients written for Claude to interact with MTPLX without modification.

### System Endpoints

- **`GET /v1/models`**: Lists available models including the default `mtplx` chat model and registered retrieval models. Each entry includes a `capability` field indicating supported operations.
- **`GET /health`**: Returns process status and optional MLX diagnostic information.
- **`GET /metrics`**: Exports Prometheus-style metrics for monitoring token throughput and request latency.

## Architecture and Implementation Details

### FastAPI Application Structure

The server initializes via `app = FastAPI(...)` in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py). All routes use standard FastAPI async function signatures, enabling high-concurrency handling of multiple concurrent chat sessions. The application internally delegates requests to the `ModelWorkScheduler` and `MTPBatchGenerationService` for actual tensor operations.

### Request Validation with Pydantic

Incoming request bodies are strictly validated using Pydantic models defined in [`openai.py`](https://github.com/youssofal/MTPLX/blob/main/openai.py). The `ChatCompletionRequest`, `CompletionRequest`, and `EmbeddingRequest` classes enforce type safety for parameters like `temperature`, `top_p`, and `presence_penalty`. Validation errors automatically serialize into standard OpenAI error response formats.

### Model Selection and Routing

The `model` field in requests resolves against the default chat model (`mtplx`) or registered retrieval models. The resolution logic lives in [`mtplx/server_urls.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server_urls.py), specifically within the `bind_label` and `local_url_for_bind` functions. This utility module handles label-to-URL mapping and filters models by capability flags.

### Multi-Token Prediction Pipeline

For generation endpoints, requests flow into the `MTPBatchGenerationService`, which implements MTPLX’s core optimization: draft-then-verify multi-token prediction. This pipeline generates multiple candidate tokens speculatively, then validates them via exact rejection sampling. This mechanism provides the latency advantages characteristic of MTPLX compared to standard autoregressive generation.

### Tool Call Handling

When requests contain OpenAI-style function definitions, the payload routes through [`mtplx/server/omlx_bridge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/omlx_bridge.py). The bridge extracts tool calls using `OMLXToolCallStreamFilter`, executes user-provided Python functions, and re-injects results into the token stream. This occurs transparently during streaming responses.

### Observability and Metrics

Request-level telemetry is captured by the `FlightRecorder` class in [`mtplx/server/flight_recorder.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/flight_recorder.py). This component records timestamps, token counts, and error states. The `/metrics` endpoint aggregates these into Prometheus counters for integration with monitoring stacks like Grafana.

### Configuration and Sampling

Generation parameters default to values managed by [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py). Users can override server-wide defaults (such as `temperature` or frequency penalties) via the CLI (`mtplx settings set ...`) or per-request JSON payloads.

## Usage Examples

### Streaming Chat Completion with curl

Request a streaming response from the local server:

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model":"mtplx",
        "messages":[{"role":"user","content":"Explain multi-token prediction"}],
        "stream":true
      }'

```

The server returns Server-Sent Events containing incremental text chunks.

### Python OpenAI Client Integration

Use the official OpenAI library to connect to MTPLX:

```python
import openai

client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1")
resp = client.chat.completions.create(
    model="mtplx",
    messages=[{"role": "user", "content": "Write a tiny Python hello-world"}],
    stream=True,
)
for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="")

```

This pattern works with LangChain, Open WebUI, and other OpenAI-compatible frameworks.

### Generating Embeddings

Compute vector embeddings for retrieval tasks:

```bash
curl http://127.0.0.1:8000/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen3-Embedding-8B-4bit-DWQ","input":["hello","world"]}'

```

### Reranking Documents

Score document relevance against a query:

```bash
curl http://127.0.0.1:8000/v1/rerank \
  -H "Content-Type: application/json" \
  -d '{
        "model":"Qwen3-Reranker-4B-4bit-MLX",
        "query":"What is the cache location?",
        "documents":["the cache lives in ~/.mtplx","unrelated text"]
      }'

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) | FastAPI application definition, route handlers, and Pydantic request models |
| [`mtplx/server/omlx_bridge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/omlx_bridge.py) | Tool call extraction, execution, and stream filtering logic |
| [`mtplx/server_urls.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server_urls.py) | Model label resolution and URL binding utilities |
| [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py) | Sampler configuration for temperature, top-p, and repetition penalties |
| [`mtplx/server/flight_recorder.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/flight_recorder.py) | Request tracing and Prometheus metrics aggregation |

## Summary

- The MTPLX API provides **OpenAI-compatible HTTP endpoints** for chat completions, embeddings, and tool calls running locally on Apple Silicon.
- The **FastAPI server** in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) handles request validation via Pydantic and delegates to the **MTP batch generation service** for speculative decoding.
- **Tool calls** process through [`omlx_bridge.py`](https://github.com/youssofal/MTPLX/blob/main/omlx_bridge.py), enabling function execution without breaking the streaming interface.
- **Retrieval endpoints** (`/v1/embeddings`, `/v1/rerank`) support RAG workflows with lazy model loading and memory caps.
- Standard OpenAI clients connect by setting `base_url` to `http://127.0.0.1:8000/v1`, requiring no code changes for migration.

## Frequently Asked Questions

### What base URL should I use to connect OpenAI clients to MTPLX?

Set the `base_url` parameter to `http://127.0.0.1:8000/v1` (or your configured host and port) when initializing the OpenAI client. This redirects all API calls to the local MTPLX server instead of OpenAI’s cloud endpoints.

### Does MTPLX support streaming responses for chat completions?

Yes. The `/v1/chat/completions` endpoint supports Server-Sent Events when you set `"stream": true` in the request payload. The implementation in [`openai.py`](https://github.com/youssofal/MTPLX/blob/main/openai.py) handles SSE formatting and chunks tokens as they are generated by the MTP pipeline.

### Can I use MTPLX for retrieval-augmented generation (RAG)?

Yes. MTPLX exposes dedicated `/v1/embeddings` and `/v1/rerank` endpoints for vector search and document reranking. These endpoints use MLX-optimized models like Qwen3-Embedding and load lazily to conserve memory, making them suitable for RAG pipelines without external services.

### How does MTPLX handle function calling and tool use?

Tool calls route through [`mtplx/server/omlx_bridge.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/omlx_bridge.py), which extracts function definitions from the request, executes the corresponding Python functions, and injects results back into the response stream via `OMLXToolCallStreamFilter`. This process is transparent to the client and supports both OpenAI and Anthropic message formats.