MTPLX API Reference: OpenAI-Compatible Server for Local LLMs on Apple Silicon
MTPLX exposes a lightweight FastAPI-based HTTP server that mirrors the OpenAI and Anthropic API specifications, enabling any standard client to interact with local large language models running on Apple Silicon via MLX.
The MTPLX API provides drop-in compatibility with existing OpenAI SDKs and tools, eliminating the need for external GPUs or Docker containers. The core implementation resides in mtplx/server/openai.py, where a FastAPI application registers routes for chat completions, embeddings, tool calls, and system observability.
Core API Endpoints
The server registers eight primary endpoints that follow OpenAI’s REST conventions. All routes are defined as asynchronous functions in mtplx/server/openai.py and delegate to MTPLX’s internal runtime components.
Chat Completions
The POST /v1/chat/completions endpoint handles conversational AI requests with support for streaming, tool calling, and optional retrieval-augmented generation. According to the source code at lines 56-69 of openai.py, this route accepts a JSON payload modeled by the ChatCompletionRequest Pydantic class. The endpoint supports the stream parameter for Server-Sent Events (SSE) and integrates with the MTPBatchGenerationService to execute multi-token prediction algorithms.
Classic Completions
The POST /v1/completions endpoint provides legacy single-prompt token generation. This route bypasses thechat message formatting and returns raw token streams, suitable for completion-style tasks.
Embeddings and Reranking
Two specialized endpoints handle retrieval tasks without invoking the MTP generation pipeline:
POST /v1/embeddings: Returns vector embeddings for input text using models likeQwen3-Embedding-8B-4bit-DWQ. The server loads embedding models lazily on first request.POST /v1/rerank: Reranks documents against a query using models such asQwen3-Reranker-4B-4bit-MLX. This endpoint shares the same underlying MLX infrastructure but applies a different forward pass.
Both endpoints respect the --retrieval-max-resident CLI flag to limit memory usage by capping the number of resident retrieval models.
Anthropic-Compatible Messages
The POST /v1/messages endpoint implements Anthropic’s API specification, supporting tool calls and streaming responses. This allows clients written for Claude to interact with MTPLX without modification.
System Endpoints
GET /v1/models: Lists available models including the defaultmtplxchat model and registered retrieval models. Each entry includes acapabilityfield indicating supported operations.GET /health: Returns process status and optional MLX diagnostic information.GET /metrics: Exports Prometheus-style metrics for monitoring token throughput and request latency.
Architecture and Implementation Details
FastAPI Application Structure
The server initializes via app = FastAPI(...) in mtplx/server/openai.py. All routes use standard FastAPI async function signatures, enabling high-concurrency handling of multiple concurrent chat sessions. The application internally delegates requests to the ModelWorkScheduler and MTPBatchGenerationService for actual tensor operations.
Request Validation with Pydantic
Incoming request bodies are strictly validated using Pydantic models defined in openai.py. The ChatCompletionRequest, CompletionRequest, and EmbeddingRequest classes enforce type safety for parameters like temperature, top_p, and presence_penalty. Validation errors automatically serialize into standard OpenAI error response formats.
Model Selection and Routing
The model field in requests resolves against the default chat model (mtplx) or registered retrieval models. The resolution logic lives in mtplx/server_urls.py, specifically within the bind_label and local_url_for_bind functions. This utility module handles label-to-URL mapping and filters models by capability flags.
Multi-Token Prediction Pipeline
For generation endpoints, requests flow into the MTPBatchGenerationService, which implements MTPLX’s core optimization: draft-then-verify multi-token prediction. This pipeline generates multiple candidate tokens speculatively, then validates them via exact rejection sampling. This mechanism provides the latency advantages characteristic of MTPLX compared to standard autoregressive generation.
Tool Call Handling
When requests contain OpenAI-style function definitions, the payload routes through mtplx/server/omlx_bridge.py. The bridge extracts tool calls using OMLXToolCallStreamFilter, executes user-provided Python functions, and re-injects results into the token stream. This occurs transparently during streaming responses.
Observability and Metrics
Request-level telemetry is captured by the FlightRecorder class in mtplx/server/flight_recorder.py. This component records timestamps, token counts, and error states. The /metrics endpoint aggregates these into Prometheus counters for integration with monitoring stacks like Grafana.
Configuration and Sampling
Generation parameters default to values managed by mtplx/sampling.py. Users can override server-wide defaults (such as temperature or frequency penalties) via the CLI (mtplx settings set ...) or per-request JSON payloads.
Usage Examples
Streaming Chat Completion with curl
Request a streaming response from the local server:
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model":"mtplx",
"messages":[{"role":"user","content":"Explain multi-token prediction"}],
"stream":true
}'
The server returns Server-Sent Events containing incremental text chunks.
Python OpenAI Client Integration
Use the official OpenAI library to connect to MTPLX:
import openai
client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1")
resp = client.chat.completions.create(
model="mtplx",
messages=[{"role": "user", "content": "Write a tiny Python hello-world"}],
stream=True,
)
for chunk in resp:
print(chunk.choices[0].delta.content or "", end="")
This pattern works with LangChain, Open WebUI, and other OpenAI-compatible frameworks.
Generating Embeddings
Compute vector embeddings for retrieval tasks:
curl http://127.0.0.1:8000/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model":"Qwen3-Embedding-8B-4bit-DWQ","input":["hello","world"]}'
Reranking Documents
Score document relevance against a query:
curl http://127.0.0.1:8000/v1/rerank \
-H "Content-Type: application/json" \
-d '{
"model":"Qwen3-Reranker-4B-4bit-MLX",
"query":"What is the cache location?",
"documents":["the cache lives in ~/.mtplx","unrelated text"]
}'
Key Source Files
| File | Purpose |
|---|---|
mtplx/server/openai.py |
FastAPI application definition, route handlers, and Pydantic request models |
mtplx/server/omlx_bridge.py |
Tool call extraction, execution, and stream filtering logic |
mtplx/server_urls.py |
Model label resolution and URL binding utilities |
mtplx/sampling.py |
Sampler configuration for temperature, top-p, and repetition penalties |
mtplx/server/flight_recorder.py |
Request tracing and Prometheus metrics aggregation |
Summary
- The MTPLX API provides OpenAI-compatible HTTP endpoints for chat completions, embeddings, and tool calls running locally on Apple Silicon.
- The FastAPI server in
mtplx/server/openai.pyhandles request validation via Pydantic and delegates to the MTP batch generation service for speculative decoding. - Tool calls process through
omlx_bridge.py, enabling function execution without breaking the streaming interface. - Retrieval endpoints (
/v1/embeddings,/v1/rerank) support RAG workflows with lazy model loading and memory caps. - Standard OpenAI clients connect by setting
base_urltohttp://127.0.0.1:8000/v1, requiring no code changes for migration.
Frequently Asked Questions
What base URL should I use to connect OpenAI clients to MTPLX?
Set the base_url parameter to http://127.0.0.1:8000/v1 (or your configured host and port) when initializing the OpenAI client. This redirects all API calls to the local MTPLX server instead of OpenAI’s cloud endpoints.
Does MTPLX support streaming responses for chat completions?
Yes. The /v1/chat/completions endpoint supports Server-Sent Events when you set "stream": true in the request payload. The implementation in openai.py handles SSE formatting and chunks tokens as they are generated by the MTP pipeline.
Can I use MTPLX for retrieval-augmented generation (RAG)?
Yes. MTPLX exposes dedicated /v1/embeddings and /v1/rerank endpoints for vector search and document reranking. These endpoints use MLX-optimized models like Qwen3-Embedding and load lazily to conserve memory, making them suitable for RAG pipelines without external services.
How does MTPLX handle function calling and tool use?
Tool calls route through mtplx/server/omlx_bridge.py, which extracts function definitions from the request, executes the corresponding Python functions, and injects results back into the response stream via OMLXToolCallStreamFilter. This process is transparent to the client and supports both OpenAI and Anthropic message formats.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →