MTPLX API Reference: OpenAI-Compatible Server for Local LLMs on Apple Silicon

MTPLX exposes a lightweight FastAPI-based HTTP server that mirrors the OpenAI and Anthropic API specifications, enabling any standard client to interact with local large language models running on Apple Silicon via MLX.

The MTPLX API provides drop-in compatibility with existing OpenAI SDKs and tools, eliminating the need for external GPUs or Docker containers. The core implementation resides in mtplx/server/openai.py, where a FastAPI application registers routes for chat completions, embeddings, tool calls, and system observability.

Core API Endpoints

The server registers eight primary endpoints that follow OpenAI’s REST conventions. All routes are defined as asynchronous functions in mtplx/server/openai.py and delegate to MTPLX’s internal runtime components.

Chat Completions

The POST /v1/chat/completions endpoint handles conversational AI requests with support for streaming, tool calling, and optional retrieval-augmented generation. According to the source code at lines 56-69 of openai.py, this route accepts a JSON payload modeled by the ChatCompletionRequest Pydantic class. The endpoint supports the stream parameter for Server-Sent Events (SSE) and integrates with the MTPBatchGenerationService to execute multi-token prediction algorithms.

Classic Completions

The POST /v1/completions endpoint provides legacy single-prompt token generation. This route bypasses thechat message formatting and returns raw token streams, suitable for completion-style tasks.

Embeddings and Reranking

Two specialized endpoints handle retrieval tasks without invoking the MTP generation pipeline:

  • POST /v1/embeddings: Returns vector embeddings for input text using models like Qwen3-Embedding-8B-4bit-DWQ. The server loads embedding models lazily on first request.
  • POST /v1/rerank: Reranks documents against a query using models such as Qwen3-Reranker-4B-4bit-MLX. This endpoint shares the same underlying MLX infrastructure but applies a different forward pass.

Both endpoints respect the --retrieval-max-resident CLI flag to limit memory usage by capping the number of resident retrieval models.

Anthropic-Compatible Messages

The POST /v1/messages endpoint implements Anthropic’s API specification, supporting tool calls and streaming responses. This allows clients written for Claude to interact with MTPLX without modification.

System Endpoints

  • GET /v1/models: Lists available models including the default mtplx chat model and registered retrieval models. Each entry includes a capability field indicating supported operations.
  • GET /health: Returns process status and optional MLX diagnostic information.
  • GET /metrics: Exports Prometheus-style metrics for monitoring token throughput and request latency.

Architecture and Implementation Details

FastAPI Application Structure

The server initializes via app = FastAPI(...) in mtplx/server/openai.py. All routes use standard FastAPI async function signatures, enabling high-concurrency handling of multiple concurrent chat sessions. The application internally delegates requests to the ModelWorkScheduler and MTPBatchGenerationService for actual tensor operations.

Request Validation with Pydantic

Incoming request bodies are strictly validated using Pydantic models defined in openai.py. The ChatCompletionRequest, CompletionRequest, and EmbeddingRequest classes enforce type safety for parameters like temperature, top_p, and presence_penalty. Validation errors automatically serialize into standard OpenAI error response formats.

Model Selection and Routing

The model field in requests resolves against the default chat model (mtplx) or registered retrieval models. The resolution logic lives in mtplx/server_urls.py, specifically within the bind_label and local_url_for_bind functions. This utility module handles label-to-URL mapping and filters models by capability flags.

Multi-Token Prediction Pipeline

For generation endpoints, requests flow into the MTPBatchGenerationService, which implements MTPLX’s core optimization: draft-then-verify multi-token prediction. This pipeline generates multiple candidate tokens speculatively, then validates them via exact rejection sampling. This mechanism provides the latency advantages characteristic of MTPLX compared to standard autoregressive generation.

Tool Call Handling

When requests contain OpenAI-style function definitions, the payload routes through mtplx/server/omlx_bridge.py. The bridge extracts tool calls using OMLXToolCallStreamFilter, executes user-provided Python functions, and re-injects results into the token stream. This occurs transparently during streaming responses.

Observability and Metrics

Request-level telemetry is captured by the FlightRecorder class in mtplx/server/flight_recorder.py. This component records timestamps, token counts, and error states. The /metrics endpoint aggregates these into Prometheus counters for integration with monitoring stacks like Grafana.

Configuration and Sampling

Generation parameters default to values managed by mtplx/sampling.py. Users can override server-wide defaults (such as temperature or frequency penalties) via the CLI (mtplx settings set ...) or per-request JSON payloads.

Usage Examples

Streaming Chat Completion with curl

Request a streaming response from the local server:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model":"mtplx",
        "messages":[{"role":"user","content":"Explain multi-token prediction"}],
        "stream":true
      }'

The server returns Server-Sent Events containing incremental text chunks.

Python OpenAI Client Integration

Use the official OpenAI library to connect to MTPLX:

import openai

client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1")
resp = client.chat.completions.create(
    model="mtplx",
    messages=[{"role": "user", "content": "Write a tiny Python hello-world"}],
    stream=True,
)
for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="")

This pattern works with LangChain, Open WebUI, and other OpenAI-compatible frameworks.

Generating Embeddings

Compute vector embeddings for retrieval tasks:

curl http://127.0.0.1:8000/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen3-Embedding-8B-4bit-DWQ","input":["hello","world"]}'

Reranking Documents

Score document relevance against a query:

curl http://127.0.0.1:8000/v1/rerank \
  -H "Content-Type: application/json" \
  -d '{
        "model":"Qwen3-Reranker-4B-4bit-MLX",
        "query":"What is the cache location?",
        "documents":["the cache lives in ~/.mtplx","unrelated text"]
      }'

Key Source Files

File Purpose
mtplx/server/openai.py FastAPI application definition, route handlers, and Pydantic request models
mtplx/server/omlx_bridge.py Tool call extraction, execution, and stream filtering logic
mtplx/server_urls.py Model label resolution and URL binding utilities
mtplx/sampling.py Sampler configuration for temperature, top-p, and repetition penalties
mtplx/server/flight_recorder.py Request tracing and Prometheus metrics aggregation

Summary

  • The MTPLX API provides OpenAI-compatible HTTP endpoints for chat completions, embeddings, and tool calls running locally on Apple Silicon.
  • The FastAPI server in mtplx/server/openai.py handles request validation via Pydantic and delegates to the MTP batch generation service for speculative decoding.
  • Tool calls process through omlx_bridge.py, enabling function execution without breaking the streaming interface.
  • Retrieval endpoints (/v1/embeddings, /v1/rerank) support RAG workflows with lazy model loading and memory caps.
  • Standard OpenAI clients connect by setting base_url to http://127.0.0.1:8000/v1, requiring no code changes for migration.

Frequently Asked Questions

What base URL should I use to connect OpenAI clients to MTPLX?

Set the base_url parameter to http://127.0.0.1:8000/v1 (or your configured host and port) when initializing the OpenAI client. This redirects all API calls to the local MTPLX server instead of OpenAI’s cloud endpoints.

Does MTPLX support streaming responses for chat completions?

Yes. The /v1/chat/completions endpoint supports Server-Sent Events when you set "stream": true in the request payload. The implementation in openai.py handles SSE formatting and chunks tokens as they are generated by the MTP pipeline.

Can I use MTPLX for retrieval-augmented generation (RAG)?

Yes. MTPLX exposes dedicated /v1/embeddings and /v1/rerank endpoints for vector search and document reranking. These endpoints use MLX-optimized models like Qwen3-Embedding and load lazily to conserve memory, making them suitable for RAG pipelines without external services.

How does MTPLX handle function calling and tool use?

Tool calls route through mtplx/server/omlx_bridge.py, which extracts function definitions from the request, executes the corresponding Python functions, and injects results back into the response stream via OMLXToolCallStreamFilter. This process is transparent to the client and supports both OpenAI and Anthropic message formats.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →