How to Start the MTPLX Daemon and Expose an OpenAI-Compatible API
Start the MTPLX daemon with mtplx serve to host models and expose OpenAI-compatible endpoints at /v1/chat/completions and /v1/completions.
MTPLX is an open-source inference engine that runs models locally and exposes them through a REST API matching the OpenAI specification. This guide shows you how to launch the MTPLX daemon, configure network access, and connect clients using standard OpenAI SDK patterns. All commands and configuration paths reference the actual implementation in youssofal/MTPLX.
Quick Start: Launch the Daemon in 3 Steps
1. Install MTPLX
Choose your preferred installation method:
# Homebrew (macOS/Linux)
brew install youssofal/mtplx/mtplx
# Or PyPI
pip install mtplx
Full installation details are documented in [docs/quickstart.md](https://github.com/youssofal/MTPLX/blob/main/docs/quickstart.md).
2. Pull and Verify Your Model
Before starting the daemon, ensure your model is downloaded and functional:
# Pull a model from the registry
mtplx pull Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed
# Verify environment and model loading
mtplx doctor --summary
3. Start the MTPLX Daemon
Run mtplx serve with appropriate flags for your use case.
Local development (localhost only):
mtplx serve --host 127.0.0.1 --port 8000 --no-stats-footer
Output shows the local base URL:
http://127.0.0.1:8000/v1
Network-wide access (other devices can connect):
mtplx serve \
--host 0.0.0.0 \
--port 8000 \
--api-key-file ~/.mtplx/api-key \
--no-stats-footer
Output includes:
- Network OpenAI API Base URL (e.g.,
http://192.168.1.20:8000/v1) - Generated API key (saved to
~/.mtplx/api-key)
The daemon creates the API key file automatically if it doesn't exist per the logic described in [docs/server.md](https://github.com/youssofal/MTPLX/blob/main/docs/server.md).
Configuring the MTPLX Daemon
Host and Port Binding
| Flag | Purpose | Recommended Use |
|---|---|---|
--host 127.0.0.1 |
Localhost only | Single-machine development, maximum security |
--host 0.0.0.0 |
All interfaces | Multi-device networks, containerized deployments |
Authentication
The MTPLX daemon uses Bearer token authentication when an API key file is specified:
--api-key-file ~/.mtplx/api-key
If the file exists, the daemon reads the key. If absent, it generates one and writes it to the path. Never expose --host 0.0.0.0 without --api-key-file.
Consuming the OpenAI-Compatible API
Using cURL
Test the MTPLX daemon directly with HTTP requests:
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Authorization: Bearer $MTPLX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed",
"messages": [{"role":"user","content":"Explain quantum tunnelling"}],
"temperature": 0.7
}'
The endpoint structure mirrors OpenAI's: /v1/chat/completions for chat models and /v1/completions for legacy completion models.
Using Open WebUI
MTPLX provides a helper command for running Open WebUI against the daemon without Ollama detection conflicts:
mtplx openwebui docker-command
See [docs/openwebui.md](https://github.com/youssofal/MTPLX/blob/main/examples/openwebui.md) for the exact Docker command and container configuration.
Using the Anthropic SDK
The MTPLX daemon also supports Anthropic's messages API format. Point the SDK at your local base URL:
import os
import anthropic
client = anthropic.Anthropic(
api_key=os.getenv("MTPLX_API_KEY"),
base_url="http://127.0.0.1:8000" # Note: no /v1 suffix required
)
resp = client.messages.create(
model="Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed",
max_tokens=256,
messages=[{"role": "user", "content": "Write a haiku about clouds"}],
)
print(resp.content)
Daemon Architecture and Key Files
Understanding these source files helps troubleshoot and extend the MTPLX daemon:
docs/server.md— Complete HTTP endpoint reference, n-gram pre-warm options, and security guidance for network exposure.apps/MTPLXApp/README.md— Daemon supervisor logic, settings persistence, and UI-to-daemon communication protocol.docs/quickstart.md— Installation prerequisites, model registry commands, and first-run verification.
The daemon implementation separates concerns between the low-level launcher (mtplx start) and the high-level server (mtplx serve), with the latter handling OpenAI API route registration and request validation.
Production Deployment Considerations
For production MTPLX daemon deployments:
- Always use TLS termination — Run behind nginx, Caddy, or a cloud load balancer
- Restrict
--host 0.0.0.0— Combine with firewall rules limiting source IPs - Rotate API keys — Regenerate
~/.mtplx/api-keyperiodically; the daemon reads the file on each request - Monitor via health endpoints — The daemon exposes health and metrics routes documented in
docs/server.md
Summary
- Install MTPLX via Homebrew or PyPI, then
mtplx pullyour target model - Start the daemon with
mtplx serve --host 127.0.0.1for local work or--host 0.0.0.0 --api-key-filefor network access - Connect any OpenAI-compatible client using the printed base URL and Bearer token authentication
- Reference
docs/server.mdfor endpoint details andapps/MTPLXApp/README.mdfor daemon internals
Frequently Asked Questions
What's the difference between mtplx start and mtplx serve?
mtplx start is the low-level launcher that initializes the model and runtime. mtplx serve wraps this with HTTP server setup, OpenAI API route registration, and graceful shutdown handling. For API exposure, always use mtplx serve.
Does the MTPLX daemon support streaming responses?
Yes. The OpenAI-compatible endpoints at /v1/chat/completions and /v1/completions support Server-Sent Events (SSE) streaming when stream: true is included in the request body, matching OpenAI's behavior. See docs/server.md for the full parameter schema.
Can I run multiple MTPLX daemons on different ports?
Yes. Specify distinct --port values and optional --api-key-file paths for each instance. Each daemon maintains independent model state and memory allocation, so ensure your hardware has sufficient VRAM for concurrent loads.
Why does network mode require an API key file?
The MTPLX daemon enforces authentication when binding to all interfaces (0.0.0.0) to prevent unauthorized access to your model and compute resources. The --api-key-file requirement is a deliberate safety guardrail; localhost-only mode (127.0.0.1) skips this check for developer convenience.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →