How to Start the MTPLX Daemon and Expose an OpenAI-Compatible API

Start the MTPLX daemon with mtplx serve to host models and expose OpenAI-compatible endpoints at /v1/chat/completions and /v1/completions.

MTPLX is an open-source inference engine that runs models locally and exposes them through a REST API matching the OpenAI specification. This guide shows you how to launch the MTPLX daemon, configure network access, and connect clients using standard OpenAI SDK patterns. All commands and configuration paths reference the actual implementation in youssofal/MTPLX.

Quick Start: Launch the Daemon in 3 Steps

1. Install MTPLX

Choose your preferred installation method:


# Homebrew (macOS/Linux)

brew install youssofal/mtplx/mtplx

# Or PyPI

pip install mtplx

Full installation details are documented in [docs/quickstart.md](https://github.com/youssofal/MTPLX/blob/main/docs/quickstart.md).

2. Pull and Verify Your Model

Before starting the daemon, ensure your model is downloaded and functional:


# Pull a model from the registry

mtplx pull Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed

# Verify environment and model loading

mtplx doctor --summary

3. Start the MTPLX Daemon

Run mtplx serve with appropriate flags for your use case.

Local development (localhost only):

mtplx serve --host 127.0.0.1 --port 8000 --no-stats-footer

Output shows the local base URL:


http://127.0.0.1:8000/v1

Network-wide access (other devices can connect):

mtplx serve \
    --host 0.0.0.0 \
    --port 8000 \
    --api-key-file ~/.mtplx/api-key \
    --no-stats-footer

Output includes:

  • Network OpenAI API Base URL (e.g., http://192.168.1.20:8000/v1)
  • Generated API key (saved to ~/.mtplx/api-key)

The daemon creates the API key file automatically if it doesn't exist per the logic described in [docs/server.md](https://github.com/youssofal/MTPLX/blob/main/docs/server.md).

Configuring the MTPLX Daemon

Host and Port Binding

Flag Purpose Recommended Use
--host 127.0.0.1 Localhost only Single-machine development, maximum security
--host 0.0.0.0 All interfaces Multi-device networks, containerized deployments

Authentication

The MTPLX daemon uses Bearer token authentication when an API key file is specified:

--api-key-file ~/.mtplx/api-key

If the file exists, the daemon reads the key. If absent, it generates one and writes it to the path. Never expose --host 0.0.0.0 without --api-key-file.

Consuming the OpenAI-Compatible API

Using cURL

Test the MTPLX daemon directly with HTTP requests:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Authorization: Bearer $MTPLX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed",
        "messages": [{"role":"user","content":"Explain quantum tunnelling"}],
        "temperature": 0.7
      }'

The endpoint structure mirrors OpenAI's: /v1/chat/completions for chat models and /v1/completions for legacy completion models.

Using Open WebUI

MTPLX provides a helper command for running Open WebUI against the daemon without Ollama detection conflicts:

mtplx openwebui docker-command

See [docs/openwebui.md](https://github.com/youssofal/MTPLX/blob/main/examples/openwebui.md) for the exact Docker command and container configuration.

Using the Anthropic SDK

The MTPLX daemon also supports Anthropic's messages API format. Point the SDK at your local base URL:

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.getenv("MTPLX_API_KEY"),
    base_url="http://127.0.0.1:8000"  # Note: no /v1 suffix required

)

resp = client.messages.create(
    model="Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed",
    max_tokens=256,
    messages=[{"role": "user", "content": "Write a haiku about clouds"}],
)
print(resp.content)

Daemon Architecture and Key Files

Understanding these source files helps troubleshoot and extend the MTPLX daemon:

  • docs/server.md — Complete HTTP endpoint reference, n-gram pre-warm options, and security guidance for network exposure.
  • apps/MTPLXApp/README.md — Daemon supervisor logic, settings persistence, and UI-to-daemon communication protocol.
  • docs/quickstart.md — Installation prerequisites, model registry commands, and first-run verification.

The daemon implementation separates concerns between the low-level launcher (mtplx start) and the high-level server (mtplx serve), with the latter handling OpenAI API route registration and request validation.

Production Deployment Considerations

For production MTPLX daemon deployments:

  1. Always use TLS termination — Run behind nginx, Caddy, or a cloud load balancer
  2. Restrict --host 0.0.0.0 — Combine with firewall rules limiting source IPs
  3. Rotate API keys — Regenerate ~/.mtplx/api-key periodically; the daemon reads the file on each request
  4. Monitor via health endpoints — The daemon exposes health and metrics routes documented in docs/server.md

Summary

  • Install MTPLX via Homebrew or PyPI, then mtplx pull your target model
  • Start the daemon with mtplx serve --host 127.0.0.1 for local work or --host 0.0.0.0 --api-key-file for network access
  • Connect any OpenAI-compatible client using the printed base URL and Bearer token authentication
  • Reference docs/server.md for endpoint details and apps/MTPLXApp/README.md for daemon internals

Frequently Asked Questions

What's the difference between mtplx start and mtplx serve?

mtplx start is the low-level launcher that initializes the model and runtime. mtplx serve wraps this with HTTP server setup, OpenAI API route registration, and graceful shutdown handling. For API exposure, always use mtplx serve.

Does the MTPLX daemon support streaming responses?

Yes. The OpenAI-compatible endpoints at /v1/chat/completions and /v1/completions support Server-Sent Events (SSE) streaming when stream: true is included in the request body, matching OpenAI's behavior. See docs/server.md for the full parameter schema.

Can I run multiple MTPLX daemons on different ports?

Yes. Specify distinct --port values and optional --api-key-file paths for each instance. Each daemon maintains independent model state and memory allocation, so ensure your hardware has sufficient VRAM for concurrent loads.

Why does network mode require an API key file?

The MTPLX daemon enforces authentication when binding to all interfaces (0.0.0.0) to prevent unauthorized access to your model and compute resources. The --api-key-file requirement is a deliberate safety guardrail; localhost-only mode (127.0.0.1) skips this check for developer convenience.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →