MTPLX Integration with Other Tools: OpenAI-Compatible Server Guide
MTPLX exposes a local OpenAI-compatible HTTP server that enables seamless integration with external tools like OpenCode, Pi, Hermes, and Android Studio through standard endpoints such as /v1/chat/completions, while preserving cross-client performance via warm-prefix session caching.
MTPLX integration with other tools relies on its implementation as a drop-in replacement for OpenAI's API server. By hosting standard HTTP endpoints defined in mtplx/server/openai.py, the repository allows AI-enabled applications to connect without code changes, automatically discovering model capabilities like vision support and tool calling through the /v1/models endpoint.
Architecture of the MTPLX OpenAI-Compatible Server
MTPLX functions as a local OpenAI-compatible server that external tools consume through three distinct integration layers. The implementation parses the OpenAI JSON schema, routes chat and completion requests, and injects MTPLX-specific headers including x-session-id and x-session-affinity for distributed tracing.
Transport Layer and Endpoints
The Transport layer exposes standard HTTP endpoints including /v1/chat/completions, /v1/completions, /v1/messages, and /v1/models. You initialize this layer by running mtplx serve, which binds to localhost:8000 by default and accepts requests following the OpenAI API contract.
Authentication Mechanisms
For non-localhost network bindings, MTPLX provides optional API-key protection through the --api-key-file flag or the MTPLX_API_KEY environment variable. The server generates an API key automatically on first run when configured for external network access, as detailed in docs/server.md.
Dynamic Capability Discovery
The server populates the /v1/models response with metadata that advertises model capabilities such as supports_vision and modalities.input. Clients query this endpoint to detect features like vision towers—derived from vision_config or model.visual entries—allowing dynamic adaptation of request payloads without manual configuration.
Practical Integration Examples
MTPLX supports diverse client ecosystems ranging from AI coding assistants to IDE plugins. Each client connects to the same daemon instance, avoiding duplicate model loads and preserving the warm-prefix cache across different tools.
Connecting OpenCode, Pi, and Hermes
Clients such as OpenCode, Pi, and Hermes enrich their provider definitions using capability data from /v1/models. When the loaded model includes vision metadata, these clients automatically enable image input support and advertise tool-type endpoints through the tools field in chat completions.
Open WebUI Docker Configuration
For Open WebUI, MTPLX provides a helper command that generates the appropriate Docker configuration. Running mtplx openwebui docker-command disables Open WebUI’s internal Ollama probe and points the UI at the local MTPLX base URL (http://127.0.0.1:8000/v1), with full instructions available in examples/openwebui.md.
Android Studio External Model Provider
Android Studio connects to MTPLX through its External Model Provider settings. Configure the IDE with the following parameters to enable code-completion requests:
URL: http://127.0.0.1:8008/v1
URL schema: OpenAI-compatible
API key: (leave blank for localhost)
This configuration is documented in docs/server.md under the Android Studio section.
Python requests Client Implementation
The following implementation from examples/python-requests-client.py demonstrates a minimal chat completion client:
import requests
response = requests.post(
"http://127.0.0.1:8000/v1/chat/completions",
json={
"model": "mtplx",
"messages": [
{"role": "user", "content": "Return a compact JSON object with one greeting key."}
],
"max_tokens": 128,
},
timeout=120,
)
response.raise_for_status()
payload = response.json()
print(payload["choices"][0]["message"]["content"])
cURL Command-Line Testing
For rapid testing or shell scripts, use the bash example from examples/curl-chat-completions.sh:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"mtplx","messages":[{"role":"user","content":"Hello MTPLX"}],"stream":true}'
Performance and Advanced Features
MTPLX maintains high performance across multiple simultaneous clients through architectural optimizations that extend beyond simple request routing.
Session Bank Warm-Prefix Caching
The session bank stores KV-cache snapshots on SSD at ~/.mtplx/…, enabling warm-prefix reuse across different client connections. Any client respecting the model field receives cached context, dramatically reducing turn-time for long-context workloads regardless of which tool initiated the session.
Embedding and Reranking Sidecars
Additional models can be registered using the --embedding-model and --reranker-model flags. These sidecars expose /v1/embeddings and /v1/rerank endpoints but bypass the MTP token-generation path, returning vectors directly for cost-effective inference while maintaining the same server interface.
Summary
- MTPLX provides an OpenAI-compatible HTTP server through
mtplx/server/openai.py, supporting standard endpoints like/v1/chat/completionsand/v1/models. - MTPLX integration with other tools requires no client-side code changes; applications such as OpenCode, Pi, and Android Studio connect using standard OpenAI API configurations.
- Capability discovery via
/v1/modelsautomatically advertises vision support and tool calling based on model metadata examination. - Authentication is optional for localhost but configurable via
--api-key-fileor theMTPLX_API_KEYenvironment variable for network exposure. - The session bank preserves warm-prefix caches across all connected clients, ensuring consistent performance regardless of tool diversity.
- Embedding and reranking models run as sidecars with dedicated endpoints, expanding utility without impacting chat completion latency.
Frequently Asked Questions
Which tools are compatible with MTPLX integration?
MTPLX supports any client implementing the OpenAI or Anthropic API contract, including OpenCode, Pi, Hermes, Open WebUI, Anthropic clients, and Android Studio. The server advertises capabilities through the /v1/models endpoint, allowing clients to automatically detect vision support and tool-calling features.
How does MTPLX handle authentication for external tools?
For localhost connections, no authentication is required. When binding to external network interfaces, MTPLX generates an API key on first run and validates requests against the --api-key-file flag or the MTPLX_API_KEY environment variable. This security layer is documented in docs/server.md under the network sharing section.
Can multiple tools use the same MTPLX server simultaneously?
Yes. Launching mtplx start creates a single daemon instance that all clients share. This architecture prevents duplicate model loads in memory and preserves the warm-prefix cache stored in ~/.mtplx/… across different tools, ensuring that context from one client remains available to others.
What endpoints does MTPLX expose for integration?
The server exposes /v1/chat/completions for chat, /v1/completions for text completion, /v1/messages for Anthropic-style requests, /v1/models for capability discovery, and optional /v1/embeddings and /v1/rerank when sidecar models are loaded via --embedding-model and --reranker-model. The router in mtplx/server/openai.py handles schema parsing and request routing for all endpoints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →