# MTPLX Integration with Other Tools: OpenAI-Compatible Server Guide

> Integrate MTPLX with tools like OpenCode and Android Studio using its OpenAI-compatible server. Leverage standard endpoints and warm-prefix caching for peak performance. Learn more now.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-11

---

**MTPLX exposes a local OpenAI-compatible HTTP server that enables seamless integration with external tools like OpenCode, Pi, Hermes, and Android Studio through standard endpoints such as `/v1/chat/completions`, while preserving cross-client performance via warm-prefix session caching.**

MTPLX integration with other tools relies on its implementation as a drop-in replacement for OpenAI's API server. By hosting standard HTTP endpoints defined in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), the repository allows AI-enabled applications to connect without code changes, automatically discovering model capabilities like vision support and tool calling through the `/v1/models` endpoint.

## Architecture of the MTPLX OpenAI-Compatible Server

MTPLX functions as a local **OpenAI-compatible** server that external tools consume through three distinct integration layers. The implementation parses the OpenAI JSON schema, routes chat and completion requests, and injects MTPLX-specific headers including `x-session-id` and `x-session-affinity` for distributed tracing.

### Transport Layer and Endpoints

The **Transport** layer exposes standard HTTP endpoints including `/v1/chat/completions`, `/v1/completions`, `/v1/messages`, and `/v1/models`. You initialize this layer by running `mtplx serve`, which binds to `localhost:8000` by default and accepts requests following the OpenAI API contract.

### Authentication Mechanisms

For non-localhost network bindings, MTPLX provides optional **API-key protection** through the `--api-key-file` flag or the `MTPLX_API_KEY` environment variable. The server generates an API key automatically on first run when configured for external network access, as detailed in [`docs/server.md`](https://github.com/youssofal/MTPLX/blob/main/docs/server.md).

### Dynamic Capability Discovery

The server populates the `/v1/models` response with metadata that advertises model capabilities such as `supports_vision` and `modalities.input`. Clients query this endpoint to detect features like vision towers—derived from `vision_config` or `model.visual` entries—allowing dynamic adaptation of request payloads without manual configuration.

## Practical Integration Examples

MTPLX supports diverse client ecosystems ranging from AI coding assistants to IDE plugins. Each client connects to the same daemon instance, avoiding duplicate model loads and preserving the **warm-prefix cache** across different tools.

### Connecting OpenCode, Pi, and Hermes

Clients such as **OpenCode**, **Pi**, and **Hermes** enrich their provider definitions using capability data from `/v1/models`. When the loaded model includes vision metadata, these clients automatically enable image input support and advertise tool-type endpoints through the `tools` field in chat completions.

### Open WebUI Docker Configuration

For **Open WebUI**, MTPLX provides a helper command that generates the appropriate Docker configuration. Running `mtplx openwebui docker-command` disables Open WebUI’s internal Ollama probe and points the UI at the local MTPLX base URL (`http://127.0.0.1:8000/v1`), with full instructions available in [`examples/openwebui.md`](https://github.com/youssofal/MTPLX/blob/main/examples/openwebui.md).

### Android Studio External Model Provider

**Android Studio** connects to MTPLX through its External Model Provider settings. Configure the IDE with the following parameters to enable code-completion requests:

```text
URL: http://127.0.0.1:8008/v1
URL schema: OpenAI-compatible
API key: (leave blank for localhost)

```

This configuration is documented in [`docs/server.md`](https://github.com/youssofal/MTPLX/blob/main/docs/server.md) under the Android Studio section.

### Python requests Client Implementation

The following implementation from [`examples/python-requests-client.py`](https://github.com/youssofal/MTPLX/blob/main/examples/python-requests-client.py) demonstrates a minimal chat completion client:

```python
import requests

response = requests.post(
    "http://127.0.0.1:8000/v1/chat/completions",
    json={
        "model": "mtplx",
        "messages": [
            {"role": "user", "content": "Return a compact JSON object with one greeting key."}
        ],
        "max_tokens": 128,
    },
    timeout=120,
)
response.raise_for_status()
payload = response.json()
print(payload["choices"][0]["message"]["content"])

```

### cURL Command-Line Testing

For rapid testing or shell scripts, use the bash example from [`examples/curl-chat-completions.sh`](https://github.com/youssofal/MTPLX/blob/main/examples/curl-chat-completions.sh):

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mtplx","messages":[{"role":"user","content":"Hello MTPLX"}],"stream":true}'

```

## Performance and Advanced Features

MTPLX maintains high performance across multiple simultaneous clients through architectural optimizations that extend beyond simple request routing.

### Session Bank Warm-Prefix Caching

The **session bank** stores KV-cache snapshots on SSD at `~/.mtplx/…`, enabling **warm-prefix** reuse across different client connections. Any client respecting the `model` field receives cached context, dramatically reducing turn-time for long-context workloads regardless of which tool initiated the session.

### Embedding and Reranking Sidecars

Additional models can be registered using the `--embedding-model` and `--reranker-model` flags. These sidecars expose `/v1/embeddings` and `/v1/rerank` endpoints but bypass the MTP token-generation path, returning vectors directly for cost-effective inference while maintaining the same server interface.

## Summary

- MTPLX provides an **OpenAI-compatible** HTTP server through [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), supporting standard endpoints like `/v1/chat/completions` and `/v1/models`.
- **MTPLX integration with other tools** requires no client-side code changes; applications such as **OpenCode**, **Pi**, and **Android Studio** connect using standard OpenAI API configurations.
- **Capability discovery** via `/v1/models` automatically advertises vision support and tool calling based on model metadata examination.
- Authentication is optional for localhost but configurable via `--api-key-file` or the `MTPLX_API_KEY` environment variable for network exposure.
- The **session bank** preserves warm-prefix caches across all connected clients, ensuring consistent performance regardless of tool diversity.
- **Embedding and reranking** models run as sidecars with dedicated endpoints, expanding utility without impacting chat completion latency.

## Frequently Asked Questions

### Which tools are compatible with MTPLX integration?

MTPLX supports any client implementing the OpenAI or Anthropic API contract, including **OpenCode**, **Pi**, **Hermes**, **Open WebUI**, **Anthropic clients**, and **Android Studio**. The server advertises capabilities through the `/v1/models` endpoint, allowing clients to automatically detect vision support and tool-calling features.

### How does MTPLX handle authentication for external tools?

For localhost connections, no authentication is required. When binding to external network interfaces, MTPLX generates an API key on first run and validates requests against the `--api-key-file` flag or the `MTPLX_API_KEY` environment variable. This security layer is documented in [`docs/server.md`](https://github.com/youssofal/MTPLX/blob/main/docs/server.md) under the network sharing section.

### Can multiple tools use the same MTPLX server simultaneously?

Yes. Launching `mtplx start` creates a single daemon instance that all clients share. This architecture prevents duplicate model loads in memory and preserves the **warm-prefix cache** stored in `~/.mtplx/…` across different tools, ensuring that context from one client remains available to others.

### What endpoints does MTPLX expose for integration?

The server exposes `/v1/chat/completions` for chat, `/v1/completions` for text completion, `/v1/messages` for Anthropic-style requests, `/v1/models` for capability discovery, and optional `/v1/embeddings` and `/v1/rerank` when sidecar models are loaded via `--embedding-model` and `--reranker-model`. The router in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) handles schema parsing and request routing for all endpoints.