# How to Configure Embeddings and Reranking Endpoints in MTPLX

> Configure embeddings and reranking endpoints in MTPLX using CLI flags. Easily set up model registration and expose OpenAI-compatible API routes for your application.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-13

---

**You configure embeddings and reranking endpoints in MTPLX by passing the `--embedding-model` and `--reranker-model` CLI flags to the `mtplx serve` command, which registers models in a retrieval registry and exposes OpenAI-compatible `/v1/embeddings` and `/v1/rerank` routes.**

MTPLX is an open-source multi-token parallel language modeling framework that provides OpenAI-compatible API endpoints for retrieval tasks. To configure embeddings and reranking endpoints in MTPLX, you must explicitly enable them at server startup using specific command-line flags that populate the retrieval registry. This guide walks through the exact CLI arguments, configuration file options, and source code architecture required to activate these endpoints.

## Enabling the Endpoints with CLI Flags

Both the `/v1/embeddings` and `/v1/rerank` endpoints are **disabled by default**. You activate them by passing model references to the `mtplx serve` command.

- **`--embedding-model <reference>`**: Registers a model for the embeddings endpoint using the syntax `<repo>/<model>` or `<repo>/<model>=<custom_id>` to set a custom served ID.
- **`--reranker-model <reference>`**: Registers a model for the rerank endpoint; the model must implement a Jina-style [`rerank.py`](https://github.com/youssofal/MTPLX/blob/main/rerank.py) or equivalent retrieval loader.

In [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py), these flags are defined and forwarded via [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) to the server subprocess.

```bash
mtplx serve \
    --model Qwen3-Chat-7B \
    --embedding-model org/jina-embed \
    --reranker-model org/jina-rerank

```

## Understanding the Retrieval Registry Architecture

When the server starts, it builds a **retrieval registry** defined in [`mtplx/retrieval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/retrieval.py) from the supplied references. The registry loads each model **lazily** on the first request and caches loaded models according to retrieval configuration options like `--retrieval-max-resident` and `--retrieval-idle-timeout`. This architecture minimizes memory usage by only keeping frequently accessed models resident in GPU memory.

## Step-by-Step Configuration Examples

### Command-Line Configuration

For Jina embedding checkpoints, you must include the security flag to allow remote code execution:

```bash
mtplx serve \
    --model Qwen3-Chat-7B \
    --embedding-model org/jina-embed \
    --reranker-model org/jina-rerank \
    --retrieval-trust-remote-code

```

After the server initializes, test the endpoints:

```bash
curl http://localhost:8080/v1/embeddings \
    -H "Content-Type: application/json" \
    -d '{ "input": ["Hello world"] }'

```

```bash
curl http://localhost:8080/v1/rerank \
    -H "Content-Type: application/json" \
    -d '{
        "query": "What is MTPLX?",
        "documents": [
            "MTPLX is a framework for multi-token-parallel-language modeling.",
            "This is unrelated text."
        ]
    }'

```

### Configuration File Setup

Persist settings in `~/.mtplx/config.toml` to avoid repeating CLI arguments:

```toml
embedding_model = ["org/jina-embed"]
reranker_model = ["org/jina-rerank"]
retrieval_trust_remote_code = true
retrieval_max_resident = 2

```

Then run the server with minimal arguments:

```bash
mtplx serve --model Qwen3-Chat-7B

```

The CLI merges configuration file values with any command-line overrides.

### Python Client Implementation

Use any OpenAI-compatible client to interact with the configured endpoints:

```python
import openai

# Embeddings

resp = openai.Embedding.create(
    model="embed-a",  # Served ID from the registry

    input=["MTPLX provides retrieval APIs"]
)
vectors = resp["data"][0]["embedding"]

# Reranking

resp = openai.Rerank.create(
    model="rerank-a",  # Served ID from the registry

    query="What does MTPLX do?",
    documents=[
        "MTPLX is a retrieval-enabled LLM runtime.",
        "Random unrelated sentence."
    ]
)
scores = [r["score"] for r in resp["results"]]

```

## Security Requirements for Jina Checkpoints

If your model originates from a **Jina embedding checkpoint**, you must enable `--retrieval-trust-remote-code` (or set `retrieval_trust_remote_code = true` in your configuration file). This flag, processed in [`mtplx/retrieval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/retrieval.py), allows execution of repository-specific loading code. Without it, the endpoint returns a **400 error** indicating that remote code execution is disabled.

## Core Implementation Files

The configuration flow spans these critical source files:

- **[`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py)**: Defines top-level CLI options including `--embedding-model`, `--reranker-model`, and retrieval-specific flags.
- **[`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py)**: Forwards embedding and reranker flags to the server subprocess.
- **[`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)**: Implements the FastAPI routes for `/v1/embeddings` and `/v1/rerank`, looking up models in the retrieval registry.
- **[`mtplx/retrieval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/retrieval.py)**: Manages the registry, validates checkpoint types, and handles lazy loading/unloading of retrieval models.

## Summary

- Both endpoints are **disabled by default** and require explicit CLI flags to enable.
- Use `--embedding-model` and `--reranker-model` with `<repo>/<model>` syntax to configure embeddings and reranking endpoints in MTPLX.
- **Jina models require** `--retrieval-trust-remote-code` for security compliance.
- Models are **lazy-loaded** through the retrieval registry defined in [`mtplx/retrieval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/retrieval.py).
- Settings can be persisted in `~/.mtplx/config.toml` and merged with CLI overrides.

## Frequently Asked Questions

### What command enables the embedding endpoint in MTPLX?

Pass `--embedding-model <reference>` to the `mtplx serve` command. The reference uses the format `<repo>/<model>` or `<repo>/<model>=<custom_id>` to register a model that serves the `/v1/embeddings` route according to the logic in [`mtplx/retrieval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/retrieval.py).

### Can I use both embeddings and reranking models simultaneously?

Yes. You can pass both `--embedding-model` and `--reranker-model` flags in the same command. The server initializes a retrieval registry that manages both endpoints independently, allowing concurrent usage of the `/v1/embeddings` and `/v1/rerank` routes through [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).

### Why am I getting a 400 error when using Jina embedding models?

Jina checkpoints require explicit permission to execute remote code. Add the `--retrieval-trust-remote-code` flag to your startup command or set `retrieval_trust_remote_code = true` in your configuration file. Without this flag, [`mtplx/retrieval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/retrieval.py) blocks the model load and returns a 400 error.

### How does MTPLX handle model loading for retrieval tasks?

The retrieval registry in [`mtplx/retrieval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/retrieval.py) loads models lazily on their first request rather than at startup. Loaded models remain cached in memory according to `--retrieval-max-resident` and `--retrieval-idle-timeout` settings, optimizing resource usage for variable traffic patterns.