How to Configure Embeddings and Reranking Endpoints in MTPLX
You configure embeddings and reranking endpoints in MTPLX by passing the --embedding-model and --reranker-model CLI flags to the mtplx serve command, which registers models in a retrieval registry and exposes OpenAI-compatible /v1/embeddings and /v1/rerank routes.
MTPLX is an open-source multi-token parallel language modeling framework that provides OpenAI-compatible API endpoints for retrieval tasks. To configure embeddings and reranking endpoints in MTPLX, you must explicitly enable them at server startup using specific command-line flags that populate the retrieval registry. This guide walks through the exact CLI arguments, configuration file options, and source code architecture required to activate these endpoints.
Enabling the Endpoints with CLI Flags
Both the /v1/embeddings and /v1/rerank endpoints are disabled by default. You activate them by passing model references to the mtplx serve command.
--embedding-model <reference>: Registers a model for the embeddings endpoint using the syntax<repo>/<model>or<repo>/<model>=<custom_id>to set a custom served ID.--reranker-model <reference>: Registers a model for the rerank endpoint; the model must implement a Jina-stylererank.pyor equivalent retrieval loader.
In mtplx/cli.py, these flags are defined and forwarded via mtplx/commands/public.py to the server subprocess.
mtplx serve \
--model Qwen3-Chat-7B \
--embedding-model org/jina-embed \
--reranker-model org/jina-rerank
Understanding the Retrieval Registry Architecture
When the server starts, it builds a retrieval registry defined in mtplx/retrieval.py from the supplied references. The registry loads each model lazily on the first request and caches loaded models according to retrieval configuration options like --retrieval-max-resident and --retrieval-idle-timeout. This architecture minimizes memory usage by only keeping frequently accessed models resident in GPU memory.
Step-by-Step Configuration Examples
Command-Line Configuration
For Jina embedding checkpoints, you must include the security flag to allow remote code execution:
mtplx serve \
--model Qwen3-Chat-7B \
--embedding-model org/jina-embed \
--reranker-model org/jina-rerank \
--retrieval-trust-remote-code
After the server initializes, test the endpoints:
curl http://localhost:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{ "input": ["Hello world"] }'
curl http://localhost:8080/v1/rerank \
-H "Content-Type: application/json" \
-d '{
"query": "What is MTPLX?",
"documents": [
"MTPLX is a framework for multi-token-parallel-language modeling.",
"This is unrelated text."
]
}'
Configuration File Setup
Persist settings in ~/.mtplx/config.toml to avoid repeating CLI arguments:
embedding_model = ["org/jina-embed"]
reranker_model = ["org/jina-rerank"]
retrieval_trust_remote_code = true
retrieval_max_resident = 2
Then run the server with minimal arguments:
mtplx serve --model Qwen3-Chat-7B
The CLI merges configuration file values with any command-line overrides.
Python Client Implementation
Use any OpenAI-compatible client to interact with the configured endpoints:
import openai
# Embeddings
resp = openai.Embedding.create(
model="embed-a", # Served ID from the registry
input=["MTPLX provides retrieval APIs"]
)
vectors = resp["data"][0]["embedding"]
# Reranking
resp = openai.Rerank.create(
model="rerank-a", # Served ID from the registry
query="What does MTPLX do?",
documents=[
"MTPLX is a retrieval-enabled LLM runtime.",
"Random unrelated sentence."
]
)
scores = [r["score"] for r in resp["results"]]
Security Requirements for Jina Checkpoints
If your model originates from a Jina embedding checkpoint, you must enable --retrieval-trust-remote-code (or set retrieval_trust_remote_code = true in your configuration file). This flag, processed in mtplx/retrieval.py, allows execution of repository-specific loading code. Without it, the endpoint returns a 400 error indicating that remote code execution is disabled.
Core Implementation Files
The configuration flow spans these critical source files:
mtplx/cli.py: Defines top-level CLI options including--embedding-model,--reranker-model, and retrieval-specific flags.mtplx/commands/public.py: Forwards embedding and reranker flags to the server subprocess.mtplx/server/openai.py: Implements the FastAPI routes for/v1/embeddingsand/v1/rerank, looking up models in the retrieval registry.mtplx/retrieval.py: Manages the registry, validates checkpoint types, and handles lazy loading/unloading of retrieval models.
Summary
- Both endpoints are disabled by default and require explicit CLI flags to enable.
- Use
--embedding-modeland--reranker-modelwith<repo>/<model>syntax to configure embeddings and reranking endpoints in MTPLX. - Jina models require
--retrieval-trust-remote-codefor security compliance. - Models are lazy-loaded through the retrieval registry defined in
mtplx/retrieval.py. - Settings can be persisted in
~/.mtplx/config.tomland merged with CLI overrides.
Frequently Asked Questions
What command enables the embedding endpoint in MTPLX?
Pass --embedding-model <reference> to the mtplx serve command. The reference uses the format <repo>/<model> or <repo>/<model>=<custom_id> to register a model that serves the /v1/embeddings route according to the logic in mtplx/retrieval.py.
Can I use both embeddings and reranking models simultaneously?
Yes. You can pass both --embedding-model and --reranker-model flags in the same command. The server initializes a retrieval registry that manages both endpoints independently, allowing concurrent usage of the /v1/embeddings and /v1/rerank routes through mtplx/server/openai.py.
Why am I getting a 400 error when using Jina embedding models?
Jina checkpoints require explicit permission to execute remote code. Add the --retrieval-trust-remote-code flag to your startup command or set retrieval_trust_remote_code = true in your configuration file. Without this flag, mtplx/retrieval.py blocks the model load and returns a 400 error.
How does MTPLX handle model loading for retrieval tasks?
The retrieval registry in mtplx/retrieval.py loads models lazily on their first request rather than at startup. Loaded models remain cached in memory according to --retrieval-max-resident and --retrieval-idle-timeout settings, optimizing resource usage for variable traffic patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →