RAG and Agent-Memory Retrieval Endpoints in MTPLX: Complete API Reference
MTPLX exposes three production-ready retrieval endpoints under /v1/mtplx/retrieval/ for embeddings, reranking, and agent-memory access.
The MTPLX open-source framework provides a FastAPI-based server with dedicated endpoints for Retrieval-Augmented Generation (RAG) pipelines and stateful agent memory. These endpoints enable developers to build context-aware LLM applications with persistent session storage and sophisticated document retrieval. This guide covers each endpoint's purpose, request format, and implementation details drawn directly from the MTPLX source code.
RAG Endpoints: Embedding and Reranking
MTPLX implements the two classic stages of RAG through separate HTTP endpoints. Both are defined in mtplx/server/openai.py and registered during server initialization.
POST /v1/mtplx/retrieval/embedding
The embedding endpoint converts raw text into dense vectors for semantic search. This is the foundation of any RAG pipeline, enabling similarity-based document retrieval.
import requests
BASE_URL = "http://localhost:8000"
# Generate embeddings for RAG query
embed_resp = requests.post(
f"{BASE_URL}/v1/mtplx/retrieval/embedding",
json={"text": "What is the capital of France?"}
)
embeddings = embed_resp.json()["embeddings"]
The endpoint accepts a JSON payload with a text field and returns a vector representation suitable for cosine similarity search against a document index.
POST /v1/mtplx/retrieval/rerank
The reranking endpoint refines retrieval results using a cross-encoder model. After initial candidate selection via vector similarity, this endpoint reorders passages by true relevance to the query.
# Rerank candidate passages for improved RAG accuracy
rerank_resp = requests.post(
f"{BASE_URL}/v1/mtplx/retrieval/rerank",
json={
"query": "What is the capital of France?",
"candidates": [
"Paris is the capital of France.",
"Berlin is the capital of Germany."
]
}
)
ranked = rerank_resp.json()["candidates"]
The request body requires:
query: The original user questioncandidates: List of passage strings from initial retrieval
The response contains the same candidates sorted by descending relevance score, with the most pertinent passage first.
Agent-Memory Retrieval Endpoint
POST /v1/mtplx/retrieval/agent_memory
The agent-memory endpoint provides per-session persistent storage, enabling stateful agent interactions across multiple API calls. Unlike the stateless RAG endpoints, this maintains a "long-term" memory store tied to a session identifier.
# Retrieve stored memories for an active agent session
mem_resp = requests.post(
f"{BASE_URL}/v1/mtplx/retrieval/agent_memory",
json={"session_id": "abc123", "limit": 10}
)
memory_items = mem_resp.json()["memories"]
The endpoint accepts:
session_id(required): Unique identifier for the agent sessionlimit(optional): Maximum number of memory items to return- Additional filters for memory type or time range
Response contents include previously generated outputs, tool-call results, and intermediate reasoning steps—enabling agents to reference prior context without re-embedding entire conversation histories.
Implementation Details and Source Files
According to the MTPLX source code, these endpoints are registered through FastAPI decorators in mtplx/server/openai.py:
@app.post("/v1/mtplx/retrieval/embedding")
@app.post("/v1/mtplx/retrieval/rerank")
@app.post("/v1/mtplx/retrieval/agent_memory")
The actual retrieval logic resides in mtplx/retrieval.py, which handles model loading, embedding computation, cross-encoder reranking, and agent-memory management. CLI interactions are available through mtplx/commands/public.py with commands like mtplx retrieval embed.
Server capabilities are advertised via GET /v1/mtplx/app/capabilities (implemented in mtplx/server/dashboard_state.py), allowing UI clients to discover available retrieval services dynamically.
Endpoint Summary Table
| Service | Method | Path | Payload | Response |
|---|---|---|---|---|
| Embedding generation | POST |
/v1/mtplx/retrieval/embedding |
{"text": "..."} |
{"embeddings": [...]} |
| Passage reranking | POST |
/v1/mtplx/retrieval/rerank |
{"query": "...", "candidates": [...]} |
{"candidates": [...]} |
| Agent memory access | POST |
/v1/mtplx/retrieval/agent_memory |
{"session_id": "...", "limit": N} |
{"memories": [...]} |
Summary
- MTPLX provides three retrieval endpoints under the
/v1/mtplx/retrieval/namespace: embedding generation, passage reranking, and agent-memory access. - RAG pipelines use the embedding and rerank endpoints sequentially for high-quality context retrieval.
- Agent memory enables persistent, session-scoped storage for stateful agent applications.
- All endpoints are POST methods accepting JSON payloads and returning structured responses.
- Source implementation is located in
mtplx/server/openai.pywith core logic inmtplx/retrieval.py.
Frequently Asked Questions
What is the base URL for MTPLX retrieval endpoints?
MTPLX retrieval endpoints are served from the FastAPI server base URL, typically http://localhost:8000 during local development. All paths are prefixed with /v1/mtplx/retrieval/ as implemented in mtplx/server/openai.py.
Can I use the embedding endpoint without the reranking endpoint?
Yes. The embedding endpoint functions independently for basic semantic search. The reranking endpoint is optional and provides quality improvements for applications requiring high-precision retrieval, particularly with large candidate sets.
How does agent memory differ from standard RAG context?
Agent memory is persistent and session-scoped, storing structured data like tool outputs and generated snippets across multiple API calls. Standard RAG context is ephemeral, computed fresh per request from external document indexes. Agent memory enables multi-turn agent workflows; RAG provides external knowledge grounding.
Where are the retrieval models loaded from in MTPLX?
Model loading and inference are handled in mtplx/retrieval.py. The specific embedding and reranking models are configured through MTPLX's configuration system and loaded at server startup, with inference executed on-demand through the FastAPI endpoints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →