RAG and Agent-Memory Retrieval Endpoints in MTPLX: Complete API Reference

MTPLX exposes three production-ready retrieval endpoints under /v1/mtplx/retrieval/ for embeddings, reranking, and agent-memory access.

The MTPLX open-source framework provides a FastAPI-based server with dedicated endpoints for Retrieval-Augmented Generation (RAG) pipelines and stateful agent memory. These endpoints enable developers to build context-aware LLM applications with persistent session storage and sophisticated document retrieval. This guide covers each endpoint's purpose, request format, and implementation details drawn directly from the MTPLX source code.

RAG Endpoints: Embedding and Reranking

MTPLX implements the two classic stages of RAG through separate HTTP endpoints. Both are defined in mtplx/server/openai.py and registered during server initialization.

POST /v1/mtplx/retrieval/embedding

The embedding endpoint converts raw text into dense vectors for semantic search. This is the foundation of any RAG pipeline, enabling similarity-based document retrieval.

import requests

BASE_URL = "http://localhost:8000"

# Generate embeddings for RAG query

embed_resp = requests.post(
    f"{BASE_URL}/v1/mtplx/retrieval/embedding",
    json={"text": "What is the capital of France?"}
)
embeddings = embed_resp.json()["embeddings"]

The endpoint accepts a JSON payload with a text field and returns a vector representation suitable for cosine similarity search against a document index.

POST /v1/mtplx/retrieval/rerank

The reranking endpoint refines retrieval results using a cross-encoder model. After initial candidate selection via vector similarity, this endpoint reorders passages by true relevance to the query.


# Rerank candidate passages for improved RAG accuracy

rerank_resp = requests.post(
    f"{BASE_URL}/v1/mtplx/retrieval/rerank",
    json={
        "query": "What is the capital of France?",
        "candidates": [
            "Paris is the capital of France.",
            "Berlin is the capital of Germany."
        ]
    }
)
ranked = rerank_resp.json()["candidates"]

The request body requires:

  • query: The original user question
  • candidates: List of passage strings from initial retrieval

The response contains the same candidates sorted by descending relevance score, with the most pertinent passage first.

Agent-Memory Retrieval Endpoint

POST /v1/mtplx/retrieval/agent_memory

The agent-memory endpoint provides per-session persistent storage, enabling stateful agent interactions across multiple API calls. Unlike the stateless RAG endpoints, this maintains a "long-term" memory store tied to a session identifier.


# Retrieve stored memories for an active agent session

mem_resp = requests.post(
    f"{BASE_URL}/v1/mtplx/retrieval/agent_memory",
    json={"session_id": "abc123", "limit": 10}
)
memory_items = mem_resp.json()["memories"]

The endpoint accepts:

  • session_id (required): Unique identifier for the agent session
  • limit (optional): Maximum number of memory items to return
  • Additional filters for memory type or time range

Response contents include previously generated outputs, tool-call results, and intermediate reasoning steps—enabling agents to reference prior context without re-embedding entire conversation histories.

Implementation Details and Source Files

According to the MTPLX source code, these endpoints are registered through FastAPI decorators in mtplx/server/openai.py:

@app.post("/v1/mtplx/retrieval/embedding")
@app.post("/v1/mtplx/retrieval/rerank")
@app.post("/v1/mtplx/retrieval/agent_memory")

The actual retrieval logic resides in mtplx/retrieval.py, which handles model loading, embedding computation, cross-encoder reranking, and agent-memory management. CLI interactions are available through mtplx/commands/public.py with commands like mtplx retrieval embed.

Server capabilities are advertised via GET /v1/mtplx/app/capabilities (implemented in mtplx/server/dashboard_state.py), allowing UI clients to discover available retrieval services dynamically.

Endpoint Summary Table

Service Method Path Payload Response
Embedding generation POST /v1/mtplx/retrieval/embedding {"text": "..."} {"embeddings": [...]}
Passage reranking POST /v1/mtplx/retrieval/rerank {"query": "...", "candidates": [...]} {"candidates": [...]}
Agent memory access POST /v1/mtplx/retrieval/agent_memory {"session_id": "...", "limit": N} {"memories": [...]}

Summary

  • MTPLX provides three retrieval endpoints under the /v1/mtplx/retrieval/ namespace: embedding generation, passage reranking, and agent-memory access.
  • RAG pipelines use the embedding and rerank endpoints sequentially for high-quality context retrieval.
  • Agent memory enables persistent, session-scoped storage for stateful agent applications.
  • All endpoints are POST methods accepting JSON payloads and returning structured responses.
  • Source implementation is located in mtplx/server/openai.py with core logic in mtplx/retrieval.py.

Frequently Asked Questions

What is the base URL for MTPLX retrieval endpoints?

MTPLX retrieval endpoints are served from the FastAPI server base URL, typically http://localhost:8000 during local development. All paths are prefixed with /v1/mtplx/retrieval/ as implemented in mtplx/server/openai.py.

Can I use the embedding endpoint without the reranking endpoint?

Yes. The embedding endpoint functions independently for basic semantic search. The reranking endpoint is optional and provides quality improvements for applications requiring high-precision retrieval, particularly with large candidate sets.

How does agent memory differ from standard RAG context?

Agent memory is persistent and session-scoped, storing structured data like tool outputs and generated snippets across multiple API calls. Standard RAG context is ephemeral, computed fresh per request from external document indexes. Agent memory enables multi-turn agent workflows; RAG provides external knowledge grounding.

Where are the retrieval models loaded from in MTPLX?

Model loading and inference are handled in mtplx/retrieval.py. The specific embedding and reranking models are configured through MTPLX's configuration system and loaded at server startup, with inference executed on-demand through the FastAPI endpoints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →