How to Set Up Headroom as an MCP Server: Complete Configuration Guide
Headroom exposes its context compression and memory capabilities as a Model Context Protocol (MCP) server through two distinct modes: a lightweight MCP-only setup for direct tool access, or a full MCP+Proxy configuration that automatically compresses HTTP traffic while providing on-demand tools.
The chopratejas/headroom repository provides a powerful context compression and memory system designed for LLM workflows. When you set up Headroom as an MCP server, you enable any MCP-compatible client—such as Claude Code, Cursor, or Codex—to access persistent semantic memory and intelligent text compression directly through standardized tool calls.
Understanding Headroom's MCP Architecture
Headroom supports two distinct deployment modes for MCP integration, each suited to different workflow requirements.
MCP-Only Mode (No Proxy)
In this configuration, the client communicates directly with Headroom's MCP server. Compression operations run locally within the MCP process, with original content stored in a short-lived local cache. This mode is ideal for lightweight deployments where you need access to Headroom's tools without the overhead of HTTP interception.
MCP + Proxy Mode (Full Setup)
This complete configuration runs Headroom's proxy alongside the MCP server. The proxy automatically compresses HTTP traffic between your client and the LLM provider, while the MCP server provides supplementary on-demand tools. Critical to this architecture: the system never double-compresses data, ensuring efficient bandwidth usage.
Installation and Prerequisites
Before configuring the server, install the appropriate dependencies based on your chosen mode.
Installing MCP Dependencies
For the full experience including the HTTP proxy:
pip install "headroom-ai[proxy]"
For a lightweight MCP-only installation:
pip install "headroom-ai[mcp]"
The [proxy] extra includes both the proxy components and MCP tools, while [mcp] provides only the lightweight Model Context Protocol server implementation.
Configuring Headroom as an MCP Server
Automatic Setup with the CLI
The simplest method uses the built-in CLI command defined in headroom/cli/mcp.py:
headroom mcp install
This command automatically registers Headroom with Claude Code by writing a configuration file to ~/.claude/mcp.json, pointing the client to the MCP server entry point.
Manual Server Launch
For testing or custom integrations, launch the server directly:
python -m headroom.memory.mcp_server --db ~/.headroom/memory.db
The server communicates via stdio and logs progress to stderr. In headroom/memory/mcp_server.py, the implementation pre-loads an ONNX embedder and re-indexes missing vectors before serving requests.
Proxy Configuration (Optional)
To enable automatic HTTP compression, start the proxy in a separate terminal:
headroom proxy
The proxy listens on 127.0.0.1:8787. Configure your client to route through it:
ANTHROPIC_BASE_URL=http://127.0.0.1:8787 claude
Available MCP Tools and Capabilities
Headroom provides two distinct MCP server implementations, each exposing specific toolsets.
Memory Management Tools (headroom/memory/mcp_server.py)
The memory server exposes semantic search capabilities:
memory_search– Executes semantic search across persisted memories using the pre-loaded ONNX embeddermemory_save– Stores new facts and embeddings for later retrieval
Compression and Retrieval Tools (headroom/ccr/mcp_server.py)
The CCR (Context Compression and Retrieval) server provides content optimization:
headroom_compress– Compresses large text inputs and returns a content hashheadroom_retrieve– Fetches original content by hash, with optional query filteringheadroom_stats– Returns compression statistics and aggregates sub-agent metrics when present
Programmatic Integration
Embed the MCP server directly in Python applications using the create_memory_server function from headroom/memory/mcp_server.py:
from headroom.memory.mcp_server import create_memory_server
from mcp.server.stdio import stdio_server
import asyncio
async def run():
server = create_memory_server(db_path="~/.headroom/memory.db")
async with stdio_server() as (read, write):
await server.run(read, write, server.create_initialization_options())
if __name__ == "__main__":
asyncio.run(run())
This mirrors the main() function implementation in the source code, giving you full control over the server's lifecycle and database location.
Debugging and Verification
Verify your installation and server health using the CLI:
headroom mcp status # Checks SDK, Claude config, and proxy health
headroom mcp serve --debug # Runs with verbose logging
Summary
- Headroom operates as an MCP server in two modes: MCP-only for direct tool access, and MCP+Proxy for automatic HTTP compression
- Install via
pip install "headroom-ai[proxy]"for full functionality or"headroom-ai[mcp]"for lightweight deployment - Register automatically with
headroom mcp installor manually launch viapython -m headroom.memory.mcp_server - The memory server (
headroom/memory/mcp_server.py) providesmemory_searchandmemory_savetools - The CCR server (
headroom/ccr/mcp_server.py) exposesheadroom_compress,headroom_retrieve, andheadroom_stats - Run
headroom mcp statusto verify configuration andheadroom mcp serve --debugfor troubleshooting
Frequently Asked Questions
What is the difference between MCP-only mode and MCP+Proxy mode?
MCP-only mode runs the server without HTTP interception, processing compression and memory operations locally through the MCP protocol. MCP+Proxy mode adds the Headroom proxy (running on 127.0.0.1:8787) which automatically compresses HTTP traffic between your client and the LLM provider, while the MCP server handles additional on-demand tool requests without double-compressing data.
Which pip extras should I install for Headroom MCP?
Install "headroom-ai[proxy]" if you need both the HTTP proxy and MCP tools, which provides the complete functionality described in wiki/mcp.md. Install "headroom-ai[mcp]" for a lightweight deployment that includes only the Model Context Protocol server components without proxy dependencies.
How do I verify my Headroom MCP server is running correctly?
Use the headroom mcp status command to verify SDK installation, Claude Code configuration, and proxy health. For detailed debugging, run headroom mcp serve --debug to start the server with verbose logging to stderr, or manually launch python -m headroom.memory.mcp_server to observe the initialization process including ONNX embedder loading and vector re-indexing.
Can I run Headroom MCP without the proxy?
Yes. The MCP server functions independently in MCP-only mode. When you run python -m headroom.memory.mcp_server or install via pip install "headroom-ai[mcp]", you get direct access to memory_search, memory_save, and compression tools without requiring the HTTP proxy component. This is suitable for environments where you cannot intercept HTTP traffic or prefer explicit tool invocation over automatic compression.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →