Headroom Library Mode vs Proxy Mode: Key Differences and When to Use Each

Library mode embeds Headroom's compress() function directly into your application for inline message processing, while proxy mode runs a standalone HTTP server that intercepts LLM requests without requiring code changes.

Headroom, the open-source LLM message compression tool developed by chopratejas/headroom, provides two distinct integration strategies that suit different architectural needs. Understanding the difference between library mode and proxy mode helps you optimize how the compression pipeline—ContentRouter, SmartCrusher, and CCR—integrates with your existing systems. Both modes execute identical transforms but differ in entry points, language support, and deployment complexity.

How Library Mode Works

Direct Function Integration

Library mode exposes the core compression logic through the compress() function defined in headroom/__init__.py. When you call this function, Headroom executes the full pipeline—ContentRouter → SmartCrusher → CCR—inside your application process without spawning external servers. This mode requires you to import the library and explicitly pass your message arrays to the compression function.

from headroom import compress

messages = [
    {"role": "user", "content": "Explain the difference between library and proxy modes."}
]

# Execute compression pipeline inline

compressed = compress(messages, model="gpt-4o")

# Send compressed payload to your LLM provider

Code Changes and Scope

This approach requires modifying your codebase to import headroom and wrap your LLM calls. It compresses only the messages you explicitly pass to the function, giving you surgical control over what gets processed through the transforms in headroom/transforms/*.

How Proxy Mode Works

HTTP Interception Architecture

Proxy mode implements an ASGI server in headroom/proxy/server.py that sits between your application and LLM providers. When you run headroom proxy --port 8787, the server starts listening for HTTP requests and routes them through provider-specific handlers located in headroom/proxy/handlers/* (supporting OpenAI, Anthropic, Gemini, and others) before forwarding to the actual API endpoints.


# Start the Headroom proxy server

headroom proxy --port 8787

# Configure your application or environment to use the proxy

export http_proxy=http://127.0.0.1:8787

# Any standard HTTP client now routes through Headroom automatically

curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d '{"model":"gpt-4o","messages":[{"role":"user","content":"Hi"}]}'

Zero-Code Integration

Unlike library mode, proxy mode requires zero code changes to your existing applications. Any client that can route HTTP requests through a proxy—or target localhost:8787 directly—will automatically have its messages compressed, including tool outputs, RAG chunks, and file reads that library mode might miss.

Key Differences Between Library Mode and Proxy Mode

Feature Library Mode Proxy Mode
Entry Point compress() function in headroom/__init__.py headroom proxy command invoking headroom/proxy/server.py
Code Changes Requires Python import and function calls No code changes; HTTP proxy configuration only
Language Support Python-specific (as implemented) Language-agnostic; works with any HTTP-speaking client
Processing Scope Only messages explicitly passed to function All traffic routed through the proxy
Architecture In-process execution Separate ASGI server process
Provider Handling Generic message compression Provider-specific handlers in headroom/proxy/handlers/*

The operational modes are formally declared in headroom/proxy/modes.py, which distinguishes between inline library usage and the proxy server configuration.

When to Use Library Mode vs Proxy Mode

Choose Library Mode When

You need explicit control over exactly which messages enter the compression pipeline. Library mode is optimal for Python-native applications where you want to minimize network overhead by running the ContentRouter and SmartCrusher transforms inside your existing process. Use this when building custom SDKs or when you cannot deploy a separate proxy service.

Choose Proxy Mode When

You are integrating with existing agents or tools that already call LLM endpoints directly and cannot be modified. Proxy mode excels when supporting multiple languages or when you need to capture auxiliary traffic—such as tool outputs and file system reads—that occurs outside your main application code. The proxy server handles the routing logic defined in headroom/proxy/modes.py transparently.

Summary

  • Library mode integrates via the compress() function in headroom/__init__.py, executing the compression pipeline in-process for Python applications requiring precise control.
  • Proxy mode deploys as an ASGI server via headroom/proxy/server.py, intercepting HTTP traffic through provider-specific handlers without code modifications.
  • Both modes utilize the identical transform chain located in headroom/transforms/* (ContentRouter, SmartCrusher, CCR).
  • Library mode suits Python deployments where you control the request generation; proxy mode suits polyglot environments and legacy system integration.

Frequently Asked Questions

Can I switch between library mode and proxy mode without changing compression behavior?

Yes. Both modes invoke the same core pipeline stored in headroom/transforms/*. The compress() function and the proxy server both route messages through ContentRouter, SmartCrusher, and CCR transforms. The operational distinction in headroom/proxy/modes.py only affects how the messages reach the pipeline, not how they are processed.

Does proxy mode support streaming responses from LLM providers?

The proxy server in headroom/proxy/server.py implements provider-specific handlers in headroom/proxy/handlers/* that maintain compatibility with standard LLM API formats, including streaming protocols. Since the proxy intercepts HTTP traffic before forwarding to providers like OpenAI or Anthropic, it can handle chunked transfer encoding and server-sent events transparently while compressing the incoming request payload.

Is there a performance penalty for using proxy mode over library mode?

Proxy mode introduces minimal latency from the additional HTTP hop to headroom/proxy/server.py, but it allows you to horizontally scale the compression service independently of your application. Library mode eliminates network overhead by running transforms inside your process, but it consumes application memory and CPU cycles directly. Choose proxy mode when you need to offload processing or run the compressor on dedicated infrastructure.

Can I use proxy mode with CLI tools like Copilot or LangChain?

Absolutely. Proxy mode requires no code changes and works with any tool that respects HTTP proxy environment variables or can be configured to use a custom endpoint. By setting http_proxy=http://127.0.0.1:8787 before running CLI tools, or by configuring LangChain to point to localhost:8787, you route all LLM traffic through Headroom's compression pipeline automatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →