Key Files in the Headroom Source Code: A Complete Technical Guide

The Headroom source code is organized into a Python package where headroom/__init__.py exposes the public compress function, headroom/transforms/pipeline.py orchestrates the compression stages, and specialized modules like headroom/proxy/server.py and headroom/transforms/smart_crusher.py handle proxy interception and content-specific algorithms.

Headroom is an open-source LLM token compression framework hosted at chopratejas/headroom that reduces API costs by 60–95%. Understanding the key files in the Headroom source code reveals how the SDK, proxy server, and transform modules work together to compress prompts without requiring application changes.

Package Entry and Public API

The headroom/__init__.py file serves as the primary entry point for the SDK. It exposes the compress function and top-level objects that developers import directly.

When you call from headroom import compress, you are invoking the logic defined in this module. The function signature accepts a messages array and model identifier, delegating execution to the transform pipeline internally.

Command-Line Interface and Client Wrappers

Headroom provides multiple ways to integrate: a direct CLI and thin client wrappers.

The headroom/cli.py module parses arguments and drives the various modes including wrap, proxy, and learn. This is the entry point when running headroom commands from the terminal.

For programmatic use, headroom/client.py contains the HeadroomChatModel class and withHeadroom wrapper. These provide LLM-agnostic abstractions that wrap existing clients (OpenAI, Anthropic, etc.) to automatically apply compression.

Proxy Server Architecture

The proxy implementation allows interception of any LLM request without code changes.

The core HTTP server resides in headroom/proxy/server.py, which runs a local server that processes outgoing requests. Supporting files include headroom/proxy/runtime_env.py for environment management and headroom/proxy/output_shaper.py for response formatting.

This architecture enables drop-in usage where existing applications can point their API endpoint to the local proxy and receive automatic compression.

Transform Pipeline and Content Routing

The compression logic flows through a pipeline architecture defined in headroom/transforms/pipeline.py. This module orchestrates the sequence of operations from input validation to final output.

Before compression begins, headroom/transforms/content_router.py analyzes the payload to select the appropriate compressor. The ContentRouter inspects content type and structure to determine whether to apply JSON crushing, code compression, or prose compression.

Compression Algorithms

The actual token reduction happens in three specialized compressor modules:

Each module operates independently and is selected by the content router based on the input data type.

Cross-Agent Memory and Storage

Headroom maintains state across multiple agents and sessions through persistent layers.

headroom/shared_context.py implements a shared store that lets multiple agents reuse the same compressed context, preventing redundant compression of identical prefixes.

For reversible compression, headroom/storage/sqlite.py provides a SQLite-backed cache that stores original payloads. This enables CCR (Context Compression and Retrieval), where the system can restore original data when requested by the LLM.

Supporting Utilities and Telemetry

Additional infrastructure modules support the core functionality:

How the Components Work Together

The following flow illustrates how these key files interact during a typical compression cycle:

  1. User code calls compress(messages, model=...) from headroom/__init__.py or runs headroom wrap, which instantiates the pipeline from headroom/transforms/pipeline.py.

  2. The pipeline uses ContentRouter in headroom/transforms/content_router.py to select the appropriate compressor based on content type.

  3. The selected compressor (from smart_crusher.py, code_compressor.py, or kompress_compressor.py) reduces the payload and adds markers created via headroom/utils.py.

  4. A hash of the compressed prefix is computed using utils.compute_prefix_hash and stored in the CCR cache via headroom/storage/sqlite.py.

  5. If the LLM requires the original data, the proxy in headroom/proxy/server.py looks up the original in the cache and returns it.

  6. Telemetry modules record statistics, while headroom/shared_context.py enables multiple agents to share compressed contexts.

Code Examples

Using the SDK Inline

from headroom import compress

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain how Headroom reduces token usage."},
]

# Compress the conversation before sending it to an LLM

compressed, stats = compress(messages, model="gpt-4o")
print("Compressed payload:", compressed)
print("Saved tokens:", stats.input_tokens_saved)

This example imports the compress function from headroom/__init__.py, which delegates to the transform pipeline.

Running the Proxy


# Start a drop-in proxy on port 8787

headroom proxy --port 8787

This command loads headroom/proxy/server.py and wires together the same pipeline used by the SDK, allowing any downstream client (Anthropic SDK, OpenAI SDK, LangChain) to automatically benefit from compression without code changes.

Summary

Frequently Asked Questions

Where is the main entry point for the Headroom SDK?

The main entry point is headroom/__init__.py, which exports the compress function and other public objects. When you write from headroom import compress, you are importing from this file, which then delegates to the transform pipeline.

How does Headroom decide which compression algorithm to use?

The headroom/transforms/content_router.py module contains the ContentRouter class that analyzes the content type and structure. It routes JSON data to smart_crusher.py, source code to code_compressor.py, and prose text to kompress_compressor.py.

What enables the proxy mode to work without code changes?

The headroom/proxy/server.py module implements an HTTP server that intercepts requests to standard LLM API endpoints. Combined with headroom/proxy/runtime_env.py for environment management, this allows existing applications to point their base URL to the local proxy, automatically triggering compression via the same pipeline used by the SDK.

How does Headroom store data for reversible compression?

The system uses headroom/storage/sqlite.py to maintain a SQLite-backed cache of original payloads. When compression occurs, headroom/utils.py computes a prefix hash that serves as the lookup key. If the LLM later needs the original data, the proxy retrieves it from this cache, enabling full round-trip compression and decompression.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →