Key Files in the Headroom Source Code: A Complete Technical Guide
The Headroom source code is organized into a Python package where headroom/__init__.py exposes the public compress function, headroom/transforms/pipeline.py orchestrates the compression stages, and specialized modules like headroom/proxy/server.py and headroom/transforms/smart_crusher.py handle proxy interception and content-specific algorithms.
Headroom is an open-source LLM token compression framework hosted at chopratejas/headroom that reduces API costs by 60–95%. Understanding the key files in the Headroom source code reveals how the SDK, proxy server, and transform modules work together to compress prompts without requiring application changes.
Package Entry and Public API
The headroom/__init__.py file serves as the primary entry point for the SDK. It exposes the compress function and top-level objects that developers import directly.
When you call from headroom import compress, you are invoking the logic defined in this module. The function signature accepts a messages array and model identifier, delegating execution to the transform pipeline internally.
Command-Line Interface and Client Wrappers
Headroom provides multiple ways to integrate: a direct CLI and thin client wrappers.
The headroom/cli.py module parses arguments and drives the various modes including wrap, proxy, and learn. This is the entry point when running headroom commands from the terminal.
For programmatic use, headroom/client.py contains the HeadroomChatModel class and withHeadroom wrapper. These provide LLM-agnostic abstractions that wrap existing clients (OpenAI, Anthropic, etc.) to automatically apply compression.
Proxy Server Architecture
The proxy implementation allows interception of any LLM request without code changes.
The core HTTP server resides in headroom/proxy/server.py, which runs a local server that processes outgoing requests. Supporting files include headroom/proxy/runtime_env.py for environment management and headroom/proxy/output_shaper.py for response formatting.
This architecture enables drop-in usage where existing applications can point their API endpoint to the local proxy and receive automatic compression.
Transform Pipeline and Content Routing
The compression logic flows through a pipeline architecture defined in headroom/transforms/pipeline.py. This module orchestrates the sequence of operations from input validation to final output.
Before compression begins, headroom/transforms/content_router.py analyzes the payload to select the appropriate compressor. The ContentRouter inspects content type and structure to determine whether to apply JSON crushing, code compression, or prose compression.
Compression Algorithms
The actual token reduction happens in three specialized compressor modules:
headroom/transforms/smart_crusher.py– Handles JSON and structured data using theSmartCrusheralgorithm.headroom/transforms/code_compressor.py– Provides AST-aware source code compression via theCodeCompressorclass.headroom/transforms/kompress_compressor.py– Implements text compression using the Kompress-v2 HuggingFace model for prose content.
Each module operates independently and is selected by the content router based on the input data type.
Cross-Agent Memory and Storage
Headroom maintains state across multiple agents and sessions through persistent layers.
headroom/shared_context.py implements a shared store that lets multiple agents reuse the same compressed context, preventing redundant compression of identical prefixes.
For reversible compression, headroom/storage/sqlite.py provides a SQLite-backed cache that stores original payloads. This enables CCR (Context Compression and Retrieval), where the system can restore original data when requested by the LLM.
Supporting Utilities and Telemetry
Additional infrastructure modules support the core functionality:
headroom/tokenizers/tiktoken_counter.py– Fast token-count utilities used for cost estimates and cache alignment.headroom/telemetry/reporter.py– Emits usage metrics and compression statistics.headroom/reporting/generator.py– Builds human-readable reports from telemetry data.headroom/image/compressor.py– ONNX-based router for image compression decisions.headroom/learn/verbosity.pyandheadroom/learn/writer.py– Mine failed sessions to automatically adjust verbosity or add correction rules.headroom/rtk/installer.py– Bundles the RTK binary used for shell-output rewriting.headroom/utils.py– Contains helpers for hashing, markers, and safe JSON parsing, includingcompute_prefix_hash.
How the Components Work Together
The following flow illustrates how these key files interact during a typical compression cycle:
-
User code calls
compress(messages, model=...)fromheadroom/__init__.pyor runsheadroom wrap, which instantiates the pipeline fromheadroom/transforms/pipeline.py. -
The pipeline uses
ContentRouterinheadroom/transforms/content_router.pyto select the appropriate compressor based on content type. -
The selected compressor (from
smart_crusher.py,code_compressor.py, orkompress_compressor.py) reduces the payload and adds markers created viaheadroom/utils.py. -
A hash of the compressed prefix is computed using
utils.compute_prefix_hashand stored in the CCR cache viaheadroom/storage/sqlite.py. -
If the LLM requires the original data, the proxy in
headroom/proxy/server.pylooks up the original in the cache and returns it. -
Telemetry modules record statistics, while
headroom/shared_context.pyenables multiple agents to share compressed contexts.
Code Examples
Using the SDK Inline
from headroom import compress
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain how Headroom reduces token usage."},
]
# Compress the conversation before sending it to an LLM
compressed, stats = compress(messages, model="gpt-4o")
print("Compressed payload:", compressed)
print("Saved tokens:", stats.input_tokens_saved)
This example imports the compress function from headroom/__init__.py, which delegates to the transform pipeline.
Running the Proxy
# Start a drop-in proxy on port 8787
headroom proxy --port 8787
This command loads headroom/proxy/server.py and wires together the same pipeline used by the SDK, allowing any downstream client (Anthropic SDK, OpenAI SDK, LangChain) to automatically benefit from compression without code changes.
Summary
headroom/__init__.pyexposes the publiccompressAPI and top-level objects.headroom/cli.pyprovides the command-line entry point forwrap,proxy, andlearnmodes.headroom/transforms/pipeline.pyorchestrates the compression stages, whilecontent_router.pyselects the appropriate algorithm.headroom/transforms/smart_crusher.py,code_compressor.py, andkompress_compressor.pyimplement the actual compression algorithms for JSON, code, and prose respectively.headroom/proxy/server.pyenables drop-in proxy deployment without application modifications.headroom/storage/sqlite.pyandshared_context.pyprovide persistent caching and cross-agent memory.headroom/utils.pycontains critical helpers likecompute_prefix_hashand marker generation.
Frequently Asked Questions
Where is the main entry point for the Headroom SDK?
The main entry point is headroom/__init__.py, which exports the compress function and other public objects. When you write from headroom import compress, you are importing from this file, which then delegates to the transform pipeline.
How does Headroom decide which compression algorithm to use?
The headroom/transforms/content_router.py module contains the ContentRouter class that analyzes the content type and structure. It routes JSON data to smart_crusher.py, source code to code_compressor.py, and prose text to kompress_compressor.py.
What enables the proxy mode to work without code changes?
The headroom/proxy/server.py module implements an HTTP server that intercepts requests to standard LLM API endpoints. Combined with headroom/proxy/runtime_env.py for environment management, this allows existing applications to point their base URL to the local proxy, automatically triggering compression via the same pipeline used by the SDK.
How does Headroom store data for reversible compression?
The system uses headroom/storage/sqlite.py to maintain a SQLite-backed cache of original payloads. When compression occurs, headroom/utils.py computes a prefix hash that serves as the lookup key. If the LLM later needs the original data, the proxy retrieves it from this cache, enabling full round-trip compression and decompression.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →