# Key Files in the Headroom Source Code: A Complete Technical Guide

> Explore the headroom source code with this technical guide. Discover key files like init py, pipeline py, and server py to understand compression and proxy functionality.

- Repository: [Tejas Chopra/headroom](https://github.com/chopratejas/headroom)
- Tags: deep-dive
- Published: 2026-06-21

---

**The Headroom source code is organized into a Python package where [`headroom/__init__.py`](https://github.com/chopratejas/headroom/blob/main/headroom/__init__.py) exposes the public `compress` function, [`headroom/transforms/pipeline.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/pipeline.py) orchestrates the compression stages, and specialized modules like [`headroom/proxy/server.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/server.py) and [`headroom/transforms/smart_crusher.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/smart_crusher.py) handle proxy interception and content-specific algorithms.**

Headroom is an open-source LLM token compression framework hosted at `chopratejas/headroom` that reduces API costs by 60–95%. Understanding the key files in the Headroom source code reveals how the SDK, proxy server, and transform modules work together to compress prompts without requiring application changes.

## Package Entry and Public API

The [`headroom/__init__.py`](https://github.com/chopratejas/headroom/blob/main/headroom/__init__.py) file serves as the primary entry point for the SDK. It exposes the **`compress`** function and top-level objects that developers import directly.

When you call `from headroom import compress`, you are invoking the logic defined in this module. The function signature accepts a messages array and model identifier, delegating execution to the transform pipeline internally.

## Command-Line Interface and Client Wrappers

Headroom provides multiple ways to integrate: a direct CLI and thin client wrappers.

The **[`headroom/cli.py`](https://github.com/chopratejas/headroom/blob/main/headroom/cli.py)** module parses arguments and drives the various modes including `wrap`, `proxy`, and `learn`. This is the entry point when running `headroom` commands from the terminal.

For programmatic use, **[`headroom/client.py`](https://github.com/chopratejas/headroom/blob/main/headroom/client.py)** contains the **`HeadroomChatModel`** class and **`withHeadroom`** wrapper. These provide LLM-agnostic abstractions that wrap existing clients (OpenAI, Anthropic, etc.) to automatically apply compression.

## Proxy Server Architecture

The proxy implementation allows interception of any LLM request without code changes.

The core HTTP server resides in **[`headroom/proxy/server.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/server.py)**, which runs a local server that processes outgoing requests. Supporting files include **[`headroom/proxy/runtime_env.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/runtime_env.py)** for environment management and **[`headroom/proxy/output_shaper.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/output_shaper.py)** for response formatting.

This architecture enables drop-in usage where existing applications can point their API endpoint to the local proxy and receive automatic compression.

## Transform Pipeline and Content Routing

The compression logic flows through a pipeline architecture defined in **[`headroom/transforms/pipeline.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/pipeline.py)**. This module orchestrates the sequence of operations from input validation to final output.

Before compression begins, **[`headroom/transforms/content_router.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/content_router.py)** analyzes the payload to select the appropriate compressor. The **`ContentRouter`** inspects content type and structure to determine whether to apply JSON crushing, code compression, or prose compression.

## Compression Algorithms

The actual token reduction happens in three specialized compressor modules:

- **[`headroom/transforms/smart_crusher.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/smart_crusher.py)** – Handles JSON and structured data using the `SmartCrusher` algorithm.
- **[`headroom/transforms/code_compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/code_compressor.py)** – Provides AST-aware source code compression via the `CodeCompressor` class.
- **[`headroom/transforms/kompress_compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/kompress_compressor.py)** – Implements text compression using the Kompress-v2 HuggingFace model for prose content.

Each module operates independently and is selected by the content router based on the input data type.

## Cross-Agent Memory and Storage

Headroom maintains state across multiple agents and sessions through persistent layers.

**[`headroom/shared_context.py`](https://github.com/chopratejas/headroom/blob/main/headroom/shared_context.py)** implements a shared store that lets multiple agents reuse the same compressed context, preventing redundant compression of identical prefixes.

For reversible compression, **[`headroom/storage/sqlite.py`](https://github.com/chopratejas/headroom/blob/main/headroom/storage/sqlite.py)** provides a SQLite-backed cache that stores original payloads. This enables CCR (Context Compression and Retrieval), where the system can restore original data when requested by the LLM.

## Supporting Utilities and Telemetry

Additional infrastructure modules support the core functionality:

- **[`headroom/tokenizers/tiktoken_counter.py`](https://github.com/chopratejas/headroom/blob/main/headroom/tokenizers/tiktoken_counter.py)** – Fast token-count utilities used for cost estimates and cache alignment.
- **[`headroom/telemetry/reporter.py`](https://github.com/chopratejas/headroom/blob/main/headroom/telemetry/reporter.py)** – Emits usage metrics and compression statistics.
- **[`headroom/reporting/generator.py`](https://github.com/chopratejas/headroom/blob/main/headroom/reporting/generator.py)** – Builds human-readable reports from telemetry data.
- **[`headroom/image/compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/image/compressor.py)** – ONNX-based router for image compression decisions.
- **[`headroom/learn/verbosity.py`](https://github.com/chopratejas/headroom/blob/main/headroom/learn/verbosity.py)** and **[`headroom/learn/writer.py`](https://github.com/chopratejas/headroom/blob/main/headroom/learn/writer.py)** – Mine failed sessions to automatically adjust verbosity or add correction rules.
- **[`headroom/rtk/installer.py`](https://github.com/chopratejas/headroom/blob/main/headroom/rtk/installer.py)** – Bundles the RTK binary used for shell-output rewriting.
- **[`headroom/utils.py`](https://github.com/chopratejas/headroom/blob/main/headroom/utils.py)** – Contains helpers for hashing, markers, and safe JSON parsing, including **`compute_prefix_hash`**.

## How the Components Work Together

The following flow illustrates how these key files interact during a typical compression cycle:

1. **User code** calls `compress(messages, model=...)` from [`headroom/__init__.py`](https://github.com/chopratejas/headroom/blob/main/headroom/__init__.py) or runs `headroom wrap`, which instantiates the pipeline from [`headroom/transforms/pipeline.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/pipeline.py).

2. The pipeline uses **`ContentRouter`** in [`headroom/transforms/content_router.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/content_router.py) to select the appropriate compressor based on content type.

3. The selected compressor (from [`smart_crusher.py`](https://github.com/chopratejas/headroom/blob/main/smart_crusher.py), [`code_compressor.py`](https://github.com/chopratejas/headroom/blob/main/code_compressor.py), or [`kompress_compressor.py`](https://github.com/chopratejas/headroom/blob/main/kompress_compressor.py)) reduces the payload and adds markers created via [`headroom/utils.py`](https://github.com/chopratejas/headroom/blob/main/headroom/utils.py).

4. A hash of the compressed prefix is computed using **`utils.compute_prefix_hash`** and stored in the CCR cache via [`headroom/storage/sqlite.py`](https://github.com/chopratejas/headroom/blob/main/headroom/storage/sqlite.py).

5. If the LLM requires the original data, the **proxy** in [`headroom/proxy/server.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/server.py) looks up the original in the cache and returns it.

6. **Telemetry** modules record statistics, while [`headroom/shared_context.py`](https://github.com/chopratejas/headroom/blob/main/headroom/shared_context.py) enables multiple agents to share compressed contexts.

## Code Examples

### Using the SDK Inline

```python
from headroom import compress

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain how Headroom reduces token usage."},
]

# Compress the conversation before sending it to an LLM

compressed, stats = compress(messages, model="gpt-4o")
print("Compressed payload:", compressed)
print("Saved tokens:", stats.input_tokens_saved)

```

This example imports the `compress` function from [`headroom/__init__.py`](https://github.com/chopratejas/headroom/blob/main/headroom/__init__.py), which delegates to the transform pipeline.

### Running the Proxy

```bash

# Start a drop-in proxy on port 8787

headroom proxy --port 8787

```

This command loads [`headroom/proxy/server.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/server.py) and wires together the same pipeline used by the SDK, allowing any downstream client (Anthropic SDK, OpenAI SDK, LangChain) to automatically benefit from compression without code changes.

## Summary

- **[`headroom/__init__.py`](https://github.com/chopratejas/headroom/blob/main/headroom/__init__.py)** exposes the public `compress` API and top-level objects.
- **[`headroom/cli.py`](https://github.com/chopratejas/headroom/blob/main/headroom/cli.py)** provides the command-line entry point for `wrap`, `proxy`, and `learn` modes.
- **[`headroom/transforms/pipeline.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/pipeline.py)** orchestrates the compression stages, while [`content_router.py`](https://github.com/chopratejas/headroom/blob/main/content_router.py) selects the appropriate algorithm.
- **[`headroom/transforms/smart_crusher.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/smart_crusher.py)**, [`code_compressor.py`](https://github.com/chopratejas/headroom/blob/main/code_compressor.py), and [`kompress_compressor.py`](https://github.com/chopratejas/headroom/blob/main/kompress_compressor.py) implement the actual compression algorithms for JSON, code, and prose respectively.
- **[`headroom/proxy/server.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/server.py)** enables drop-in proxy deployment without application modifications.
- **[`headroom/storage/sqlite.py`](https://github.com/chopratejas/headroom/blob/main/headroom/storage/sqlite.py)** and [`shared_context.py`](https://github.com/chopratejas/headroom/blob/main/shared_context.py) provide persistent caching and cross-agent memory.
- **[`headroom/utils.py`](https://github.com/chopratejas/headroom/blob/main/headroom/utils.py)** contains critical helpers like `compute_prefix_hash` and marker generation.

## Frequently Asked Questions

### Where is the main entry point for the Headroom SDK?

The main entry point is **[`headroom/__init__.py`](https://github.com/chopratejas/headroom/blob/main/headroom/__init__.py)**, which exports the `compress` function and other public objects. When you write `from headroom import compress`, you are importing from this file, which then delegates to the transform pipeline.

### How does Headroom decide which compression algorithm to use?

The **[`headroom/transforms/content_router.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/content_router.py)** module contains the `ContentRouter` class that analyzes the content type and structure. It routes JSON data to [`smart_crusher.py`](https://github.com/chopratejas/headroom/blob/main/smart_crusher.py), source code to [`code_compressor.py`](https://github.com/chopratejas/headroom/blob/main/code_compressor.py), and prose text to [`kompress_compressor.py`](https://github.com/chopratejas/headroom/blob/main/kompress_compressor.py).

### What enables the proxy mode to work without code changes?

The **[`headroom/proxy/server.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/server.py)** module implements an HTTP server that intercepts requests to standard LLM API endpoints. Combined with [`headroom/proxy/runtime_env.py`](https://github.com/chopratejas/headroom/blob/main/headroom/proxy/runtime_env.py) for environment management, this allows existing applications to point their base URL to the local proxy, automatically triggering compression via the same pipeline used by the SDK.

### How does Headroom store data for reversible compression?

The system uses **[`headroom/storage/sqlite.py`](https://github.com/chopratejas/headroom/blob/main/headroom/storage/sqlite.py)** to maintain a SQLite-backed cache of original payloads. When compression occurs, [`headroom/utils.py`](https://github.com/chopratejas/headroom/blob/main/headroom/utils.py) computes a prefix hash that serves as the lookup key. If the LLM later needs the original data, the proxy retrieves it from this cache, enabling full round-trip compression and decompression.