# How to Measure Compression Accuracy and What Benchmarks Are Preserved in Headroom

> Learn how Headroom measures compression accuracy using compression_ratio and validates benchmark preservation across GSM8K, TruthfulQA, SQuAD v2, and BFCL datasets. Understand LLM performance post-compression.

- Repository: [Tejas Chopra/headroom](https://github.com/chopratejas/headroom)
- Tags: benchmarks
- Published: 2026-06-18

---

**Headroom measures compression accuracy through a `compression_ratio` attribute calculated as `1 - (compressed_len / original_len)` and validates benchmark preservation across GSM8K, TruthfulQA, SQuAD v2, and BFCL datasets using a tiered evaluation suite that compares LLM performance before and after compression.**

The `chopratejas/headroom` repository implements CCR (Contextual Compression and Reconstruction) to reduce token counts in LLM prompts while maintaining full recoverability of the original text. Understanding how to quantify compression savings and verify that downstream task accuracy remains intact is essential for production deployments.

## How Headroom Calculates Compression Accuracy

### The Compression Ratio Formula

Every compression operation in Headroom returns a result object containing a **`compression_ratio`** field. The ratio quantifies the percentage of token reduction achieved by the compressor:

```text
compression_ratio = 1 - (compressed_len / original_len)

```

Here, `original_len` represents the token count (or character count, depending on the specific compressor) before compression, while `compressed_len` is the size after the transform completes. A ratio of `0.0` indicates no reduction occurred, whereas `0.9` signifies a 90% size reduction. Because Headroom’s CCR architecture preserves the original content for reconstruction, this ratio measures pure information density gain without lossy truncation.

### Result Objects in Transform Modules

The `compression_ratio` value is populated in the generic result handling code used by all compressors. When you invoke the pipeline, the result object is constructed in transform modules such as **[`headroom/transforms/kompress_compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/kompress_compressor.py)** and **[`headroom/transforms/code_compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/code_compressor.py)**. These modules calculate the ratio immediately after compression and attach it to the returned result object, making the metric available for logging or downstream decision-making.

## Benchmark Preservation and Accuracy Verification

Headroom ships with a comprehensive benchmark suite under the **`benchmarks/`** directory and the **`headroom/evals`** package. The suite runs the same LLM prompts used in standard academic datasets and compares answers generated from uncompressed text versus text that underwent compression and subsequent reconstruction.

### Tier-1 Benchmark Datasets

The evaluation suite specifically preserves accuracy across four primary benchmarks. According to the README and evaluation harness in **[`headroom/evals/comprehensive_benchmark.py`](https://github.com/chopratejas/headroom/blob/main/headroom/evals/comprehensive_benchmark.py)**, the following results are maintained:

- **GSM8K** — Math questions (100 examples): Baseline accuracy 0.870 versus compressed 0.870 (Δ ±0.000)
- **TruthfulQA** — Factual Q&A (100 examples): Baseline 0.530 versus compressed 0.560 (Δ +0.030)
- **SQuAD v2** — Extractive QA (100 examples): 97% compression achieved with no loss of answer quality
- **BFCL** — Tool-use tasks (100 examples): 97% compression with no measurable degradation in function-calling accuracy

### Measuring Accuracy Delta with CCR

The benchmark harness measures **accuracy delta** by invoking the original dataset evaluation logic (such as the GSM8K grader) on both the uncompressed output and the compressed-then-recovered output. Because the CCR layer restores the original text exactly on demand, the LLM’s answer accuracy remains identical to the uncompressed baseline, which explains why delta values are zero or marginally positive (as seen in TruthfulQA’s slight improvement).

## Running Evaluations and Inspecting Results

You can execute the full benchmark suite from the command line to reproduce the preserved accuracy metrics:

```bash

# Run the full tier-1 benchmark suite (covers the four datasets above)

python -m headroom.evals suite --tier 1

```

To access compression metrics programmatically after compressing a single message:

```python
from headroom import compress

messages = [{"role": "user", "content": "Explain the quicksort algorithm in detail."}]
result = compress(messages)                # runs the full pipeline

print(f"Compression ratio: {result.compression_ratio:.2%}")

# → Compression ratio: 78.4% (example)

```

For programmatic benchmark execution and reporting:

```python
from headroom.evals import suite_runner

# Equivalent to CLI: python -m headroom.evals suite --tier 1

report = suite_runner.run(tier=1)

print("Overall token savings:", f"{report.avg_compression_ratio:.2%}")
print("GSM8K accuracy delta:", report.metrics["gsm8k"].delta)
print("TruthfulQA accuracy delta:", report.metrics["truthfulqa"].delta)

```

To inspect individual benchmark results or run specific evaluators like the compression-only check:

```python
from headroom.evals import compression_only

bench = compression_only.run()   # runs only the compression-only evaluator

for entry in bench.entries:
    print(entry.dataset, entry.original_score, entry.compressed_score,
          f"ratio={entry.compression_ratio:.2%}")

```

## Summary

- **Compression accuracy** is quantified via the `compression_ratio` field available on all result objects, calculated as `1 - (compressed_len / original_len)`.
- **Benchmark preservation** is verified across GSM8K, TruthfulQA, SQuAD v2, and BFCL, with accuracy deltas measured by comparing LLM outputs before and after CCR reconstruction.
- **Key source files** include [`headroom/transforms/kompress_compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/kompress_compressor.py), [`headroom/transforms/code_compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/code_compressor.py), and the evaluation harness in [`headroom/evals/comprehensive_benchmark.py`](https://github.com/chopratejas/headroom/blob/main/headroom/evals/comprehensive_benchmark.py).
- **Evaluation methods** include both CLI execution (`python -m headroom.evals suite --tier 1`) and Python API access via `suite_runner` and `compression_only` modules.

## Frequently Asked Questions

### What is the exact formula for calculating compression ratio in Headroom?

The compression ratio is calculated as `compression_ratio = 1 - (compressed_len / original_len)`, where `original_len` is the token or character count before compression and `compressed_len` is the size after the compressor runs. This value is stored in the result object returned by compression operations in modules like [`headroom/transforms/kompress_compressor.py`](https://github.com/chopratejas/headroom/blob/main/headroom/transforms/kompress_compressor.py).

### Which benchmarks does Headroom preserve without accuracy loss?

Headroom preserves accuracy across **GSM8K** (math reasoning), **TruthfulQA** (factual QA), **SQuAD v2** (extractive question answering), and **BFCL** (tool-use tasks). Specifically, GSM8K shows zero accuracy delta (0.870 vs 0.870), while SQuAD v2 and BFCL both achieve 97% compression with no measurable degradation in task performance.

### How does Headroom ensure no degradation in LLM answer quality after compression?

Headroom uses a CCR (Contextual Compression and Reconstruction) layer that compresses text for transmission but restores the exact original text before presenting it to the LLM. The benchmark suite verifies this by running the original evaluation logic (e.g., GSM8K graders) on both uncompressed and reconstructed outputs, ensuring the accuracy delta remains zero or positive.

### Can I run individual benchmark datasets instead of the full tier-1 suite?

Yes. While the command `python -m headroom.evals suite --tier 1` runs all four datasets, you can import specific evaluators from `headroom.evals` such as `compression_only` to run targeted evaluations. The `suite_runner` module also supports configuration options to select specific datasets when running programmatically.