How to Measure Compression Accuracy and What Benchmarks Are Preserved in Headroom
Headroom measures compression accuracy through a compression_ratio attribute calculated as 1 - (compressed_len / original_len) and validates benchmark preservation across GSM8K, TruthfulQA, SQuAD v2, and BFCL datasets using a tiered evaluation suite that compares LLM performance before and after compression.
The chopratejas/headroom repository implements CCR (Contextual Compression and Reconstruction) to reduce token counts in LLM prompts while maintaining full recoverability of the original text. Understanding how to quantify compression savings and verify that downstream task accuracy remains intact is essential for production deployments.
How Headroom Calculates Compression Accuracy
The Compression Ratio Formula
Every compression operation in Headroom returns a result object containing a compression_ratio field. The ratio quantifies the percentage of token reduction achieved by the compressor:
compression_ratio = 1 - (compressed_len / original_len)
Here, original_len represents the token count (or character count, depending on the specific compressor) before compression, while compressed_len is the size after the transform completes. A ratio of 0.0 indicates no reduction occurred, whereas 0.9 signifies a 90% size reduction. Because Headroom’s CCR architecture preserves the original content for reconstruction, this ratio measures pure information density gain without lossy truncation.
Result Objects in Transform Modules
The compression_ratio value is populated in the generic result handling code used by all compressors. When you invoke the pipeline, the result object is constructed in transform modules such as headroom/transforms/kompress_compressor.py and headroom/transforms/code_compressor.py. These modules calculate the ratio immediately after compression and attach it to the returned result object, making the metric available for logging or downstream decision-making.
Benchmark Preservation and Accuracy Verification
Headroom ships with a comprehensive benchmark suite under the benchmarks/ directory and the headroom/evals package. The suite runs the same LLM prompts used in standard academic datasets and compares answers generated from uncompressed text versus text that underwent compression and subsequent reconstruction.
Tier-1 Benchmark Datasets
The evaluation suite specifically preserves accuracy across four primary benchmarks. According to the README and evaluation harness in headroom/evals/comprehensive_benchmark.py, the following results are maintained:
- GSM8K — Math questions (100 examples): Baseline accuracy 0.870 versus compressed 0.870 (Δ ±0.000)
- TruthfulQA — Factual Q&A (100 examples): Baseline 0.530 versus compressed 0.560 (Δ +0.030)
- SQuAD v2 — Extractive QA (100 examples): 97% compression achieved with no loss of answer quality
- BFCL — Tool-use tasks (100 examples): 97% compression with no measurable degradation in function-calling accuracy
Measuring Accuracy Delta with CCR
The benchmark harness measures accuracy delta by invoking the original dataset evaluation logic (such as the GSM8K grader) on both the uncompressed output and the compressed-then-recovered output. Because the CCR layer restores the original text exactly on demand, the LLM’s answer accuracy remains identical to the uncompressed baseline, which explains why delta values are zero or marginally positive (as seen in TruthfulQA’s slight improvement).
Running Evaluations and Inspecting Results
You can execute the full benchmark suite from the command line to reproduce the preserved accuracy metrics:
# Run the full tier-1 benchmark suite (covers the four datasets above)
python -m headroom.evals suite --tier 1
To access compression metrics programmatically after compressing a single message:
from headroom import compress
messages = [{"role": "user", "content": "Explain the quicksort algorithm in detail."}]
result = compress(messages) # runs the full pipeline
print(f"Compression ratio: {result.compression_ratio:.2%}")
# → Compression ratio: 78.4% (example)
For programmatic benchmark execution and reporting:
from headroom.evals import suite_runner
# Equivalent to CLI: python -m headroom.evals suite --tier 1
report = suite_runner.run(tier=1)
print("Overall token savings:", f"{report.avg_compression_ratio:.2%}")
print("GSM8K accuracy delta:", report.metrics["gsm8k"].delta)
print("TruthfulQA accuracy delta:", report.metrics["truthfulqa"].delta)
To inspect individual benchmark results or run specific evaluators like the compression-only check:
from headroom.evals import compression_only
bench = compression_only.run() # runs only the compression-only evaluator
for entry in bench.entries:
print(entry.dataset, entry.original_score, entry.compressed_score,
f"ratio={entry.compression_ratio:.2%}")
Summary
- Compression accuracy is quantified via the
compression_ratiofield available on all result objects, calculated as1 - (compressed_len / original_len). - Benchmark preservation is verified across GSM8K, TruthfulQA, SQuAD v2, and BFCL, with accuracy deltas measured by comparing LLM outputs before and after CCR reconstruction.
- Key source files include
headroom/transforms/kompress_compressor.py,headroom/transforms/code_compressor.py, and the evaluation harness inheadroom/evals/comprehensive_benchmark.py. - Evaluation methods include both CLI execution (
python -m headroom.evals suite --tier 1) and Python API access viasuite_runnerandcompression_onlymodules.
Frequently Asked Questions
What is the exact formula for calculating compression ratio in Headroom?
The compression ratio is calculated as compression_ratio = 1 - (compressed_len / original_len), where original_len is the token or character count before compression and compressed_len is the size after the compressor runs. This value is stored in the result object returned by compression operations in modules like headroom/transforms/kompress_compressor.py.
Which benchmarks does Headroom preserve without accuracy loss?
Headroom preserves accuracy across GSM8K (math reasoning), TruthfulQA (factual QA), SQuAD v2 (extractive question answering), and BFCL (tool-use tasks). Specifically, GSM8K shows zero accuracy delta (0.870 vs 0.870), while SQuAD v2 and BFCL both achieve 97% compression with no measurable degradation in task performance.
How does Headroom ensure no degradation in LLM answer quality after compression?
Headroom uses a CCR (Contextual Compression and Reconstruction) layer that compresses text for transmission but restores the exact original text before presenting it to the LLM. The benchmark suite verifies this by running the original evaluation logic (e.g., GSM8K graders) on both uncompressed and reconstructed outputs, ensuring the accuracy delta remains zero or positive.
Can I run individual benchmark datasets instead of the full tier-1 suite?
Yes. While the command python -m headroom.evals suite --tier 1 runs all four datasets, you can import specific evaluators from headroom.evals such as compression_only to run targeted evaluations. The suite_runner module also supports configuration options to select specific datasets when running programmatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →