Context Compression Strategies Compared in Experiment 2-IO: A Technical Analysis

Experiment 2-IO evaluates six distinct context compression strategies—including no compression, individual non-context-aware, combined non-context-aware, context-aware, context-aware with citations, and windowed context—to measure their impact on token usage, execution time, success rate, and compression ratio.

The bojieli/ai-agent-book repository implements Experiment 2-IO (Context Compression Strategies Comparison) to benchmark how different approaches handle context window limitations in AI agent systems. Located in chapter2/context-compression/experiment.py, this experiment provides a standardized framework for comparing compression techniques through a unified CLI interface.

The Six Context Compression Strategies

According to the source code in chapter2/context-compression/experiment.py, the experiment defines six compression approaches in the STRATEGY_CHOICES dictionary (lines 25-33). Each strategy maps to a specific CLI alias for flexible testing workflows.

1. No Compression (no_compression)

Strategy Enum: NO_COMPRESSION

This baseline approach retains the full context without applying any compression algorithms. It serves as the control group to measure the performance impact of other strategies against uncompressed transcripts.

2. Individual Non-Context-Aware (individual)

Strategy Enum: NON_CONTEXT_AWARE_INDIVIDUAL

Applies compression on a per-tool basis independently. Each tool's output is compressed in isolation without considering the broader conversation context, potentially losing cross-tool dependencies.

3. Combined Non-Context-Aware (combined)

Strategy Enum: NON_CONTEXT_AWARE_COMBINED

Compresses the entire transcript as a single unit using non-context-aware methods. Unlike the individual approach, this processes the whole context together but still lacks semantic understanding of what information is salient.

4. Context-Aware (context_aware)

Strategy Enum: CONTEXT_AWARE

Intelligently identifies and retains salient information while compressing less relevant content. This strategy uses context-aware algorithms to preserve semantically important tokens that are likely needed for future reasoning steps.

5. Context-Aware with Citations (citations)

Strategy Enum: CONTEXT_AWARE_CITATIONS

Extends the base context-aware approach by additionally preserving citation metadata. This ensures that source references and attribution information remain accessible even after aggressive compression of the main content.

6. Windowed Context (windowed)

Strategy Enum: WINDOWED_CONTEXT

Implements a sliding window approach that retains only a fixed number of recent tokens. Older content beyond the window size is discarded regardless of relevance, mimicking traditional fixed-window attention mechanisms.

Implementation Files

The experiment architecture relies on three core components:

Running the Comparison

The CLI accepts the -s or --strategy flag to select specific approaches using the aliases defined above.

Run the full benchmark comparing all six strategies:

python chapter2/context-compression/experiment.py

Execute a single strategy (context-aware only):

python chapter2/context-compression/experiment.py -s context_aware

Test a subset of strategies (individual and combined):

python chapter2/context-compression/experiment.py -s individual combined

When executed without explicit --strategy arguments, the experiment runs all six strategies defined in STRATEGY_CHOICES by default.

Summary

  • Experiment 2-IO compares six distinct compression strategies ranging from no compression to advanced context-aware methods with citation preservation.
  • Strategies are defined in chapter2/context-compression/experiment.py within the STRATEGY_CHOICES dictionary and implemented as a CompressionStrategy enum in compression_strategies.py.
  • CLI aliases (no_compression, individual, combined, context_aware, citations, windowed) enable flexible testing of individual or multiple strategies.
  • The experiment measures token usage, execution time, success rate, and compression ratio across all approaches.

Frequently Asked Questions

What is the difference between individual and combined non-context-aware compression?

Individual compression processes each tool's output separately without cross-tool awareness, while combined compression treats the entire transcript as a single unit. The individual approach may lose inter-tool dependencies, whereas combined compression maintains the full sequence but applies uniform compression regardless of content importance.

How does context-aware compression differ from windowed context?

Context-aware compression uses semantic understanding to preserve salient information throughout the transcript, potentially keeping older but important tokens. Windowed context simply retains the most recent N tokens regardless of content value, discarding older information automatically when the window slides forward.

Where are the compression strategies defined in the codebase?

The strategies are defined as Python enums in chapter2/context-compression/compression_strategies.py and mapped to CLI aliases in the STRATEGY_CHOICES dictionary located at lines 25-33 of chapter2/context-compression/experiment.py. The configuration settings affecting all strategies reside in chapter2/context-compression/config.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →