Context Compression in LLM Agents: When and How to Apply It
Context compression keeps LLM agent conversations within token limits by automatically truncating or summarizing messages when configurable thresholds are exceeded.
Context compression is essential for building reliable LLM agents that handle long-running conversations without hitting model context windows. In the bojieli/ai-agent-book repository, the AWorld agent framework implements a sophisticated compression pipeline that balances token efficiency with information preservation. This article examines the technical implementation and provides concrete guidance on when to enable different compression strategies.
How Context Compression Works in LLM Agents
The AWorld framework approaches context compression through a multi-layered system that evaluates token budgets, detects overflow conditions, and applies the appropriate compression algorithm.
Core Components
| Component | Purpose | Source Location |
|---|---|---|
LlmCompressionConfig |
Configuration for enabling compression, selecting algorithm type, and setting trigger thresholds | aworld/config/conf.py (line 135) |
PromptProcessor |
Orchestrates detection and execution of compression strategies | aworld/core/context/processor/prompt_processor.py |
TruncateCompressor |
Fast, token-aware truncation fallback | aworld/core/context/processor/truncate_compressor.py (lines 18-66) |
LLMCompressor / LLMLinguaCompressor |
LLM-based or library-based summarization | Imported in prompt_processor.py (lines 12-14) |
The Compression Decision Pipeline
The PromptProcessor class implements a clear decision hierarchy in prompt_processor.py (lines 26-62):
- Calculate token budget —
get_max_tokens()multipliesmodel_config.max_model_lenbycontext_rule.optimization_config.max_token_budget_ratio - Detect overflow —
is_out_of_context()compares current tokens against this budget - Determine strategy —
decide_compression_strategy()checks if the chunk exceedstrigger_compress_token_length - Execute compression — based on
compress_type, invokeLLMLinguaCompressor,LLMCompressor, or fall back toTruncateCompressor
The result is a CompressionResult containing compressed content, compression ratio, and method metadata.
Configuring Context Compression
Enabling LLM-Based Compression
The LlmCompressionConfig class in conf.py controls all compression behavior:
from aworld.config.conf import AgentConfig, ModelConfig, LlmCompressionConfig, ContextRuleConfig, OptimizationConfig
# Define the model used for compression (can be smaller/cheaper than main LLM)
compress_model = ModelConfig(
llm_model_name="gpt-4o-mini",
llm_provider="openai",
max_model_len=8192,
)
# Build agent configuration with compression enabled
agent_cfg = AgentConfig(
llm_config=ModelConfig(llm_model_name="gpt-4o", llm_provider="openai"),
context_rule=ContextRuleConfig(
llm_compression_config=LlmCompressionConfig(
enabled=True,
compress_type="llm", # Options: "llm", "llmlingua"
trigger_compress_token_length=8000,
compress_model=compress_model,
),
optimization_config=OptimizationConfig(
enabled=True,
max_token_budget_ratio=0.5, # Use 50% of model's max context
),
),
)
Direct Processor Usage
For fine-grained control, instantiate PromptProcessor directly:
from aworld.core.context.processor.prompt_processor import PromptProcessor
processor = PromptProcessor(context_rule=ctx_rule, model_config=model_cfg)
# Check if conversation exceeds budget
if processor.is_out_of_context(messages, is_last_message_in_memory=False):
if processor.should_compress_conversation(messages):
compressed = processor.compress_pipeline.compress_messages(messages)
messages = eval(compressed.compressed_content)
Compression Algorithms: Three Approaches
TruncateCompressor: Fast and Deterministic
The TruncateCompressor in truncate_compressor.py (lines 18-66) provides a reliable fallback when summarization isn't available. It performs token-aware truncation—preserving recent messages while dropping older ones to meet the token limit.
When to use: Resource-constrained environments, latency-sensitive applications, or when deterministic behavior is required.
LLMCompressor: Intelligent Summarization
When compress_type="llm", the framework invokes an LLM to generate condensed versions of conversation history. This preserves semantic meaning better than truncation but adds compute overhead.
When to use: Complex dialogues where context relationships matter more than exact phrasing.
LLMLinguaCompressor: Specialized Library Integration
Setting compress_type="llmlingua" enables integration with the LLMLingua library—optimized specifically for prompt compression with minimal semantic loss.
When to use: Production deployments requiring efficient compression without maintaining a separate compression model.
When to Apply Context Compression
| Scenario | Recommended Configuration | Rationale |
|---|---|---|
| Long-running multi-turn conversations | enabled=True, trigger_compress_token_length=5000 |
Prevents token exhaustion as turn count grows |
| Large tool outputs (JSON, logs, code) | Enable should_compress_tool_result with moderate threshold |
Tool results often dominate context window |
| Models with ≤8K token windows | max_token_budget_ratio=0.4, consider TruncateCompressor |
Strict budgeting essential with limited headroom |
| CPU-only or edge deployments | TruncateCompressor fallback |
Eliminates dependency on additional LLM calls |
| Precision-critical domains (legal, medical) | enabled=False or very high threshold |
Avoid information loss from summarization |
Fallback Behavior: Guaranteed Operation
Even when compression is disabled, the system ensures functionality. The TruncateCompressor serves as a hard guarantee:
# Force fallback to truncation
ctx_rule.llm_compression_config.enabled = False
processor = PromptProcessor(context_rule=ctx_rule, model_config=model_cfg)
# Truncation still occurs to prevent context overflow
truncated = processor.truncate_compressor.truncate_messages(
messages,
max_tokens=4000
)
This design ensures agents never fail due to context window exhaustion regardless of configuration.
Key Implementation Files
aworld/config/conf.py— Configuration dataclasses includingLlmCompressionConfigaworld/core/context/processor/prompt_processor.py— Core orchestration logic (decide_compression_strategy,is_out_of_context)aworld/core/context/processor/truncate_compressor.py— Token-aware truncation implementationaworld/core/context/processor/llm_compressor.py— LLM-based summarizationaworld/core/context/processor/llmlingua_compressor.py— LLMLingua integration
Summary
- Context compression maintains LLM agent conversations within token limits through configurable truncation or summarization pipelines
- The
PromptProcessorclass orchestrates detection and execution, withLlmCompressionConfigcontrolling all parameters - Three compression modes exist:
LLMLinguaCompressor(library-based),LLMCompressor(model-based), andTruncateCompressor(deterministic fallback) - Apply compression when handling long dialogues, large tool outputs, or limited context windows—but disable it when exact phrasing is critical
- The
TruncateCompressorguarantee ensures agents function even when primary compressors are disabled
Frequently Asked Questions
What triggers context compression in the AWorld framework?
Compression triggers when the total token count exceeds the product of max_model_len and max_token_budget_ratio, and when an individual chunk surpasses trigger_compress_token_length. The decide_compression_strategy() method in prompt_processor.py (lines 26-38) performs this evaluation before invoking the selected compressor.
How does context compression differ from simple truncation?
Truncation (TruncateCompressor) drops tokens deterministically from the oldest messages to meet budget constraints—fast but potentially lossy. LLM-based compression generates semantic summaries that preserve relationships and key information—slower but more intelligent. The framework supports both with automatic fallback.
Can I use the same model for compression and the main agent task?
Yes, but it's inefficient. The compress_model parameter in LlmCompressionConfig allows specifying a smaller model (e.g., gpt-4o-mini) for compression while reserving the full model for complex reasoning. This reduces cost without significantly impacting summary quality.
What happens if context compression is disabled but tokens still exceed the window?
The TruncateCompressor operates as a safety fallback regardless of enabled status. According to the implementation in prompt_processor.py, the truncation path ensures the agent never sends requests that would cause a model error—though explicit configuration is recommended for predictable behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →