How to Use Context Compression in Agno to Optimize Token Usage
Agno automatically compresses verbose tool results before adding them to the LLM context using CompressionManager, reducing token consumption and API costs while preserving essential information for reasoning.
Context compression in Agno is a built-in optimization that shrinks oversized tool outputs before they reach the language model. The feature centers on the CompressionManager class and integrates seamlessly into the Agent, Team, and Model pipelines. This guide explains how to enable, configure, and monitor context compression using the actual implementation from the agno-agi/agno repository.
How Context Compression Works in Agno
The compression pipeline hooks into every model execution cycle. When an Agno model produces a response via Model._run in libs/agno/agno/models/base.py, the system checks whether tool results should be compressed before the next LLM call.
The Compression Decision
At lines 662–665 (synchronous) and 874–877 (asynchronous) in base.py, the model consults the active CompressionManager:
if _compression_manager is not None and _compression_manager.should_compress(messages):
_compression_manager.compress(messages)
The should_compress method evaluates two criteria:
compress_tool_results_limit– triggers compression when the count of un-compressed tool messages exceeds the threshold (default:3).compress_token_limit– optional hard ceiling that forces compression when token count exceeds the specified value.
Compression Execution
When triggered, CompressionManager.compress (lines 28–53 in libs/agno/agno/compression/manager.py) iterates over tool-role messages lacking compressed_content. It constructs a compression prompt—either the user-provided compress_tool_call_instructions or the built-in DEFAULT_COMPRESSION_PROMPT—and calls the configured LLM to generate a condensed summary.
The result is stored in msg.compressed_content while the original msg.content remains intact for auditability. Downstream token counting automatically uses the compressed version.
Enabling and Configuring Context Compression
Default Behavior
Both Agent and Team constructors automatically instantiate a CompressionManager when compress_tool_results=True (the default). Compression activates automatically once un-compressed tool messages exceed the default limit of 3.
Custom Configuration
For fine-grained control, instantiate CompressionManager explicitly and pass it to your Agent or Team:
from agno.compression.manager import CompressionManager
from agno.agent import Agent
compression_mgr = CompressionManager(
compress_tool_results=True,
compress_tool_results_limit=1, # Compress after every tool call
compress_token_limit=2000, # Force compression above 2000 tokens
)
agent = Agent(
model="gpt-4o-mini",
tools=[search_tool],
compression_manager=compression_mgr,
)
The same pattern applies to Team configurations in libs/agno/agno/team/_init.py (lines 548–560):
from agno.team import Team
from agno.compression.manager import CompressionManager
team = Team(
members=[agent_a, agent_b],
compress_tool_results=True,
compression_manager=CompressionManager(compress_tool_results_limit=2),
)
Custom Compression Prompts
Override the default summarization behavior by providing domain-specific instructions:
custom_prompt = """
You are a financial analyst. Summarize the numeric data, keep all percentages, dates and ticker symbols.
Remove any narrative fluff.
"""
compression = CompressionManager(
compress_tool_call_instructions=custom_prompt,
compress_tool_results_limit=1,
)
Accessing Compression Statistics
After execution, inspect RunResponse.compression_stats to verify token savings:
response = agent.run("Summarize the latest research on quantum computing")
if response.compression_stats:
stats = response.compression_stats
print(f"Tool results compressed: {stats['tool_results_compressed']}")
print(f"Original size: {stats['original_size']} tokens")
print(f"Compressed size: {stats['compressed_size']} tokens")
print(f"Tokens saved: {stats['original_size'] - stats['compressed_size']}")
The built-in CLI and HTML renderers in libs/agno/agno/utils/print_response/ automatically display these statistics after each run (see team.py lines 280–287 and the corresponding agent utility).
Summary
- Context compression in Agno automatically summarizes verbose tool outputs before they enter the LLM context, reducing token usage and API costs.
- The
CompressionManagerclass inlibs/agno/agno/compression/manager.pyorchestrates compression based on message count thresholds (compress_tool_results_limit) or token limits (compress_token_limit). - Compression triggers inside
Model._run(libs/agno/agno/models/base.py) at lines 662–665 (sync) and 874–877 (async), ensuring every tool result is evaluated before the next LLM call. - Configure compression by passing a custom
CompressionManagertoAgentorTeam, or rely on the default settings (compress_tool_results=True, limit of3). - Access runtime statistics via
response.compression_statsto measure original versus compressed token counts.
Frequently Asked Questions
What is the default compression threshold in Agno?
By default, Agno compresses tool results when the number of un-compressed tool messages exceeds 3. This is controlled by the compress_tool_results_limit parameter in CompressionManager, which defaults to 3 when compress_tool_results is enabled.
Can I use a custom LLM prompt for compression?
Yes. Pass a custom string to the compress_tool_call_instructions parameter when instantiating CompressionManager. This prompt overrides the built-in DEFAULT_COMPRESSION_PROMPT and allows domain-specific summarization rules—for example, preserving financial metrics while removing narrative text.
How do I disable context compression entirely?
Set compress_tool_results=False in your Agent or Team constructor. This prevents the automatic instantiation of CompressionManager and ensures all tool outputs are passed to the LLM in their original form, regardless of length.
Does compression work with asynchronous agents?
Yes. The compression check runs in both synchronous and asynchronous execution paths. In libs/agno/agno/models/base.py, the asynchronous run method includes the same compression logic at lines 874–877 as the synchronous path at lines 662–665, ensuring consistent behavior across async agents and teams.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →