Optimizing Token Usage for AI Agent Context Management: A Complete Guide Based on ai-agent-book
AI agents can reduce token consumption by up to 76% through context compression, budget-aware truncation, and policy-based gating—without sacrificing response quality. The ai-agent-book repository by bojieli demonstrates a production-ready architecture that separates measurement, compression, and enforcement into modular, testable components.
Built for researchers and engineers building scalable LLM systems, this codebase shows how to keep agents within budget while preserving the most semantically valuable context. Whether you're running cost-sensitive APIs or managing constrained edge deployments, the patterns in bojieli/ai-agent-book provide a battle-tested foundation.
Why Token Optimization Matters for AI Agents
Every LLM call incurs costs proportional to token count. For autonomous agents that may chain dozens of tool calls across extended sessions, unbounded context growth leads to:
- Runaway API costs from ever-expanding dialogue history
- Latency degradation as prompts exceed model processing windows
- Truncation errors when prompts hit hard provider limits
The ai-agent-book solution treats token usage as a first-class resource to be measured, compressed, and governed—rather than an afterthought.
Four Core Mechanisms for Token-Aware Context Management
The repository implements four interconnected systems. Each operates independently but composes into a unified pipeline.
Token Accounting with compute_token_usage
At tests/test_ch8_evaluate_multilingual.py, the codebase defines a lightweight TokenUsageObj pattern that serializes consumption across prompt, completion, reasoning, and total token categories.
from utils.token_usage import compute_token_usage
prompt = "Summarize the following article in two sentences."
completion = "The article discusses..."
token_counts = {"prompt_tokens": 12, "completion_tokens": 24, "total_tokens": 36}
usage = compute_token_usage(
prompt_text=prompt,
reasoning_text="I considered the key points…",
answer_text=completion,
model_output=token_counts,
)
# Result: {'prompt_tokens': 12, 'completion_tokens': 24,
# 'reasoning_tokens': 5, 'total_tokens': 41}
This accounting layer enables per-call tracking and aggregate budget monitoring. The helper aggregates counts across an agent's full execution trace, making it possible to attribute costs to specific tool invocations or dialogue branches.
Safety-Policy Gating with SafetyPolicyGate
Hard limits prevent budget overruns. In tests/test_ch9_safety_policy_gate.py, the SafetyPolicyGate class enforces max_tokens, max_timeout, and max_file_bytes thresholds:
from safety_gate import SafetyPolicyGate
gate = SafetyPolicyGate(max_tokens=10_000, token_ttl=60.0)
# Within budget—proceeds immediately
decision = gate.validate_tool_call("generate_text", {"max_tokens": 4_000})
assert decision.confirmation_token is None
# Would exceed budget—issues single-use confirmation token
decision = gate.validate_tool_call("generate_text", {"max_tokens": 7_000})
token = decision.confirmation_token
# Must re-validate with token before execution
decision = gate.validate_tool_call(
"generate_text",
{"max_tokens": 7_000},
confirm_token=token,
)
Key design decisions in this implementation:
- Single-use confirmation tokens prevent replay attacks on budget overrides
- Configurable TTL (
token_ttl=60.0) limits the window for explicit approval - Hard return of
Decisionobjects forces callers to handle denial cases explicitly
The production implementation lives at chapter9/harness-safety-gate/llm_generator.py, used in the GAIA benchmark experiments.
Context Compression with compress_key_sentence and compress_observation_filtering
Before tokens are counted or budget-checked, raw context can be semantically compressed. The benchmark utilities in chapter2/context-compression/benchmark.py provide two strategies:
from benchmark import compress_key_sentence, compress_observation_filtering
raw_context = """
The user asked about token optimisation.
The system replied with a long explanation.
The user then inquired about safety policies.
...
"""
# Strategy 1: Extract query-relevant sentence
key = compress_key_sentence(raw_context, query="token optimisation")
# Strategy 2: Filter redundant observations
compressed = compress_observation_filtering(key)
# Result: "The user asked about token optimisation."
The scripts/sync_chapter2_figures.py visualization shows 76% token reduction versus uncompressed baselines, with minimal impact on downstream task accuracy. Compression runs before tokenization, so savings compound across the full pipeline.
Dialogue Truncation with InterruptionManager
When compression alone cannot fit the budget, InterruptionManager (demonstrated in tests/test_ch9_interruption_manager.py) implements oldest-first eviction:
# Inside safe_llm_call workflow
ctx = InterruptionManager().get_dialogue_context()
# Returns trimmed list when token budget exceeded
This prioritizes recency—the assumption being that recent turns carry more predictive value for the agent's next action than distant history.
Assembling the Full Pipeline
These four mechanisms compose into a complete token-aware request flow:
def safe_llm_call(prompt, gate, tokenizer, query):
# 1. Retrieve and truncate dialogue if needed
ctx = InterruptionManager().get_dialogue_context()
# 2. Compress remaining context semantically
compact = compress_key_sentence(ctx, query)
# 3. Count tokens for budget check
usage = compute_token_usage(
prompt, compact, "",
model_output=tokenizer(compact)
)
# 4. Enforce policy; obtain confirmation if over budget
headroom = 500 # reserve tokens for completion
decision = gate.validate_tool_call(
"generate_text",
{"max_tokens": usage["total_tokens"] + headroom},
)
if decision.confirmation_token:
decision = gate.validate_tool_call(
"generate_text",
{"max_tokens": usage["total_tokens"] + headroom},
confirm_token=decision.confirmation_token,
)
return decision # approved output or denial
This pattern appears throughout the repository's evaluation harnesses, adapted for multilingual benchmarks (test_ch8_evaluate_multilingual.py) and safety-critical pathways (test_ch9_safety_policy_gate.py).
Performance Trade-offs and Benchmarking Results
The chapter2/context-compression/benchmark.py module provides empirical guidance on compression strategy selection:
| Strategy | Token Reduction | Use Case |
|---|---|---|
compress_key_sentence |
~60-80% | Query-focused retrieval, Q&A agents |
compress_observation_filtering |
~40-50% | Multi-step tool use with redundant observations |
| Combined pipeline | ~76% | Production agents with variable context |
Key insight from scripts/sync_chapter2_figures.py: compression before truncation preserves more semantic information than truncation alone, because truncation operates on raw token sequences without awareness of content relevance.
Key Source Files for Token Optimization
| File | Purpose |
|---|---|
tests/test_ch8_evaluate_multilingual.py |
Token usage objects and compute_token_usage |
tests/test_ch9_safety_policy_gate.py |
SafetyPolicyGate with budget enforcement |
tests/test_ch9_interruption_manager.py |
Dialogue truncation on budget overflow |
chapter2/context-compression/benchmark.py |
Compression algorithms (compress_key_sentence, compress_observation_filtering) |
chapter9/harness-safety-gate/llm_generator.py |
Production SafetyPolicyGate implementation |
scripts/sync_chapter2_figures.py |
Visualization of token savings benchmarks |
Summary
- Measure first:
compute_token_usageintest_ch8_evaluate_multilingual.pyprovides granular, serializable token accounting across prompt, reasoning, and completion categories - Compress early:
chapter2/context-compression/benchmark.pydemonstrates 76% token reduction through semantic compression before counting - Truncate smartly:
InterruptionManagerevicts oldest turns when budgets force hard limits - Enforce strictly:
SafetyPolicyGateintest_ch9_safety_policy_gate.pyissues single-use confirmation tokens for override workflows, with configurable TTL
Frequently Asked Questions
What is the most effective way to reduce token usage in AI agents?
Semantic context compression before tokenization. The ai-agent-book benchmark shows that compress_key_sentence and compress_observation_filtering achieve 76% reduction versus uncompressed baselines. This outperforms naive truncation because it preserves query-relevant information while eliminating redundancy.
How does SafetyPolicyGate prevent token budget overruns?
By validating every tool call against max_tokens before execution. In tests/test_ch9_safety_policy_gate.py, the gate returns a Decision object—either approving immediately or requiring a single-use confirmation_token when the request would exceed quota. The token expires after token_ttl seconds, forcing explicit human or system approval for over-budget operations.
When should I use dialogue truncation versus context compression?
Compression first, truncation as fallback. chapter2/context-compression/benchmark.py strategies preserve semantic meaning by selecting informative content. Only when compressed context still exceeds max_tokens should InterruptionManager (tests/test_ch9_interruption_manager.py) discard oldest turns. This two-tier approach maximizes information density within budget.
Where is the production-ready implementation of these token controls?
The GAIA benchmark harness at chapter9/harness-safety-gate/llm_generator.py contains the hardened SafetyPolicyGate with full error handling and logging. Test suites in tests/test_ch9_*.py demonstrate integration patterns, while chapter2/context-compression/benchmark.py provides research-grade compression utilities suitable for adaptation to production tokenizers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →