Optimizing Token Usage for AI Agent Context Management: A Complete Guide Based on ai-agent-book

AI agents can reduce token consumption by up to 76% through context compression, budget-aware truncation, and policy-based gating—without sacrificing response quality. The ai-agent-book repository by bojieli demonstrates a production-ready architecture that separates measurement, compression, and enforcement into modular, testable components.

Built for researchers and engineers building scalable LLM systems, this codebase shows how to keep agents within budget while preserving the most semantically valuable context. Whether you're running cost-sensitive APIs or managing constrained edge deployments, the patterns in bojieli/ai-agent-book provide a battle-tested foundation.


Why Token Optimization Matters for AI Agents

Every LLM call incurs costs proportional to token count. For autonomous agents that may chain dozens of tool calls across extended sessions, unbounded context growth leads to:

  • Runaway API costs from ever-expanding dialogue history
  • Latency degradation as prompts exceed model processing windows
  • Truncation errors when prompts hit hard provider limits

The ai-agent-book solution treats token usage as a first-class resource to be measured, compressed, and governed—rather than an afterthought.


Four Core Mechanisms for Token-Aware Context Management

The repository implements four interconnected systems. Each operates independently but composes into a unified pipeline.

Token Accounting with compute_token_usage

At tests/test_ch8_evaluate_multilingual.py, the codebase defines a lightweight TokenUsageObj pattern that serializes consumption across prompt, completion, reasoning, and total token categories.

from utils.token_usage import compute_token_usage

prompt = "Summarize the following article in two sentences."
completion = "The article discusses..."
token_counts = {"prompt_tokens": 12, "completion_tokens": 24, "total_tokens": 36}

usage = compute_token_usage(
    prompt_text=prompt,
    reasoning_text="I considered the key points…",
    answer_text=completion,
    model_output=token_counts,
)

# Result: {'prompt_tokens': 12, 'completion_tokens': 24,

#          'reasoning_tokens': 5, 'total_tokens': 41}

This accounting layer enables per-call tracking and aggregate budget monitoring. The helper aggregates counts across an agent's full execution trace, making it possible to attribute costs to specific tool invocations or dialogue branches.

Safety-Policy Gating with SafetyPolicyGate

Hard limits prevent budget overruns. In tests/test_ch9_safety_policy_gate.py, the SafetyPolicyGate class enforces max_tokens, max_timeout, and max_file_bytes thresholds:

from safety_gate import SafetyPolicyGate

gate = SafetyPolicyGate(max_tokens=10_000, token_ttl=60.0)

# Within budget—proceeds immediately

decision = gate.validate_tool_call("generate_text", {"max_tokens": 4_000})
assert decision.confirmation_token is None

# Would exceed budget—issues single-use confirmation token

decision = gate.validate_tool_call("generate_text", {"max_tokens": 7_000})
token = decision.confirmation_token

# Must re-validate with token before execution

decision = gate.validate_tool_call(
    "generate_text",
    {"max_tokens": 7_000},
    confirm_token=token,
)

Key design decisions in this implementation:

  • Single-use confirmation tokens prevent replay attacks on budget overrides
  • Configurable TTL (token_ttl=60.0) limits the window for explicit approval
  • Hard return of Decision objects forces callers to handle denial cases explicitly

The production implementation lives at chapter9/harness-safety-gate/llm_generator.py, used in the GAIA benchmark experiments.

Context Compression with compress_key_sentence and compress_observation_filtering

Before tokens are counted or budget-checked, raw context can be semantically compressed. The benchmark utilities in chapter2/context-compression/benchmark.py provide two strategies:

from benchmark import compress_key_sentence, compress_observation_filtering

raw_context = """
  The user asked about token optimisation.
  The system replied with a long explanation.
  The user then inquired about safety policies.
  ...
"""

# Strategy 1: Extract query-relevant sentence

key = compress_key_sentence(raw_context, query="token optimisation")

# Strategy 2: Filter redundant observations

compressed = compress_observation_filtering(key)

# Result: "The user asked about token optimisation."

The scripts/sync_chapter2_figures.py visualization shows 76% token reduction versus uncompressed baselines, with minimal impact on downstream task accuracy. Compression runs before tokenization, so savings compound across the full pipeline.

Dialogue Truncation with InterruptionManager

When compression alone cannot fit the budget, InterruptionManager (demonstrated in tests/test_ch9_interruption_manager.py) implements oldest-first eviction:


# Inside safe_llm_call workflow

ctx = InterruptionManager().get_dialogue_context()

# Returns trimmed list when token budget exceeded

This prioritizes recency—the assumption being that recent turns carry more predictive value for the agent's next action than distant history.


Assembling the Full Pipeline

These four mechanisms compose into a complete token-aware request flow:

def safe_llm_call(prompt, gate, tokenizer, query):
    # 1. Retrieve and truncate dialogue if needed

    ctx = InterruptionManager().get_dialogue_context()
    
    # 2. Compress remaining context semantically

    compact = compress_key_sentence(ctx, query)
    
    # 3. Count tokens for budget check

    usage = compute_token_usage(
        prompt, compact, "", 
        model_output=tokenizer(compact)
    )
    
    # 4. Enforce policy; obtain confirmation if over budget

    headroom = 500  # reserve tokens for completion

    decision = gate.validate_tool_call(
        "generate_text",
        {"max_tokens": usage["total_tokens"] + headroom},
    )
    
    if decision.confirmation_token:
        decision = gate.validate_tool_call(
            "generate_text",
            {"max_tokens": usage["total_tokens"] + headroom},
            confirm_token=decision.confirmation_token,
        )
    
    return decision  # approved output or denial

This pattern appears throughout the repository's evaluation harnesses, adapted for multilingual benchmarks (test_ch8_evaluate_multilingual.py) and safety-critical pathways (test_ch9_safety_policy_gate.py).


Performance Trade-offs and Benchmarking Results

The chapter2/context-compression/benchmark.py module provides empirical guidance on compression strategy selection:

Strategy Token Reduction Use Case
compress_key_sentence ~60-80% Query-focused retrieval, Q&A agents
compress_observation_filtering ~40-50% Multi-step tool use with redundant observations
Combined pipeline ~76% Production agents with variable context

Key insight from scripts/sync_chapter2_figures.py: compression before truncation preserves more semantic information than truncation alone, because truncation operates on raw token sequences without awareness of content relevance.


Key Source Files for Token Optimization

File Purpose
tests/test_ch8_evaluate_multilingual.py Token usage objects and compute_token_usage
tests/test_ch9_safety_policy_gate.py SafetyPolicyGate with budget enforcement
tests/test_ch9_interruption_manager.py Dialogue truncation on budget overflow
chapter2/context-compression/benchmark.py Compression algorithms (compress_key_sentence, compress_observation_filtering)
chapter9/harness-safety-gate/llm_generator.py Production SafetyPolicyGate implementation
scripts/sync_chapter2_figures.py Visualization of token savings benchmarks

Summary

  • Measure first: compute_token_usage in test_ch8_evaluate_multilingual.py provides granular, serializable token accounting across prompt, reasoning, and completion categories
  • Compress early: chapter2/context-compression/benchmark.py demonstrates 76% token reduction through semantic compression before counting
  • Truncate smartly: InterruptionManager evicts oldest turns when budgets force hard limits
  • Enforce strictly: SafetyPolicyGate in test_ch9_safety_policy_gate.py issues single-use confirmation tokens for override workflows, with configurable TTL

Frequently Asked Questions

What is the most effective way to reduce token usage in AI agents?

Semantic context compression before tokenization. The ai-agent-book benchmark shows that compress_key_sentence and compress_observation_filtering achieve 76% reduction versus uncompressed baselines. This outperforms naive truncation because it preserves query-relevant information while eliminating redundancy.

How does SafetyPolicyGate prevent token budget overruns?

By validating every tool call against max_tokens before execution. In tests/test_ch9_safety_policy_gate.py, the gate returns a Decision object—either approving immediately or requiring a single-use confirmation_token when the request would exceed quota. The token expires after token_ttl seconds, forcing explicit human or system approval for over-budget operations.

When should I use dialogue truncation versus context compression?

Compression first, truncation as fallback. chapter2/context-compression/benchmark.py strategies preserve semantic meaning by selecting informative content. Only when compressed context still exceeds max_tokens should InterruptionManager (tests/test_ch9_interruption_manager.py) discard oldest turns. This two-tier approach maximizes information density within budget.

Where is the production-ready implementation of these token controls?

The GAIA benchmark harness at chapter9/harness-safety-gate/llm_generator.py contains the hardened SafetyPolicyGate with full error handling and logging. Test suites in tests/test_ch9_*.py demonstrate integration patterns, while chapter2/context-compression/benchmark.py provides research-grade compression utilities suitable for adaptation to production tokenizers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →