# Optimizing Token Usage for AI Agent Context Management: A Complete Guide Based on ai-agent-book

> Master AI agent token optimization with context compression and budget-aware truncation. This guide shows how to reduce consumption by up to 76% without losing quality. Learn from bojieli ai agent book.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-18

---

**AI agents can reduce token consumption by up to 76% through context compression, budget-aware truncation, and policy-based gating—without sacrificing response quality.** The `ai-agent-book` repository by `bojieli` demonstrates a production-ready architecture that separates **measurement**, **compression**, and **enforcement** into modular, testable components.

Built for researchers and engineers building scalable LLM systems, this codebase shows how to keep agents within budget while preserving the most semantically valuable context. Whether you're running cost-sensitive APIs or managing constrained edge deployments, the patterns in `bojieli/ai-agent-book` provide a battle-tested foundation.

---

## Why Token Optimization Matters for AI Agents

Every LLM call incurs costs proportional to token count. For autonomous agents that may chain dozens of tool calls across extended sessions, **unbounded context growth** leads to:

- **Runaway API costs** from ever-expanding dialogue history
- **Latency degradation** as prompts exceed model processing windows
- **Truncation errors** when prompts hit hard provider limits

The `ai-agent-book` solution treats token usage as a **first-class resource** to be measured, compressed, and governed—rather than an afterthought.

---

## Four Core Mechanisms for Token-Aware Context Management

The repository implements four interconnected systems. Each operates independently but composes into a unified pipeline.

### Token Accounting with `compute_token_usage`

At [`tests/test_ch8_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch8_evaluate_multilingual.py), the codebase defines a lightweight `TokenUsageObj` pattern that serializes consumption across prompt, completion, reasoning, and total token categories.

```python
from utils.token_usage import compute_token_usage

prompt = "Summarize the following article in two sentences."
completion = "The article discusses..."
token_counts = {"prompt_tokens": 12, "completion_tokens": 24, "total_tokens": 36}

usage = compute_token_usage(
    prompt_text=prompt,
    reasoning_text="I considered the key points…",
    answer_text=completion,
    model_output=token_counts,
)

# Result: {'prompt_tokens': 12, 'completion_tokens': 24,

#          'reasoning_tokens': 5, 'total_tokens': 41}

```

This accounting layer enables **per-call tracking** and **aggregate budget monitoring**. The helper aggregates counts across an agent's full execution trace, making it possible to attribute costs to specific tool invocations or dialogue branches.

### Safety-Policy Gating with `SafetyPolicyGate`

Hard limits prevent budget overruns. In [`tests/test_ch9_safety_policy_gate.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_safety_policy_gate.py), the `SafetyPolicyGate` class enforces `max_tokens`, `max_timeout`, and `max_file_bytes` thresholds:

```python
from safety_gate import SafetyPolicyGate

gate = SafetyPolicyGate(max_tokens=10_000, token_ttl=60.0)

# Within budget—proceeds immediately

decision = gate.validate_tool_call("generate_text", {"max_tokens": 4_000})
assert decision.confirmation_token is None

# Would exceed budget—issues single-use confirmation token

decision = gate.validate_tool_call("generate_text", {"max_tokens": 7_000})
token = decision.confirmation_token

# Must re-validate with token before execution

decision = gate.validate_tool_call(
    "generate_text",
    {"max_tokens": 7_000},
    confirm_token=token,
)

```

Key design decisions in this implementation:

- **Single-use confirmation tokens** prevent replay attacks on budget overrides
- **Configurable TTL** (`token_ttl=60.0`) limits the window for explicit approval
- **Hard return of `Decision` objects** forces callers to handle denial cases explicitly

The production implementation lives at [`chapter9/harness-safety-gate/llm_generator.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/harness-safety-gate/llm_generator.py), used in the GAIA benchmark experiments.

### Context Compression with `compress_key_sentence` and `compress_observation_filtering`

Before tokens are counted or budget-checked, raw context can be semantically compressed. The benchmark utilities in [`chapter2/context-compression/benchmark.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/context-compression/benchmark.py) provide two strategies:

```python
from benchmark import compress_key_sentence, compress_observation_filtering

raw_context = """
  The user asked about token optimisation.
  The system replied with a long explanation.
  The user then inquired about safety policies.
  ...
"""

# Strategy 1: Extract query-relevant sentence

key = compress_key_sentence(raw_context, query="token optimisation")

# Strategy 2: Filter redundant observations

compressed = compress_observation_filtering(key)

# Result: "The user asked about token optimisation."

```

The [`scripts/sync_chapter2_figures.py`](https://github.com/bojieli/ai-agent-book/blob/main/scripts/sync_chapter2_figures.py) visualization shows **76% token reduction** versus uncompressed baselines, with minimal impact on downstream task accuracy. Compression runs **before tokenization**, so savings compound across the full pipeline.

### Dialogue Truncation with `InterruptionManager`

When compression alone cannot fit the budget, `InterruptionManager` (demonstrated in [`tests/test_ch9_interruption_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_interruption_manager.py)) implements **oldest-first eviction**:

```python

# Inside safe_llm_call workflow

ctx = InterruptionManager().get_dialogue_context()

# Returns trimmed list when token budget exceeded

```

This prioritizes **recency**—the assumption being that recent turns carry more predictive value for the agent's next action than distant history.

---

## Assembling the Full Pipeline

These four mechanisms compose into a complete token-aware request flow:

```python
def safe_llm_call(prompt, gate, tokenizer, query):
    # 1. Retrieve and truncate dialogue if needed

    ctx = InterruptionManager().get_dialogue_context()
    
    # 2. Compress remaining context semantically

    compact = compress_key_sentence(ctx, query)
    
    # 3. Count tokens for budget check

    usage = compute_token_usage(
        prompt, compact, "", 
        model_output=tokenizer(compact)
    )
    
    # 4. Enforce policy; obtain confirmation if over budget

    headroom = 500  # reserve tokens for completion

    decision = gate.validate_tool_call(
        "generate_text",
        {"max_tokens": usage["total_tokens"] + headroom},
    )
    
    if decision.confirmation_token:
        decision = gate.validate_tool_call(
            "generate_text",
            {"max_tokens": usage["total_tokens"] + headroom},
            confirm_token=decision.confirmation_token,
        )
    
    return decision  # approved output or denial

```

This pattern appears throughout the repository's evaluation harnesses, adapted for multilingual benchmarks ([`test_ch8_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/test_ch8_evaluate_multilingual.py)) and safety-critical pathways ([`test_ch9_safety_policy_gate.py`](https://github.com/bojieli/ai-agent-book/blob/main/test_ch9_safety_policy_gate.py)).

---

## Performance Trade-offs and Benchmarking Results

The [`chapter2/context-compression/benchmark.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/context-compression/benchmark.py) module provides empirical guidance on compression strategy selection:

| Strategy | Token Reduction | Use Case |
|----------|-----------------|----------|
| `compress_key_sentence` | ~60-80% | Query-focused retrieval, Q&A agents |
| `compress_observation_filtering` | ~40-50% | Multi-step tool use with redundant observations |
| Combined pipeline | ~76% | Production agents with variable context |

Key insight from [`scripts/sync_chapter2_figures.py`](https://github.com/bojieli/ai-agent-book/blob/main/scripts/sync_chapter2_figures.py): **compression before truncation** preserves more semantic information than truncation alone, because truncation operates on raw token sequences without awareness of content relevance.

---

## Key Source Files for Token Optimization

| File | Purpose |
|------|---------|
| [`tests/test_ch8_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch8_evaluate_multilingual.py) | Token usage objects and `compute_token_usage` |
| [`tests/test_ch9_safety_policy_gate.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_safety_policy_gate.py) | `SafetyPolicyGate` with budget enforcement |
| [`tests/test_ch9_interruption_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_interruption_manager.py) | Dialogue truncation on budget overflow |
| [`chapter2/context-compression/benchmark.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/context-compression/benchmark.py) | Compression algorithms (`compress_key_sentence`, `compress_observation_filtering`) |
| [`chapter9/harness-safety-gate/llm_generator.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/harness-safety-gate/llm_generator.py) | Production `SafetyPolicyGate` implementation |
| [`scripts/sync_chapter2_figures.py`](https://github.com/bojieli/ai-agent-book/blob/main/scripts/sync_chapter2_figures.py) | Visualization of token savings benchmarks |

---

## Summary

- **Measure first**: `compute_token_usage` in [`test_ch8_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/test_ch8_evaluate_multilingual.py) provides granular, serializable token accounting across prompt, reasoning, and completion categories
- **Compress early**: [`chapter2/context-compression/benchmark.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/context-compression/benchmark.py) demonstrates 76% token reduction through semantic compression before counting
- **Truncate smartly**: `InterruptionManager` evicts oldest turns when budgets force hard limits
- **Enforce strictly**: `SafetyPolicyGate` in [`test_ch9_safety_policy_gate.py`](https://github.com/bojieli/ai-agent-book/blob/main/test_ch9_safety_policy_gate.py) issues single-use confirmation tokens for override workflows, with configurable TTL

---

## Frequently Asked Questions

### What is the most effective way to reduce token usage in AI agents?

**Semantic context compression before tokenization.** The `ai-agent-book` benchmark shows that `compress_key_sentence` and `compress_observation_filtering` achieve 76% reduction versus uncompressed baselines. This outperforms naive truncation because it preserves query-relevant information while eliminating redundancy.

### How does `SafetyPolicyGate` prevent token budget overruns?

By validating every tool call against `max_tokens` before execution. In [`tests/test_ch9_safety_policy_gate.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_safety_policy_gate.py), the gate returns a `Decision` object—either approving immediately or requiring a single-use `confirmation_token` when the request would exceed quota. The token expires after `token_ttl` seconds, forcing explicit human or system approval for over-budget operations.

### When should I use dialogue truncation versus context compression?

**Compression first, truncation as fallback.** [`chapter2/context-compression/benchmark.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/context-compression/benchmark.py) strategies preserve semantic meaning by selecting informative content. Only when compressed context still exceeds `max_tokens` should `InterruptionManager` ([`tests/test_ch9_interruption_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch9_interruption_manager.py)) discard oldest turns. This two-tier approach maximizes information density within budget.

### Where is the production-ready implementation of these token controls?

The GAIA benchmark harness at [`chapter9/harness-safety-gate/llm_generator.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/harness-safety-gate/llm_generator.py) contains the hardened `SafetyPolicyGate` with full error handling and logging. Test suites in `tests/test_ch9_*.py` demonstrate integration patterns, while [`chapter2/context-compression/benchmark.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/context-compression/benchmark.py) provides research-grade compression utilities suitable for adaptation to production tokenizers.