# How to Optimize Token Usage in LLM Context Windows: 7 Strategies from 12-Factor Agents

> Optimize LLM context window token usage with 7 strategies from 12-Factor Agents. Learn to pack data, serialize events, filter history, and use token-budget-aware trimming for efficient LLM performance.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: best-practices
- Published: 2026-05-19

---

**To optimize token usage in LLM context windows, pack related information into dense, custom-formatted messages using XML-style tags, serialize events efficiently with helper functions like `event_to_prompt`, filter low-priority historical data, and implement token-budget-aware trimming using libraries like `tiktoken`.**

Every token in an LLM's context window consumes finite space that could otherwise hold relevant information. The `humanlayer/12-factor-agents` repository provides a systematic approach to "owning" your context window, enabling you to shape exactly what the model sees at inference time for maximum information density and minimal token waste.

## Structure Context for Density

The default chat completion format repeats `"role"` and `"content"` fields for every message, consuming overhead tokens that carry no semantic value. Instead of scattering information across multiple system, user, and assistant messages, consolidate related items into a single message using a concise, custom format.

### Use XML-Style Tags for Logical Separation

Pack tool calls, results, and conversation history into one user message using XML-style tags. According to [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md) (lines 73-112), wrapping content in tags like `<slack_message>`, `<list_git_tags>`, and `<list_git_tags_result>` preserves semantic boundaries while eliminating the per-message overhead of standard role fields.

This approach allows the model to parse structured data without the token cost of repeating field names like `"role": "user"` and `"content":` for every interaction.

## Serialize Events Efficiently

When storing thread history, convert each event into a minimal string representation that eliminates repetitive wording. The `event_to_prompt` helper function in the 12-Factor Agents codebase (lines 31-35) transforms an `Event` object into a compact XML-wrapped block.

```python
from typing import List, Literal, Union
from dataclasses import dataclass

@dataclass
class Event:
    type: Literal[
        "list_git_tags", "deploy_backend", "error",
        # … other intents …

    ]
    data: Union[str, dict]

def event_to_prompt(event: Event) -> str:
    """Compact XML-style representation of an event."""
    payload = (
        event.data
        if isinstance(event.data, str)
        else "\n".join(f"{k}: {v}" for k, v in event.data.items())
    )
    return f"<{event.type}>\n{payload}\n</{event.type}>"

def thread_to_prompt(events: List[Event]) -> str:
    """Join all events into a single context block."""
    return "\n\n".join(event_to_prompt(e) for e in events)

```

By concatenating these compact blocks, you create a dense prompt that forces the model to focus on the core payload rather than parsing verbose JSON structures.

## Filter and Prioritize Information

Not every historical detail warrants inclusion in the active context window. As implemented in [`workshops/2025-07-16/walkthrough/07-agent.py`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-07-16/walkthrough/07-agent.py), you should dynamically hide resolved errors, stale intermediate results, and large raw documents that the model no longer needs to reference.

After an error is recovered, remove it from the context window to maintain the causal chain while keeping token counts low. This selective retention ensures only high-value information competes for the model's limited attention span.

## Use Compact Data Representations

Replace verbose JSON with space-efficient formats like YAML or key-value lines. For example, a list of Git tags represented in JSON consumes significantly more tokens than the equivalent YAML format:

```yaml
tags:
  - v1.2.3 abc123 2024-03-15
  - v1.2.2 def456 2024-03-14
  - v1.2.1 ghi789 2024-03-13

```

This format omits repeated field names while remaining human-readable for the model, directly reducing per-entry token overhead as demonstrated in the repository (lines 118-129).

## Implement Token-Budget-Aware Generation

Compute token counts dynamically during prompt construction to guarantee you never exceed model limits. Use `tiktoken` for OpenAI models or equivalent tokenizers to measure each component before assembly.

```python
import tiktoken

MAX_TOKENS = 8192
ENCODER = tiktoken.encoding_for_model("gpt-4")

def build_prompt(system_msg: str, events: List[Event]) -> str:
    # Serialize everything first

    context = thread_to_prompt(events)
    full_prompt = f"{system_msg}\n\n{context}"
    
    # Trim if needed by dropping lowest-priority items

    while len(ENCODER.encode(full_prompt)) > MAX_TOKENS:
        events.pop(0)  # Remove oldest event

        context = thread_to_prompt(events)
        full_prompt = f"{system_msg}\n\n{context}"
    return full_prompt

```

This dynamic trimming strategy ensures the most relevant data remains in view while respecting the hard constraints of the model's context window.

## Leverage Structured Tool Outputs

Define tool call schemas that output the same tag-based format you use for context. When the model returns `<list_git_tags>` or `<error>` tags directly, you eliminate conversion steps that would otherwise add tokens to the prompt. This symmetry between input context and output format streamlines the entire pipeline as documented in the tool-call schema examples (lines 53-58).

## Summary

- **Consolidate messages** into single user messages with XML-style tags to remove per-message overhead from role fields.
- **Serialize events** using `event_to_prompt` to pack type and data into minimal token representations.
- **Filter dynamically** to remove resolved errors and stale data before they reach the prompt.
- **Prefer compact formats** like YAML over verbose JSON to represent structured data.
- **Measure and trim** using `tiktoken` to stay within model-specific context limits while preserving high-priority information.
- **Align tool outputs** with your context format to eliminate parsing conversion costs.

## Frequently Asked Questions

### What is the most effective way to reduce token overhead in LLM prompts?

The most effective technique is replacing the standard message list format with a single user message containing XML-style tagged content. This eliminates the repeated `"role"` and `"content"` JSON keys that consume tokens without adding semantic value, as demonstrated in [`content/factor-03-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md).

### How does the `event_to_prompt` function save tokens?

The `event_to_prompt` function (lines 31-35) converts Event objects into compact XML-wrapped strings like `<list_git_tags>\nkey: value\n</list_git_tags>`, avoiding the verbose nesting and quotes required for JSON representation. This reduces the token count per event while maintaining clear structure.

### When should I filter events from the context window?

Remove events from the context window when they become irrelevant to future reasoning, such as resolved errors, superseded intermediate results, or large raw documents the model has already processed. This filtering should occur before the prompt reaches the model, ensuring only causally relevant information consumes token budget.

### Why use XML-style tags instead of standard JSON for LLM context?

XML-style tags provide clear semantic boundaries with fewer tokens than JSON's required quotes, braces, and field names. The model can parse `<tag>content</tag>` structures efficiently, and the format aligns naturally with custom prompt templates that combine system instructions and serialized context into a single dense message.