How to Optimize Token Usage in LLM Context Windows: 7 Strategies from 12-Factor Agents
To optimize token usage in LLM context windows, pack related information into dense, custom-formatted messages using XML-style tags, serialize events efficiently with helper functions like event_to_prompt, filter low-priority historical data, and implement token-budget-aware trimming using libraries like tiktoken.
Every token in an LLM's context window consumes finite space that could otherwise hold relevant information. The humanlayer/12-factor-agents repository provides a systematic approach to "owning" your context window, enabling you to shape exactly what the model sees at inference time for maximum information density and minimal token waste.
Structure Context for Density
The default chat completion format repeats "role" and "content" fields for every message, consuming overhead tokens that carry no semantic value. Instead of scattering information across multiple system, user, and assistant messages, consolidate related items into a single message using a concise, custom format.
Use XML-Style Tags for Logical Separation
Pack tool calls, results, and conversation history into one user message using XML-style tags. According to content/factor-03-own-your-context-window.md (lines 73-112), wrapping content in tags like <slack_message>, <list_git_tags>, and <list_git_tags_result> preserves semantic boundaries while eliminating the per-message overhead of standard role fields.
This approach allows the model to parse structured data without the token cost of repeating field names like "role": "user" and "content": for every interaction.
Serialize Events Efficiently
When storing thread history, convert each event into a minimal string representation that eliminates repetitive wording. The event_to_prompt helper function in the 12-Factor Agents codebase (lines 31-35) transforms an Event object into a compact XML-wrapped block.
from typing import List, Literal, Union
from dataclasses import dataclass
@dataclass
class Event:
type: Literal[
"list_git_tags", "deploy_backend", "error",
# … other intents …
]
data: Union[str, dict]
def event_to_prompt(event: Event) -> str:
"""Compact XML-style representation of an event."""
payload = (
event.data
if isinstance(event.data, str)
else "\n".join(f"{k}: {v}" for k, v in event.data.items())
)
return f"<{event.type}>\n{payload}\n</{event.type}>"
def thread_to_prompt(events: List[Event]) -> str:
"""Join all events into a single context block."""
return "\n\n".join(event_to_prompt(e) for e in events)
By concatenating these compact blocks, you create a dense prompt that forces the model to focus on the core payload rather than parsing verbose JSON structures.
Filter and Prioritize Information
Not every historical detail warrants inclusion in the active context window. As implemented in workshops/2025-07-16/walkthrough/07-agent.py, you should dynamically hide resolved errors, stale intermediate results, and large raw documents that the model no longer needs to reference.
After an error is recovered, remove it from the context window to maintain the causal chain while keeping token counts low. This selective retention ensures only high-value information competes for the model's limited attention span.
Use Compact Data Representations
Replace verbose JSON with space-efficient formats like YAML or key-value lines. For example, a list of Git tags represented in JSON consumes significantly more tokens than the equivalent YAML format:
tags:
- v1.2.3 abc123 2024-03-15
- v1.2.2 def456 2024-03-14
- v1.2.1 ghi789 2024-03-13
This format omits repeated field names while remaining human-readable for the model, directly reducing per-entry token overhead as demonstrated in the repository (lines 118-129).
Implement Token-Budget-Aware Generation
Compute token counts dynamically during prompt construction to guarantee you never exceed model limits. Use tiktoken for OpenAI models or equivalent tokenizers to measure each component before assembly.
import tiktoken
MAX_TOKENS = 8192
ENCODER = tiktoken.encoding_for_model("gpt-4")
def build_prompt(system_msg: str, events: List[Event]) -> str:
# Serialize everything first
context = thread_to_prompt(events)
full_prompt = f"{system_msg}\n\n{context}"
# Trim if needed by dropping lowest-priority items
while len(ENCODER.encode(full_prompt)) > MAX_TOKENS:
events.pop(0) # Remove oldest event
context = thread_to_prompt(events)
full_prompt = f"{system_msg}\n\n{context}"
return full_prompt
This dynamic trimming strategy ensures the most relevant data remains in view while respecting the hard constraints of the model's context window.
Leverage Structured Tool Outputs
Define tool call schemas that output the same tag-based format you use for context. When the model returns <list_git_tags> or <error> tags directly, you eliminate conversion steps that would otherwise add tokens to the prompt. This symmetry between input context and output format streamlines the entire pipeline as documented in the tool-call schema examples (lines 53-58).
Summary
- Consolidate messages into single user messages with XML-style tags to remove per-message overhead from role fields.
- Serialize events using
event_to_promptto pack type and data into minimal token representations. - Filter dynamically to remove resolved errors and stale data before they reach the prompt.
- Prefer compact formats like YAML over verbose JSON to represent structured data.
- Measure and trim using
tiktokento stay within model-specific context limits while preserving high-priority information. - Align tool outputs with your context format to eliminate parsing conversion costs.
Frequently Asked Questions
What is the most effective way to reduce token overhead in LLM prompts?
The most effective technique is replacing the standard message list format with a single user message containing XML-style tagged content. This eliminates the repeated "role" and "content" JSON keys that consume tokens without adding semantic value, as demonstrated in content/factor-03-own-your-context-window.md.
How does the event_to_prompt function save tokens?
The event_to_prompt function (lines 31-35) converts Event objects into compact XML-wrapped strings like <list_git_tags>\nkey: value\n</list_git_tags>, avoiding the verbose nesting and quotes required for JSON representation. This reduces the token count per event while maintaining clear structure.
When should I filter events from the context window?
Remove events from the context window when they become irrelevant to future reasoning, such as resolved errors, superseded intermediate results, or large raw documents the model has already processed. This filtering should occur before the prompt reaches the model, ensuring only causally relevant information consumes token budget.
Why use XML-style tags instead of standard JSON for LLM context?
XML-style tags provide clear semantic boundaries with fewer tokens than JSON's required quotes, braces, and field names. The model can parse <tag>content</tag> structures efficiently, and the format aligns naturally with custom prompt templates that combine system instructions and serialized context into a single dense message.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →