How to Manage LLM Context Window Size for Long-Running Agents: A 12-Factor Approach
To manage LLM context window size for long-running agents, treat the context as a first-class data structure by serializing events into compact XML-style tags, implementing token budget guards like willExceedTokenLimit(), and compacting or pausing execution when approaching model limits.
The humanlayer/12-factor-agents framework provides architectural patterns specifically designed to prevent context window overflow in autonomous agent systems. When agents run for extended periods, they accumulate tool calls, user interactions, and environmental data that can quickly exceed token limits. By owning your context window rather than relying on default message formats, you maintain control over information density and agent continuity.
Own the Context Window (Factor 3)
Traditional message-based formats waste tokens with verbose JSON schemas. According to [factor-3-own-your-context-window.md](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-3-own-your-context-window.md), you should serialize the entire thread state into a single user message using compact, XML-style tags.
Key implementation strategies:
- Token estimation: Approximate 4 characters per token to preemptively check
serialized.length / 4against your model limit before each LLM call - Information density: Replace JSON with compact tags like
<tool_call>intent: deploy_service\nstatus: pending</tool_call> - Safety filtering: Strip sensitive data during the serialization phase in
serializeOneEvent()rather than post-processing
The framework recommends calculating token budget at the serialization layer. In src/agent.ts, the willExceedTokenLimit() method checks the serialized string length before invoking the LLM:
// src/agent.ts
willExceedTokenLimit(serialized: string, limit = 8192): boolean {
// Approximate 4 characters ≈ 1 token
return Math.ceil(serialized.length / 4) > limit;
}
Compact the History (Factor 8)
When the context window grows, collapse older events into summaries rather than dropping them entirely. As documented in [factor-8-own-your-control-flow.md](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-8-own-your-control-flow.md), implement a compact_context() function that preserves recent events while summarizing historical ones.
def compact_context(thread):
# Keep the last N events (e.g., 10) and summarize the rest
recent = thread.events[-10:]
older = thread.events[:-10]
summary = summarize_events(older) # LLM-generated summary
return recent + [Event(type="summary", data=summary)]
This approach maintains semantic continuity while reclaiming token budget. The summarization should occur in the control-flow loop before calling determine_next_step(), ensuring the LLM receives only relevant context.
Pause and Resume (Factor 6)
When compaction is insufficient, pause the agent gracefully rather than crashing or truncating critical information. The architecture in [factor-6-launch-pause-resume.md](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-6-launch-pause-resume.md) uses HTTP 202 status codes to signal that the agent requires external intervention.
// src/agent.ts
async step() {
const ctx = this.serializeForLLM();
if (this.willExceedTokenLimit(ctx)) {
await db.saveThread(thread); // Persist state
await pauseAgent(thread.id); // Return HTTP 202
return; // Break the LLM loop
}
const next = await determine_next_step(ctx);
// Process next step...
}
The orchestrator persists the thread state to the database when willExceedTokenLimit() returns true, allowing the agent to resume via webhook after the user or an external service trims the context.
Leverage Human-in-the-Loop (Factor 7)
Offload large data processing to humans through structured tools. Rather than injecting entire raw logs into the context window, use the pattern from [factor-7-contact-humans-with-tools.md](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-7-contact-humans-with-tools.md) to request concise summaries:
// webhook endpoint - after user provides summary
app.post('/webhook', async (req, res) => {
const { thread_id, summary } = req.body;
const thread = await db.loadThread(thread_id);
thread.events.push({ type: 'summary', data: summary });
await thread.step(); // Continue the LLM loop
res.sendStatus(200);
});
This technique reduces token usage while preserving the semantic gist of complex operations that would otherwise consume hundreds of tokens.
Compact Errors (Factor 9)
Errors consume disproportionate space when rendered with full stack traces. Following [factor-9-compact-errors.md](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-9-compact-errors.md), inject short <error> tags and remove them once resolved:
<error>
Failed to connect to deployment service – retry later
</error>
Implement garbage collection in your serialization logic to drop resolved error tags, immediately reclaiming those tokens for productive context.
Complete Implementation Example
Combine these patterns in the Agent class to manage context window size proactively:
// src/agent.ts
class Agent {
events: Event[] = [];
private trimLeadingWhitespace(s: string) {
return s.replace(/^[ \t]+/gm, '');
}
private serializeOneEvent(e: Event): string {
const tag = e.type === 'tool_call' ? e.data.intent : e.type;
const body = typeof e.data === 'object'
? Object.entries(e.data)
.filter(([k]) => k !== 'intent')
.map(([k, v]) => `${k}: ${v}`).join('\n')
: e.data;
return this.trimLeadingWhitespace(`
<${tag}>
${body}
</${tag}>
`);
}
serializeForLLM(): string {
return this.events.map(e => this.serializeOneEvent(e)).join('\n');
}
willExceedTokenLimit(serialized: string, limit = 8192): boolean {
return Math.ceil(serialized.length / 4) > limit;
}
async compactOrPause(ctx: string) {
// Try compaction first
if (this.events.length > 10) {
this.events = compact_context(this);
const newCtx = this.serializeForLLM();
if (!this.willExceedTokenLimit(newCtx)) {
return;
}
}
// If still over limit, pause
await db.saveThread(this);
await pauseAgent(this.id);
}
async step() {
const ctx = this.serializeForLLM();
if (this.willExceedTokenLimit(ctx)) {
await this.compactOrPause(ctx);
return;
}
const next = await determine_next_step(ctx);
// Process response...
}
}
This implementation demonstrates the XML serialization strategy shown in the workshop materials at [workshops/2025-05/sections/07-context-window/README.md](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-05/sections/07-context-window/README.md).
Summary
- Treat context as a data structure: Serialize events into compact XML-style tags using
serializeOneEvent()rather than verbose JSON arrays - Implement token guards: Use
willExceedTokenLimit()with a 4:1 character-to-token ratio to prevent overflow before LLM invocation - Compact strategically: Summarize events older than N steps while preserving recent context for continuity
- Pause gracefully: Persist thread state and return HTTP 202 when compaction fails, resuming via webhook after external trimming
- Minimize error overhead: Use brief
<error>tags and remove them after resolution to reclaim tokens immediately
Frequently Asked Questions
How do you calculate token usage before sending to the LLM?
Use the willExceedTokenLimit() method with a conservative estimate of 4 characters per token. Measure the serialized context string length, divide by 4, and compare against your model's context limit. This approach, implemented in src/agent.ts, catches oversized contexts before the API call fails.
What is the best format for serializing agent events?
XML-style tags provide superior token efficiency compared to JSON. Wrap each event in <event_type> tags with compact key-value pairs inside, filtering out metadata fields like intent that duplicate the tag name. This format reduces token count by approximately 30-40% compared to standard OpenAI message schemas.
When should you pause versus compact the context window?
Attempt compaction first when you exceed 80% of your token budget by summarizing events older than 10 steps. If the compacted context still exceeds limits—or if the compaction would remove critical decision-making context—pause the agent with pauseAgent(), persist to db.saveThread(), and wait for external intervention via webhook.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →