# How to Manage LLM Context Window Size for Long-Running Agents: A 12-Factor Approach

> Master LLM context window size for long-running agents. Learn to serialize events, implement token budget guards, and optimize context for enhanced performance.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: how-to-guide
- Published: 2026-05-19

---

**To manage LLM context window size for long-running agents, treat the context as a first-class data structure by serializing events into compact XML-style tags, implementing token budget guards like `willExceedTokenLimit()`, and compacting or pausing execution when approaching model limits.**

The `humanlayer/12-factor-agents` framework provides architectural patterns specifically designed to prevent context window overflow in autonomous agent systems. When agents run for extended periods, they accumulate tool calls, user interactions, and environmental data that can quickly exceed token limits. By owning your context window rather than relying on default message formats, you maintain control over information density and agent continuity.

## Own the Context Window (Factor 3)

Traditional message-based formats waste tokens with verbose JSON schemas. According to [[`factor-3-own-your-context-window.md`](https://github.com/humanlayer/12-factor-agents/blob/main/factor-3-own-your-context-window.md)](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-3-own-your-context-window.md), you should serialize the entire thread state into a single user message using compact, XML-style tags.

**Key implementation strategies:**

- **Token estimation**: Approximate 4 characters per token to preemptively check `serialized.length / 4` against your model limit before each LLM call
- **Information density**: Replace JSON with compact tags like `<tool_call>intent: deploy_service\nstatus: pending</tool_call>`
- **Safety filtering**: Strip sensitive data during the serialization phase in `serializeOneEvent()` rather than post-processing

The framework recommends calculating token budget at the serialization layer. In [`src/agent.ts`](https://github.com/humanlayer/12-factor-agents/blob/main/src/agent.ts), the `willExceedTokenLimit()` method checks the serialized string length before invoking the LLM:

```typescript
// src/agent.ts
willExceedTokenLimit(serialized: string, limit = 8192): boolean {
  // Approximate 4 characters ≈ 1 token
  return Math.ceil(serialized.length / 4) > limit;
}

```

## Compact the History (Factor 8)

When the context window grows, collapse older events into summaries rather than dropping them entirely. As documented in [[`factor-8-own-your-control-flow.md`](https://github.com/humanlayer/12-factor-agents/blob/main/factor-8-own-your-control-flow.md)](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-8-own-your-control-flow.md), implement a `compact_context()` function that preserves recent events while summarizing historical ones.

```python
def compact_context(thread):
    # Keep the last N events (e.g., 10) and summarize the rest

    recent = thread.events[-10:]
    older = thread.events[:-10]
    summary = summarize_events(older)  # LLM-generated summary

    return recent + [Event(type="summary", data=summary)]

```

This approach maintains semantic continuity while reclaiming token budget. The summarization should occur in the control-flow loop before calling `determine_next_step()`, ensuring the LLM receives only relevant context.

## Pause and Resume (Factor 6)

When compaction is insufficient, pause the agent gracefully rather than crashing or truncating critical information. The architecture in [[`factor-6-launch-pause-resume.md`](https://github.com/humanlayer/12-factor-agents/blob/main/factor-6-launch-pause-resume.md)](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-6-launch-pause-resume.md) uses HTTP 202 status codes to signal that the agent requires external intervention.

```typescript
// src/agent.ts
async step() {
  const ctx = this.serializeForLLM();

  if (this.willExceedTokenLimit(ctx)) {
    await db.saveThread(thread);   // Persist state
    await pauseAgent(thread.id);   // Return HTTP 202
    return;                        // Break the LLM loop
  }
  
  const next = await determine_next_step(ctx);
  // Process next step...
}

```

The orchestrator persists the thread state to the database when `willExceedTokenLimit()` returns true, allowing the agent to resume via webhook after the user or an external service trims the context.

## Leverage Human-in-the-Loop (Factor 7)

Offload large data processing to humans through structured tools. Rather than injecting entire raw logs into the context window, use the pattern from [[`factor-7-contact-humans-with-tools.md`](https://github.com/humanlayer/12-factor-agents/blob/main/factor-7-contact-humans-with-tools.md)](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-7-contact-humans-with-tools.md) to request concise summaries:

```typescript
// webhook endpoint - after user provides summary
app.post('/webhook', async (req, res) => {
  const { thread_id, summary } = req.body;
  const thread = await db.loadThread(thread_id);
  thread.events.push({ type: 'summary', data: summary });
  await thread.step(); // Continue the LLM loop
  res.sendStatus(200);
});

```

This technique reduces token usage while preserving the semantic gist of complex operations that would otherwise consume hundreds of tokens.

## Compact Errors (Factor 9)

Errors consume disproportionate space when rendered with full stack traces. Following [[`factor-9-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/factor-9-compact-errors.md)](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-9-compact-errors.md), inject short `<error>` tags and remove them once resolved:

```xml
<error>
  Failed to connect to deployment service – retry later
</error>

```

Implement garbage collection in your serialization logic to drop resolved error tags, immediately reclaiming those tokens for productive context.

## Complete Implementation Example

Combine these patterns in the Agent class to manage context window size proactively:

```typescript
// src/agent.ts
class Agent {
  events: Event[] = [];

  private trimLeadingWhitespace(s: string) {
    return s.replace(/^[ \t]+/gm, '');
  }

  private serializeOneEvent(e: Event): string {
    const tag = e.type === 'tool_call' ? e.data.intent : e.type;
    const body = typeof e.data === 'object'
      ? Object.entries(e.data)
          .filter(([k]) => k !== 'intent')
          .map(([k, v]) => `${k}: ${v}`).join('\n')
      : e.data;
    return this.trimLeadingWhitespace(`
      <${tag}>
      ${body}
      </${tag}>
    `);
  }

  serializeForLLM(): string {
    return this.events.map(e => this.serializeOneEvent(e)).join('\n');
  }

  willExceedTokenLimit(serialized: string, limit = 8192): boolean {
    return Math.ceil(serialized.length / 4) > limit;
  }

  async compactOrPause(ctx: string) {
    // Try compaction first
    if (this.events.length > 10) {
      this.events = compact_context(this);
      const newCtx = this.serializeForLLM();
      if (!this.willExceedTokenLimit(newCtx)) {
        return;
      }
    }
    // If still over limit, pause
    await db.saveThread(this);
    await pauseAgent(this.id);
  }

  async step() {
    const ctx = this.serializeForLLM();
    
    if (this.willExceedTokenLimit(ctx)) {
      await this.compactOrPause(ctx);
      return;
    }
    
    const next = await determine_next_step(ctx);
    // Process response...
  }
}

```

This implementation demonstrates the XML serialization strategy shown in the workshop materials at [[`workshops/2025-05/sections/07-context-window/README.md`](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-05/sections/07-context-window/README.md)](https://github.com/humanlayer/12-factor-agents/blob/main/workshops/2025-05/sections/07-context-window/README.md).

## Summary

- **Treat context as a data structure**: Serialize events into compact XML-style tags using `serializeOneEvent()` rather than verbose JSON arrays
- **Implement token guards**: Use `willExceedTokenLimit()` with a 4:1 character-to-token ratio to prevent overflow before LLM invocation
- **Compact strategically**: Summarize events older than *N* steps while preserving recent context for continuity
- **Pause gracefully**: Persist thread state and return HTTP 202 when compaction fails, resuming via webhook after external trimming
- **Minimize error overhead**: Use brief `<error>` tags and remove them after resolution to reclaim tokens immediately

## Frequently Asked Questions

### How do you calculate token usage before sending to the LLM?

Use the `willExceedTokenLimit()` method with a conservative estimate of 4 characters per token. Measure the serialized context string length, divide by 4, and compare against your model's context limit. This approach, implemented in [`src/agent.ts`](https://github.com/humanlayer/12-factor-agents/blob/main/src/agent.ts), catches oversized contexts before the API call fails.

### What is the best format for serializing agent events?

XML-style tags provide superior token efficiency compared to JSON. Wrap each event in `<event_type>` tags with compact key-value pairs inside, filtering out metadata fields like `intent` that duplicate the tag name. This format reduces token count by approximately 30-40% compared to standard OpenAI message schemas.

### When should you pause versus compact the context window?

Attempt compaction first when you exceed 80% of your token budget by summarizing events older than 10 steps. If the compacted context still exceeds limits—or if the compaction would remove critical decision-making context—pause the agent with `pauseAgent()`, persist to `db.saveThread()`, and wait for external intervention via webhook.