How to Compact Errors into Context Window for Agent Self-Healing

Wrap tool execution in try/except blocks and append compact error representations to the agent's thread events so the LLM can see failures and decide whether to retry, adjust parameters, or escalate.

When building LLM-driven agents, the context window serves as the working memory for reasoning and decision-making. According to the humanlayer/12-factor-agents framework, Factor 9 specifies that agents must compact errors into the context window to enable self-healing behavior. By treating exceptions as serialized events within the thread's message history, you allow the model to inspect failures and autonomously determine recovery strategies without external intervention.

Why Error Compacting Matters for Self-Healing Agents

The context window is the only source of information an LLM can reason over at each turn. When a tool call fails—whether due to invalid parameters, API timeouts, or logic errors—the agent cannot recover unless those error details are fed back into the context stream. Without this feedback loop, the agent remains unaware that the previous action failed, leading to stalled or inconsistent state.

By compacting errors into the context window, you transform exceptions into first-class events that the LLM can analyze. The model can then issue deterministic retry commands, adjust parameters based on error messages, or escalate to human operators when error thresholds are exceeded.

The Thread Event Model

The 12-Factor Agents architecture relies on a mutable thread (or context) object that maintains a chronological list of events. Each event follows a strict schema with a type field and a data payload:

  • Tool events: Record the intent to invoke a tool
  • Result events: Store successful execution outputs
  • Error events: Capture serialized exception data

This event log becomes the complete state representation passed to the LLM on each inference cycle. According to content/factor-09-compact-errors.md, the agent appends error events just like any other message, ensuring the failure context is visible in subsequent turns.

Implementing Error Compacting in Python

Basic Error Capture and Serialization

The core pattern requires wrapping tool execution in a try/except block and immediately serializing any caught exceptions into the thread. In workshops/2025-07-16/walkthrough/07-agent.py, this is implemented as a continuous loop where errors are appended as distinct events:

thread = {"events": [initial_message]}

while True:
    # Ask the LLM what to do next, feeding the whole thread as context

    next_step = await determine_next_step(thread_to_prompt(thread))

    thread["events"].append({
        "type": next_step.intent,
        "data": next_step,
    })

    try:
        # Execute the chosen tool / step

        result = await handle_next_step(thread, next_step)
    except Exception as e:
        # Put the error back into the context so the LLM can see it

        thread["events"].append({
            "type": "error",
            "data": format_error(e),          # compact, user‑friendly string

        })
        # The loop continues – the LLM will now receive the error event

        continue
    else:
        # Store the successful result and reset any error counters

        thread["events"].append({
            "type": f"{next_step.intent}_result",
            "data": result,
        })
        # (optional) break when a “done” intent is returned

        if next_step.intent == "done":
            break

By appending the error event before continuing the loop, the next call to determine_next_step() includes the failure context in the prompt, allowing the model to react intelligently.

Creating a Compact Error Formatter

To respect token limits and prevent context window overflow, errors must be compacted before serialization. The format_error() function implemented in the workshops truncates stack traces while preserving essential diagnostic information:

def format_error(exc: Exception) -> str:
    # Keep only the first 3 lines of the traceback for brevity

    import traceback, io
    buf = io.StringIO()
    traceback.print_exception(type(exc), exc, exc.__traceback__, limit=3, file=buf)
    return buf.getvalue().strip()

This approach balances informativeness with brevity. You can customize this function to strip sensitive internal paths or summarize error categories if your agent operates with strict token budgets.

Handling Retry Logic and Escalation

Pure self-healing requires safeguards against infinite loops. The reference implementation uses a consecutive_errors counter to enforce maximum retry thresholds before escalating to human operators:

consecutive_errors = 0
MAX_RETRIES = 3

while True:
    # ... same as above up to the try block ...

    try:
        result = await handle_next_step(thread, next_step)
        # Success – reset error counter

        consecutive_errors = 0
    except Exception as e:
        consecutive_errors += 1
        thread["events"].append({
            "type": "error",
            "data": format_error(e),
        })
        if consecutive_errors >= MAX_RETRIES:
            # Escalate or abort

            thread["events"].append({
                "type": "escalation",
                "data": "Too many retries – notifying a human.",
            })
            break
        # otherwise loop again and let the LLM decide the next move

        continue

This pattern integrates cleanly with Factor 7 (human-in-the-loop), allowing deterministic fallback when agent autonomy reaches its limits.

Complete Implementation Example

Combining these elements yields a robust self-healing agent loop. The workshops/2025-07-16/walkthrough/07-agent.py file demonstrates the complete implementation, showing how error events coexist with tool results and system messages in the thread. When errors occur, the LLM receives output like:


Traceback (most recent call last):
  File "agent.py", line 45, in handle_next_step
    result = tool(**params)
TypeError: calculate_distance() got an unexpected keyword argument 'unit'

The model can then emit a new tool call with corrected parameters {"lat": 40.7, "lon": -74.0, "units": "km"} to resolve the issue autonomously.

Key Files and References

The humanlayer/12-factor-agents repository provides canonical implementations and documentation for this pattern:

Summary

  • Compact errors into context window by wrapping tool execution in try/except blocks and appending exception data to the thread events list.
  • Use a format_error() function to truncate stack traces and limit token consumption while preserving diagnostic value.
  • Maintain a consecutive_errors counter to prevent infinite retry loops and trigger escalation when thresholds are exceeded.
  • Store errors as typed events with "type": "error" so the LLM can distinguish failures from successful results in the conversation history.
  • Reference workshops/2025-07-16/walkthrough/07-agent.py for a complete working implementation following the 12-Factor Agents methodology.

Frequently Asked Questions

Why not simply raise exceptions to the orchestrator instead of compacting them?

Raising exceptions to an external orchestrator breaks the own-your-context-window principle. When you compact errors into the context window, the LLM retains agency to decide whether to retry with different parameters, attempt an alternative tool, or escalate. External exception handling forces the developer to hard-code recovery logic rather than letting the model reason about the failure context.

How much error detail should I include to keep it truly compact?

Limit stack traces to 3-5 lines using traceback.print_exception() with a limit parameter, and strip file paths that don't belong to your application code. Include the exception type and message, but omit framework internals. The goal is to give the LLM enough context to diagnose the issue without consuming excessive tokens that could displace other critical conversation history.

How does error compacting relate to Factor 8 (Own Your Control Flow)?

According to content/factor-08-own-your-control-flow.md, agents must maintain deterministic control over execution flow. By compacting errors into the context window, you enable the LLM to issue deterministic "retry" or "abort" decisions as tool calls rather than relying on imperative exception handling. This keeps the control flow explicit and inspectable in the event log rather than hidden in exception stack unwinding.

Can I use this pattern with pre-built agent frameworks like LangChain or CrewAI?

Yes, but you must ensure the framework exposes the raw context window or message history for mutation. Many high-level frameworks abstract away the thread model, making it difficult to inject custom error events. For full compatibility with Factor 9, implement your own thread management as shown in workshops/2025-07-16/walkthrough/07-agent.py, or verify that your chosen framework allows custom event types in the message history.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →