How to Compact Errors into Context Window for Agent Self-Healing
Wrap tool execution in try/except blocks and append compact error representations to the agent's thread events so the LLM can see failures and decide whether to retry, adjust parameters, or escalate.
When building LLM-driven agents, the context window serves as the working memory for reasoning and decision-making. According to the humanlayer/12-factor-agents framework, Factor 9 specifies that agents must compact errors into the context window to enable self-healing behavior. By treating exceptions as serialized events within the thread's message history, you allow the model to inspect failures and autonomously determine recovery strategies without external intervention.
Why Error Compacting Matters for Self-Healing Agents
The context window is the only source of information an LLM can reason over at each turn. When a tool call fails—whether due to invalid parameters, API timeouts, or logic errors—the agent cannot recover unless those error details are fed back into the context stream. Without this feedback loop, the agent remains unaware that the previous action failed, leading to stalled or inconsistent state.
By compacting errors into the context window, you transform exceptions into first-class events that the LLM can analyze. The model can then issue deterministic retry commands, adjust parameters based on error messages, or escalate to human operators when error thresholds are exceeded.
The Thread Event Model
The 12-Factor Agents architecture relies on a mutable thread (or context) object that maintains a chronological list of events. Each event follows a strict schema with a type field and a data payload:
- Tool events: Record the intent to invoke a tool
- Result events: Store successful execution outputs
- Error events: Capture serialized exception data
This event log becomes the complete state representation passed to the LLM on each inference cycle. According to content/factor-09-compact-errors.md, the agent appends error events just like any other message, ensuring the failure context is visible in subsequent turns.
Implementing Error Compacting in Python
Basic Error Capture and Serialization
The core pattern requires wrapping tool execution in a try/except block and immediately serializing any caught exceptions into the thread. In workshops/2025-07-16/walkthrough/07-agent.py, this is implemented as a continuous loop where errors are appended as distinct events:
thread = {"events": [initial_message]}
while True:
# Ask the LLM what to do next, feeding the whole thread as context
next_step = await determine_next_step(thread_to_prompt(thread))
thread["events"].append({
"type": next_step.intent,
"data": next_step,
})
try:
# Execute the chosen tool / step
result = await handle_next_step(thread, next_step)
except Exception as e:
# Put the error back into the context so the LLM can see it
thread["events"].append({
"type": "error",
"data": format_error(e), # compact, user‑friendly string
})
# The loop continues – the LLM will now receive the error event
continue
else:
# Store the successful result and reset any error counters
thread["events"].append({
"type": f"{next_step.intent}_result",
"data": result,
})
# (optional) break when a “done” intent is returned
if next_step.intent == "done":
break
By appending the error event before continuing the loop, the next call to determine_next_step() includes the failure context in the prompt, allowing the model to react intelligently.
Creating a Compact Error Formatter
To respect token limits and prevent context window overflow, errors must be compacted before serialization. The format_error() function implemented in the workshops truncates stack traces while preserving essential diagnostic information:
def format_error(exc: Exception) -> str:
# Keep only the first 3 lines of the traceback for brevity
import traceback, io
buf = io.StringIO()
traceback.print_exception(type(exc), exc, exc.__traceback__, limit=3, file=buf)
return buf.getvalue().strip()
This approach balances informativeness with brevity. You can customize this function to strip sensitive internal paths or summarize error categories if your agent operates with strict token budgets.
Handling Retry Logic and Escalation
Pure self-healing requires safeguards against infinite loops. The reference implementation uses a consecutive_errors counter to enforce maximum retry thresholds before escalating to human operators:
consecutive_errors = 0
MAX_RETRIES = 3
while True:
# ... same as above up to the try block ...
try:
result = await handle_next_step(thread, next_step)
# Success – reset error counter
consecutive_errors = 0
except Exception as e:
consecutive_errors += 1
thread["events"].append({
"type": "error",
"data": format_error(e),
})
if consecutive_errors >= MAX_RETRIES:
# Escalate or abort
thread["events"].append({
"type": "escalation",
"data": "Too many retries – notifying a human.",
})
break
# otherwise loop again and let the LLM decide the next move
continue
This pattern integrates cleanly with Factor 7 (human-in-the-loop), allowing deterministic fallback when agent autonomy reaches its limits.
Complete Implementation Example
Combining these elements yields a robust self-healing agent loop. The workshops/2025-07-16/walkthrough/07-agent.py file demonstrates the complete implementation, showing how error events coexist with tool results and system messages in the thread. When errors occur, the LLM receives output like:
Traceback (most recent call last):
File "agent.py", line 45, in handle_next_step
result = tool(**params)
TypeError: calculate_distance() got an unexpected keyword argument 'unit'
The model can then emit a new tool call with corrected parameters {"lat": 40.7, "lon": -74.0, "units": "km"} to resolve the issue autonomously.
Key Files and References
The humanlayer/12-factor-agents repository provides canonical implementations and documentation for this pattern:
content/factor-09-compact-errors.md– Core documentation explaining the philosophy of error serialization and thread management.content/factor-08-own-your-control-flow.md– Explains how error compacting integrates with deterministic control flow decisions.workshops/2025-07-16/walkthrough/07-agent.py– Reference implementation demonstrating the error-compact loop in a production-style agent.workshops/2025-07-16/hack/inspect_notebook.py– Utility showing error detection patterns for notebook-based tool execution.
Summary
- Compact errors into context window by wrapping tool execution in try/except blocks and appending exception data to the thread events list.
- Use a
format_error()function to truncate stack traces and limit token consumption while preserving diagnostic value. - Maintain a
consecutive_errorscounter to prevent infinite retry loops and trigger escalation when thresholds are exceeded. - Store errors as typed events with
"type": "error"so the LLM can distinguish failures from successful results in the conversation history. - Reference
workshops/2025-07-16/walkthrough/07-agent.pyfor a complete working implementation following the 12-Factor Agents methodology.
Frequently Asked Questions
Why not simply raise exceptions to the orchestrator instead of compacting them?
Raising exceptions to an external orchestrator breaks the own-your-context-window principle. When you compact errors into the context window, the LLM retains agency to decide whether to retry with different parameters, attempt an alternative tool, or escalate. External exception handling forces the developer to hard-code recovery logic rather than letting the model reason about the failure context.
How much error detail should I include to keep it truly compact?
Limit stack traces to 3-5 lines using traceback.print_exception() with a limit parameter, and strip file paths that don't belong to your application code. Include the exception type and message, but omit framework internals. The goal is to give the LLM enough context to diagnose the issue without consuming excessive tokens that could displace other critical conversation history.
How does error compacting relate to Factor 8 (Own Your Control Flow)?
According to content/factor-08-own-your-control-flow.md, agents must maintain deterministic control over execution flow. By compacting errors into the context window, you enable the LLM to issue deterministic "retry" or "abort" decisions as tool calls rather than relying on imperative exception handling. This keeps the control flow explicit and inspectable in the event log rather than hidden in exception stack unwinding.
Can I use this pattern with pre-built agent frameworks like LangChain or CrewAI?
Yes, but you must ensure the framework exposes the raw context window or message history for mutation. Many high-level frameworks abstract away the thread model, making it difficult to inject custom error events. For full compatibility with Factor 9, implement your own thread management as shown in workshops/2025-07-16/walkthrough/07-agent.py, or verify that your chosen framework allows custom event types in the message history.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →