Agent Error Recovery and Retry Logic in 12-Factor Agents: A Complete Guide
Treat tool failures as structured events appended to the agent's context window, allowing the LLM to self-correct on the next iteration while using a consecutive-error counter to limit retries and enable escalation.
When building autonomous agents in the humanlayer/12-factor-agents framework, handling transient failures gracefully is critical for robust operation. Instead of aborting workflows when a tool fails, the recommended pattern treats errors as first-class context, giving the LLM visibility into what went wrong so it can propose corrected actions. This approach combines structured error recording with intelligent retry limits to create self-healing agents that know when to persist and when to escalate.
The Core Pattern: Treat Errors as Context
According to content/factor-09-compact-errors.md, the fundamental principle is to record every error as an event in the thread's history. When a tool invocation fails, you append a formatted error event to thread["events"] rather than raising an exception that breaks the execution loop.
This keeps the LLM's context complete and machine-readable, allowing the model to see the exact error message, error type, and stack trace on the next iteration. The agent can then reason about the failure and select a different tool or adjust its parameters.
The pattern requires maintaining a consecutive-error counter in the same thread structure that holds all execution state. As noted in content/factor-05-unify-execution-state.md, keeping this counter in the unified thread state (lines 11-14) ensures all execution metadata lives in one place, simplifying serialization and debugging.
Implementing the Retry Loop
The basic implementation follows a four-step cycle inside your control flow (as described in content/factor-08-own-your-control-flow.md, lines 10-18):
- Determine the next step from the current context window.
- Execute the step inside a
try / exceptblock. - On success, record the result event and reset
consecutive_errorsto zero. - On failure, record an error event, increment the counter, and decide whether to retry or abort.
Basic Error Handling Loop
Here is the canonical pattern from the repository showing how to structure your execution loop:
thread = {"events": [initial_message]} # initial context
while True:
next_step = await determine_next_step(thread_to_prompt(thread))
thread["events"].append({
"type": next_step.intent,
"data": next_step,
})
try:
# Run the tool associated with the intent
result = await handle_next_step(thread, next_step)
thread["events"].append({
"type": f"{next_step.intent}_result",
"data": result,
})
consecutive_errors = 0 # success → reset counter
except Exception as e:
# Record the error and decide whether to retry
consecutive_errors += 1
thread["events"].append({
"type": "error",
"data": format_error(e),
})
if consecutive_errors >= 3:
# Too many failures – break or escalate
await notify_human(thread, e) # see factor-07
break
# otherwise, loop again and let the LLM pick a new tool
This loop demonstrates the Compact Errors factor (lines 21-28 in content/factor-09-compact-errors.md), where the format_error() function ensures the exception is serialized into a clean, readable format for the LLM.
Using a Dedicated Retry Helper
For reusable logic across multiple tools, extract the retry mechanism into a helper function:
MAX_RETRIES = 3
async def run_with_retry(step, thread):
for attempt in range(1, MAX_RETRIES + 1):
try:
return await handle_next_step(thread, step)
except Exception as exc:
thread["events"].append({
"type": "error",
"data": format_error(exc, attempt=attempt),
})
if attempt == MAX_RETRIES:
raise # bubble up after final failure
Note that this approach still records each failure attempt to the thread, giving the LLM full visibility into repeated failures while enforcing a hard limit on retries.
Escalation and Control Flow
When the consecutive-error counter exceeds your threshold (typically three attempts), you must decide whether to continue looping, reset context, or escalate to human oversight (lines 33-58 and line 62 in content/factor-09-compact-errors.md).
When to Escalate to Humans
The Contact Humans With Tools factor (content/factor-07-contact-humans-with-tools.md) provides the pattern for handling persistent failures:
if consecutive_errors >= MAX_RETRIES:
await send_message_to_human(
f"Tool `{next_step.intent}` failed {consecutive_errors} times. "
"Human review required."
)
break # pause the loop until a human resolves the issue
This creates a clean handoff point where the agent stops spinning on a broken operation and waits for human intervention. The thread state remains intact, allowing the human to inspect the error history and either fix the underlying issue or guide the agent forward.
Integrating with Control Flow
As documented in content/factor-08-own-your-control-flow.md, error handling should integrate with your custom control-flow logic. Errors can trigger specific branches—such as pausing execution, triggering a back-off policy, or switching to a fallback tool. The key is that the error event becomes part of the decision context, not a termination signal.
State Management Benefits
By storing the consecutive_errors counter within the thread structure alongside all events, you achieve full deterministic state as described in content/factor-05-unify-execution-state.md. This means:
- Serialization: The entire execution state, including retry history, can be persisted to a database or passed between services.
- Resumability: If your agent process crashes, you can resume exactly where it left off, including knowledge of how many consecutive errors occurred before the crash.
- Debugging: Inspecting the thread reveals the complete history of attempts, failures, and retries without needing external logs.
Summary
- Record errors as events: Append formatted exceptions to
thread["events"]with type"error"to give the LLM visibility into failures. - Use a consecutive-error counter: Maintain
consecutive_errorsin the unified thread state to prevent infinite retry loops (seecontent/factor-09-compact-errors.md, lines 33-58). - Reset on success: Always reset the error counter to zero after successful tool execution to distinguish between transient and persistent failures.
- Escalate gracefully: When
consecutive_errorsexceeds your threshold (e.g., three), break the loop and usesend_message_to_human()to escalate (seecontent/factor-07-contact-humans-with-tools.md). - Keep state unified: Store retry counters in the same
threadstructure as events to maintain deterministic, serializable execution state.
Frequently Asked Questions
How does the 12-Factor Agents framework prevent infinite retry loops?
The framework uses a consecutive-error counter stored in the unified thread state to track repeated failures. As implemented in content/factor-09-compact-errors.md, the agent increments this counter on each failure and resets it to zero on success. When the counter exceeds a defined threshold (typically three), the loop terminates or escalates to a human rather than continuing indefinitely.
What should I include in the error event data?
According to content/factor-09-compact-errors.md (lines 21-28), the error event should contain a structured, formatted representation of the exception. Use a format_error() function to extract the error type, message, and relevant context while avoiding overly verbose stack traces that might waste context window space. The goal is machine-readability for the LLM while preserving diagnostic information.
Can I implement exponential back-off instead of immediate retries?
Yes. Since the error counter lives in the thread structure, you can implement any back-off policy by checking consecutive_errors before continuing. For example, you could add await asyncio.sleep(2 ** consecutive_errors) inside the except block, or pause execution entirely until external conditions change. The Own Your Control Flow factor (content/factor-08-own-your-control-flow.md) explicitly supports custom logic like back-off, circuit breakers, or conditional branching based on error counts.
Where should the retry logic live in my codebase?
The retry logic should reside in your main execution loop or control-flow layer, not inside individual tools. As shown in the examples from humanlayer/12-factor-agents, the loop calling determine_next_step() and handle_next_step() owns the try/except block. This keeps tools simple and deterministic while allowing the orchestration layer to handle recovery strategies, state tracking, and escalation logic consistently across all tool invocations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →