Agent Error Recovery and Retry Logic in 12-Factor Agents: A Complete Guide

Treat tool failures as structured events appended to the agent's context window, allowing the LLM to self-correct on the next iteration while using a consecutive-error counter to limit retries and enable escalation.

When building autonomous agents in the humanlayer/12-factor-agents framework, handling transient failures gracefully is critical for robust operation. Instead of aborting workflows when a tool fails, the recommended pattern treats errors as first-class context, giving the LLM visibility into what went wrong so it can propose corrected actions. This approach combines structured error recording with intelligent retry limits to create self-healing agents that know when to persist and when to escalate.

The Core Pattern: Treat Errors as Context

According to content/factor-09-compact-errors.md, the fundamental principle is to record every error as an event in the thread's history. When a tool invocation fails, you append a formatted error event to thread["events"] rather than raising an exception that breaks the execution loop.

This keeps the LLM's context complete and machine-readable, allowing the model to see the exact error message, error type, and stack trace on the next iteration. The agent can then reason about the failure and select a different tool or adjust its parameters.

The pattern requires maintaining a consecutive-error counter in the same thread structure that holds all execution state. As noted in content/factor-05-unify-execution-state.md, keeping this counter in the unified thread state (lines 11-14) ensures all execution metadata lives in one place, simplifying serialization and debugging.

Implementing the Retry Loop

The basic implementation follows a four-step cycle inside your control flow (as described in content/factor-08-own-your-control-flow.md, lines 10-18):

  1. Determine the next step from the current context window.
  2. Execute the step inside a try / except block.
  3. On success, record the result event and reset consecutive_errors to zero.
  4. On failure, record an error event, increment the counter, and decide whether to retry or abort.

Basic Error Handling Loop

Here is the canonical pattern from the repository showing how to structure your execution loop:

thread = {"events": [initial_message]}          # initial context

while True:
    next_step = await determine_next_step(thread_to_prompt(thread))
    thread["events"].append({
        "type": next_step.intent,
        "data": next_step,
    })
    try:
        # Run the tool associated with the intent

        result = await handle_next_step(thread, next_step)
        thread["events"].append({
            "type": f"{next_step.intent}_result",
            "data": result,
        })
        consecutive_errors = 0                 # success → reset counter

    except Exception as e:
        # Record the error and decide whether to retry

        consecutive_errors += 1
        thread["events"].append({
            "type": "error",
            "data": format_error(e),
        })
        if consecutive_errors >= 3:
            # Too many failures – break or escalate

            await notify_human(thread, e)      # see factor-07

            break
        # otherwise, loop again and let the LLM pick a new tool

This loop demonstrates the Compact Errors factor (lines 21-28 in content/factor-09-compact-errors.md), where the format_error() function ensures the exception is serialized into a clean, readable format for the LLM.

Using a Dedicated Retry Helper

For reusable logic across multiple tools, extract the retry mechanism into a helper function:

MAX_RETRIES = 3

async def run_with_retry(step, thread):
    for attempt in range(1, MAX_RETRIES + 1):
        try:
            return await handle_next_step(thread, step)
        except Exception as exc:
            thread["events"].append({
                "type": "error",
                "data": format_error(exc, attempt=attempt),
            })
            if attempt == MAX_RETRIES:
                raise   # bubble up after final failure

Note that this approach still records each failure attempt to the thread, giving the LLM full visibility into repeated failures while enforcing a hard limit on retries.

Escalation and Control Flow

When the consecutive-error counter exceeds your threshold (typically three attempts), you must decide whether to continue looping, reset context, or escalate to human oversight (lines 33-58 and line 62 in content/factor-09-compact-errors.md).

When to Escalate to Humans

The Contact Humans With Tools factor (content/factor-07-contact-humans-with-tools.md) provides the pattern for handling persistent failures:

if consecutive_errors >= MAX_RETRIES:
    await send_message_to_human(
        f"Tool `{next_step.intent}` failed {consecutive_errors} times. "
        "Human review required."
    )
    break   # pause the loop until a human resolves the issue

This creates a clean handoff point where the agent stops spinning on a broken operation and waits for human intervention. The thread state remains intact, allowing the human to inspect the error history and either fix the underlying issue or guide the agent forward.

Integrating with Control Flow

As documented in content/factor-08-own-your-control-flow.md, error handling should integrate with your custom control-flow logic. Errors can trigger specific branches—such as pausing execution, triggering a back-off policy, or switching to a fallback tool. The key is that the error event becomes part of the decision context, not a termination signal.

State Management Benefits

By storing the consecutive_errors counter within the thread structure alongside all events, you achieve full deterministic state as described in content/factor-05-unify-execution-state.md. This means:

  • Serialization: The entire execution state, including retry history, can be persisted to a database or passed between services.
  • Resumability: If your agent process crashes, you can resume exactly where it left off, including knowledge of how many consecutive errors occurred before the crash.
  • Debugging: Inspecting the thread reveals the complete history of attempts, failures, and retries without needing external logs.

Summary

  • Record errors as events: Append formatted exceptions to thread["events"] with type "error" to give the LLM visibility into failures.
  • Use a consecutive-error counter: Maintain consecutive_errors in the unified thread state to prevent infinite retry loops (see content/factor-09-compact-errors.md, lines 33-58).
  • Reset on success: Always reset the error counter to zero after successful tool execution to distinguish between transient and persistent failures.
  • Escalate gracefully: When consecutive_errors exceeds your threshold (e.g., three), break the loop and use send_message_to_human() to escalate (see content/factor-07-contact-humans-with-tools.md).
  • Keep state unified: Store retry counters in the same thread structure as events to maintain deterministic, serializable execution state.

Frequently Asked Questions

How does the 12-Factor Agents framework prevent infinite retry loops?

The framework uses a consecutive-error counter stored in the unified thread state to track repeated failures. As implemented in content/factor-09-compact-errors.md, the agent increments this counter on each failure and resets it to zero on success. When the counter exceeds a defined threshold (typically three), the loop terminates or escalates to a human rather than continuing indefinitely.

What should I include in the error event data?

According to content/factor-09-compact-errors.md (lines 21-28), the error event should contain a structured, formatted representation of the exception. Use a format_error() function to extract the error type, message, and relevant context while avoiding overly verbose stack traces that might waste context window space. The goal is machine-readability for the LLM while preserving diagnostic information.

Can I implement exponential back-off instead of immediate retries?

Yes. Since the error counter lives in the thread structure, you can implement any back-off policy by checking consecutive_errors before continuing. For example, you could add await asyncio.sleep(2 ** consecutive_errors) inside the except block, or pause execution entirely until external conditions change. The Own Your Control Flow factor (content/factor-08-own-your-control-flow.md) explicitly supports custom logic like back-off, circuit breakers, or conditional branching based on error counts.

Where should the retry logic live in my codebase?

The retry logic should reside in your main execution loop or control-flow layer, not inside individual tools. As shown in the examples from humanlayer/12-factor-agents, the loop calling determine_next_step() and handle_next_step() owns the try/except block. This keeps tools simple and deterministic while allowing the orchestration layer to handle recovery strategies, state tracking, and escalation logic consistently across all tool invocations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →