# Agent Error Recovery and Retry Logic in 12-Factor Agents: A Complete Guide

> Master agent error recovery and retry logic in 12-factor agents. Learn to treat tool failures as events for LLM self-correction and implement robust retry counters. Ensure agent resilience.

- Repository: [HumanLayer/12-factor-agents](https://github.com/humanlayer/12-factor-agents)
- Tags: how-to-guide
- Published: 2026-05-19

---

**Treat tool failures as structured events appended to the agent's context window, allowing the LLM to self-correct on the next iteration while using a consecutive-error counter to limit retries and enable escalation.**

When building autonomous agents in the `humanlayer/12-factor-agents` framework, handling transient failures gracefully is critical for robust operation. Instead of aborting workflows when a tool fails, the recommended pattern treats errors as first-class context, giving the LLM visibility into what went wrong so it can propose corrected actions. This approach combines structured error recording with intelligent retry limits to create self-healing agents that know when to persist and when to escalate.

## The Core Pattern: Treat Errors as Context

According to [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md), the fundamental principle is to record every error as an event in the thread's history. When a tool invocation fails, you append a formatted `error` event to `thread["events"]` rather than raising an exception that breaks the execution loop.

This keeps the LLM's context complete and machine-readable, allowing the model to see the exact error message, error type, and stack trace on the next iteration. The agent can then reason about the failure and select a different tool or adjust its parameters.

The pattern requires maintaining a **consecutive-error counter** in the same thread structure that holds all execution state. As noted in [`content/factor-05-unify-execution-state.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-05-unify-execution-state.md), keeping this counter in the unified thread state (lines 11-14) ensures all execution metadata lives in one place, simplifying serialization and debugging.

## Implementing the Retry Loop

The basic implementation follows a four-step cycle inside your control flow (as described in [`content/factor-08-own-your-control-flow.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md), lines 10-18):

1. **Determine the next step** from the current context window.
2. **Execute the step** inside a `try / except` block.
3. **On success**, record the result event and reset `consecutive_errors` to zero.
4. **On failure**, record an error event, increment the counter, and decide whether to retry or abort.

### Basic Error Handling Loop

Here is the canonical pattern from the repository showing how to structure your execution loop:

```python
thread = {"events": [initial_message]}          # initial context

while True:
    next_step = await determine_next_step(thread_to_prompt(thread))
    thread["events"].append({
        "type": next_step.intent,
        "data": next_step,
    })
    try:
        # Run the tool associated with the intent

        result = await handle_next_step(thread, next_step)
        thread["events"].append({
            "type": f"{next_step.intent}_result",
            "data": result,
        })
        consecutive_errors = 0                 # success → reset counter

    except Exception as e:
        # Record the error and decide whether to retry

        consecutive_errors += 1
        thread["events"].append({
            "type": "error",
            "data": format_error(e),
        })
        if consecutive_errors >= 3:
            # Too many failures – break or escalate

            await notify_human(thread, e)      # see factor-07

            break
        # otherwise, loop again and let the LLM pick a new tool

```

This loop demonstrates the **Compact Errors** factor (lines 21-28 in [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md)), where the `format_error()` function ensures the exception is serialized into a clean, readable format for the LLM.

### Using a Dedicated Retry Helper

For reusable logic across multiple tools, extract the retry mechanism into a helper function:

```python
MAX_RETRIES = 3

async def run_with_retry(step, thread):
    for attempt in range(1, MAX_RETRIES + 1):
        try:
            return await handle_next_step(thread, step)
        except Exception as exc:
            thread["events"].append({
                "type": "error",
                "data": format_error(exc, attempt=attempt),
            })
            if attempt == MAX_RETRIES:
                raise   # bubble up after final failure

```

Note that this approach still records each failure attempt to the thread, giving the LLM full visibility into repeated failures while enforcing a hard limit on retries.

## Escalation and Control Flow

When the consecutive-error counter exceeds your threshold (typically three attempts), you must decide whether to continue looping, reset context, or escalate to human oversight (lines 33-58 and line 62 in [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md)).

### When to Escalate to Humans

The **Contact Humans With Tools** factor ([`content/factor-07-contact-humans-with-tools.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-07-contact-humans-with-tools.md)) provides the pattern for handling persistent failures:

```python
if consecutive_errors >= MAX_RETRIES:
    await send_message_to_human(
        f"Tool `{next_step.intent}` failed {consecutive_errors} times. "
        "Human review required."
    )
    break   # pause the loop until a human resolves the issue

```

This creates a clean handoff point where the agent stops spinning on a broken operation and waits for human intervention. The thread state remains intact, allowing the human to inspect the error history and either fix the underlying issue or guide the agent forward.

### Integrating with Control Flow

As documented in [`content/factor-08-own-your-control-flow.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md), error handling should integrate with your custom control-flow logic. Errors can trigger specific branches—such as pausing execution, triggering a back-off policy, or switching to a fallback tool. The key is that the error event becomes part of the decision context, not a termination signal.

## State Management Benefits

By storing the `consecutive_errors` counter within the `thread` structure alongside all events, you achieve full **deterministic state** as described in [`content/factor-05-unify-execution-state.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-05-unify-execution-state.md). This means:

- **Serialization**: The entire execution state, including retry history, can be persisted to a database or passed between services.
- **Resumability**: If your agent process crashes, you can resume exactly where it left off, including knowledge of how many consecutive errors occurred before the crash.
- **Debugging**: Inspecting the thread reveals the complete history of attempts, failures, and retries without needing external logs.

## Summary

- **Record errors as events**: Append formatted exceptions to `thread["events"]` with type `"error"` to give the LLM visibility into failures.
- **Use a consecutive-error counter**: Maintain `consecutive_errors` in the unified thread state to prevent infinite retry loops (see [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md), lines 33-58).
- **Reset on success**: Always reset the error counter to zero after successful tool execution to distinguish between transient and persistent failures.
- **Escalate gracefully**: When `consecutive_errors` exceeds your threshold (e.g., three), break the loop and use `send_message_to_human()` to escalate (see [`content/factor-07-contact-humans-with-tools.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-07-contact-humans-with-tools.md)).
- **Keep state unified**: Store retry counters in the same `thread` structure as events to maintain deterministic, serializable execution state.

## Frequently Asked Questions

### How does the 12-Factor Agents framework prevent infinite retry loops?

The framework uses a **consecutive-error counter** stored in the unified thread state to track repeated failures. As implemented in [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md), the agent increments this counter on each failure and resets it to zero on success. When the counter exceeds a defined threshold (typically three), the loop terminates or escalates to a human rather than continuing indefinitely.

### What should I include in the error event data?

According to [`content/factor-09-compact-errors.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-09-compact-errors.md) (lines 21-28), the error event should contain a **structured, formatted representation** of the exception. Use a `format_error()` function to extract the error type, message, and relevant context while avoiding overly verbose stack traces that might waste context window space. The goal is machine-readability for the LLM while preserving diagnostic information.

### Can I implement exponential back-off instead of immediate retries?

Yes. Since the error counter lives in the `thread` structure, you can implement any back-off policy by checking `consecutive_errors` before continuing. For example, you could add `await asyncio.sleep(2 ** consecutive_errors)` inside the `except` block, or pause execution entirely until external conditions change. The **Own Your Control Flow** factor ([`content/factor-08-own-your-control-flow.md`](https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-08-own-your-control-flow.md)) explicitly supports custom logic like back-off, circuit breakers, or conditional branching based on error counts.

### Where should the retry logic live in my codebase?

The retry logic should reside in your main execution loop or control-flow layer, not inside individual tools. As shown in the examples from `humanlayer/12-factor-agents`, the loop calling `determine_next_step()` and `handle_next_step()` owns the `try/except` block. This keeps tools simple and deterministic while allowing the orchestration layer to handle recovery strategies, state tracking, and escalation logic consistently across all tool invocations.