# How GenericAgent Implements Error Recovery and Retry Logic in Its Agent Loop

> Discover how GenericAgent uses its LLM client for error recovery and retry logic with exponential backoff. Learn how it prevents agent crashes by treating residual errors as text.

- Repository: [LJQ/GenericAgent](https://github.com/lsdefine/GenericAgent)
- Tags: how-to-guide
- Published: 2026-04-16

---

**GenericAgent delegates error recovery to its LLM client in [`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py), which implements exponential backoff with jitter for retryable HTTP status codes, while the main `agent_runner_loop` treats any residual errors as regular text to ensure the agent never crashes mid-execution.**

The GenericAgent repository provides a lightweight, extensible framework for building LLM-driven agents. Understanding how it handles transient failures—such as rate limits or network timeouts—is critical for running production workloads. This article examines the specific implementation of error recovery and retry logic within the GenericAgent architecture, tracing the path from the high-level agent loop down to the resilient HTTP client.

## Where Error Recovery Lives in the GenericAgent Architecture

Error recovery in GenericAgent is intentionally separated between two core modules. The **agent loop** ([`agent_loop.py`](https://github.com/lsdefine/GenericAgent/blob/main/agent_loop.py)) focuses on orchestration and tool execution, while the **LLM client** ([`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py)) owns all network resilience.

In [`agent_loop.py`](https://github.com/lsdefine/GenericAgent/blob/main/agent_loop.py), the function `agent_runner_loop` manages the conversation turn-by-turn. It does not contain explicit try-catch blocks for HTTP errors. Instead, it consumes a generator returned by `client.chat()`, trusting that the client handles retries internally. If the LLM client ultimately fails after exhausting retries, it yields an error block that the loop treats as normal assistant text.

## The LLM Client Retry Mechanism in llmcore.py

The retry logic resides in [`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py), specifically within the `_openai_stream` helper (and its Claude counterpart `_parse_claude_sse`). This function wraps the raw HTTP request to OpenAI or Anthropic APIs with a robust retry strategy.

### Retryable HTTP Status Codes and Network Failures

The client defines a constant set of retryable status codes that indicate transient failures:

```python
RETRYABLE = {408, 409, 425, 429, 500, 502, 503, 504}

```

These cover request timeouts, rate limits (`429`), and various server errors (`5xx`). Additionally, the client catches network-level exceptions `(requests.Timeout, requests.ConnectionError)` and subjects them to the same retry logic.

### Exponential Backoff with Retry-After Header Support

When a retryable error occurs, the client computes a delay using the private `_delay` function:

```python
def _delay(resp, attempt):
    try: 
        ra = float((resp.headers or {}).get("retry-after"))
    except: 
        ra = None
    return max(0.5, ra if ra is not None else min(30.0, 1.5 * (2 ** attempt)))

```

This implements **exponential backoff** with a base of `1.5` and a cap of `30` seconds. If the API returns a `Retry-After` header, the client respects that value instead of calculating its own delay, ensuring compliance with rate-limit directives.

### Maximum Retry Limits and Graceful Degradation

The caller can specify `max_retries` (defaulting to `0`). The loop attempts the request up to `max_retries + 1` times:

```python
for attempt in range(max_retries + 1):
    try:
        with requests.post(url, headers=headers, json=payload,
                           stream=True,
                           timeout=(connect_timeout, read_timeout),
                           proxies=proxies) as r:
            if r.status_code >= 400:
                if r.status_code in RETRYABLE and attempt < max_retries:
                    d = _delay(r, attempt)
                    print(f"[LLM Retry] HTTP {r.status_code}, retry in {d:.1f}s "
                          f"({attempt+1}/{max_retries+1})")
                    time.sleep(d)
                    continue
                # Non-retryable error handling...

    except (requests.Timeout, requests.ConnectionError) as e:
        if attempt < max_retries:
            # ... retry logic ...

            continue
        # Final failure: yield error block instead of crashing

        err = f"Error: {type(e).__name__}: {e}"
        yield err
        return [{"type": "text", "text": err}]

```

If all retries are exhausted, the client yields a formatted error string and returns a text block. This **graceful degradation** ensures that `agent_runner_loop` receives a valid generator item rather than an unhandled exception, allowing the agent to decide how to proceed.

## How the Agent Loop Handles Persistent Failures

The `agent_runner_loop` in [`agent_loop.py`](https://github.com/lsdefine/GenericAgent/blob/main/agent_loop.py) consumes the generator returned by `client.chat()`:

```python
response_gen = client.chat(messages=messages, tools=tools_schema)
if verbose:
    response = yield from response_gen          # streams text & errors

else:
    response = exhaust(response_gen)            # collect all at once

```

Because the LLM client guarantees that the generator always yields valid content (even if it is an error message), the loop treats the response as normal assistant text. If the response contains no tool calls, the loop continues to the next turn or exits based on the `StepOutcome` dataclass, which signals `should_exit` or provides a `next_prompt`.

The loop also enforces a `max_turns` parameter, providing an upper bound on the entire interaction to prevent infinite loops if persistent failures occur.

## Summary

- **Error recovery is layered**: The LLM client ([`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py)) handles transient network and HTTP errors, while the agent loop ([`agent_loop.py`](https://github.com/lsdefine/GenericAgent/blob/main/agent_loop.py)) handles final failures as content.
- **Retry logic uses exponential backoff**: The `_delay` function in [`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py) implements capped exponential backoff with `1.5 * (2 ** attempt)` and respects `Retry-After` headers.
- **Retryable status codes are explicitly defined**: `RETRYABLE = {408, 409, 425, 429, 500, 502, 503, 504}` covers transient server and rate-limit errors.
- **Graceful degradation prevents crashes**: After exhausting `max_retries`, the client yields an error string rather than raising an exception, allowing `agent_runner_loop` to continue.
- **Safety limits prevent infinite loops**: The `max_turns` parameter in the agent loop caps total interaction length.

## Frequently Asked Questions

### How does GenericAgent handle rate limiting from OpenAI or Anthropic APIs?

GenericAgent handles rate limiting through the `_openai_stream` function in [`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py). When it receives an HTTP `429` status code (which is included in the `RETRYABLE` set), it checks for a `Retry-After` header. If present, it waits exactly that many seconds; otherwise, it falls back to exponential backoff with a base of 1.5 seconds. This ensures compliance with provider rate limits while minimizing unnecessary delays.

### What happens when the LLM client exhausts all retry attempts?

When the maximum retry count is reached, the LLM client does not raise an exception that would crash the agent. Instead, [`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py) yields a formatted error string (for example, `"Error: requests.Timeout: ..."`) and returns a text block structure. The `agent_runner_loop` in [`agent_loop.py`](https://github.com/lsdefine/GenericAgent/blob/main/agent_loop.py) consumes this as regular assistant content, allowing the handler to inspect the failure and decide whether to terminate via `StepOutcome.should_exit` or attempt recovery with a `next_prompt`.

### Can developers configure the retry behavior for specific HTTP errors?

Yes, developers can configure retry behavior through the `max_retries` parameter passed to the LLM client's `chat` method, which defaults to `0` (no retries). While the set of `RETRYABLE` status codes is hardcoded in [`llmcore.py`](https://github.com/lsdefine/GenericAgent/blob/main/llmcore.py) as `{408, 409, 425, 429, 500, 502, 503, 504}`, the exponential backoff timing and respect for `Retry-After` headers apply automatically to all retryable errors. Developers can also wrap the client call in custom logic if they need to handle specific non-retryable errors differently.

### How does the agent loop prevent infinite loops during persistent LLM failures?

The `agent_runner_loop` enforces a `max_turns` parameter that sets an absolute upper bound on the number of interaction turns. Even if the LLM client returns error blocks repeatedly (treating them as regular content), the loop increments its turn counter each iteration. Once `max_turns` is reached, the loop terminates regardless of the conversation state, preventing infinite resource consumption when persistent network or service failures occur.