How GenericAgent Implements Error Recovery and Retry Logic in Its Agent Loop
GenericAgent delegates error recovery to its LLM client in llmcore.py, which implements exponential backoff with jitter for retryable HTTP status codes, while the main agent_runner_loop treats any residual errors as regular text to ensure the agent never crashes mid-execution.
The GenericAgent repository provides a lightweight, extensible framework for building LLM-driven agents. Understanding how it handles transient failures—such as rate limits or network timeouts—is critical for running production workloads. This article examines the specific implementation of error recovery and retry logic within the GenericAgent architecture, tracing the path from the high-level agent loop down to the resilient HTTP client.
Where Error Recovery Lives in the GenericAgent Architecture
Error recovery in GenericAgent is intentionally separated between two core modules. The agent loop (agent_loop.py) focuses on orchestration and tool execution, while the LLM client (llmcore.py) owns all network resilience.
In agent_loop.py, the function agent_runner_loop manages the conversation turn-by-turn. It does not contain explicit try-catch blocks for HTTP errors. Instead, it consumes a generator returned by client.chat(), trusting that the client handles retries internally. If the LLM client ultimately fails after exhausting retries, it yields an error block that the loop treats as normal assistant text.
The LLM Client Retry Mechanism in llmcore.py
The retry logic resides in llmcore.py, specifically within the _openai_stream helper (and its Claude counterpart _parse_claude_sse). This function wraps the raw HTTP request to OpenAI or Anthropic APIs with a robust retry strategy.
Retryable HTTP Status Codes and Network Failures
The client defines a constant set of retryable status codes that indicate transient failures:
RETRYABLE = {408, 409, 425, 429, 500, 502, 503, 504}
These cover request timeouts, rate limits (429), and various server errors (5xx). Additionally, the client catches network-level exceptions (requests.Timeout, requests.ConnectionError) and subjects them to the same retry logic.
Exponential Backoff with Retry-After Header Support
When a retryable error occurs, the client computes a delay using the private _delay function:
def _delay(resp, attempt):
try:
ra = float((resp.headers or {}).get("retry-after"))
except:
ra = None
return max(0.5, ra if ra is not None else min(30.0, 1.5 * (2 ** attempt)))
This implements exponential backoff with a base of 1.5 and a cap of 30 seconds. If the API returns a Retry-After header, the client respects that value instead of calculating its own delay, ensuring compliance with rate-limit directives.
Maximum Retry Limits and Graceful Degradation
The caller can specify max_retries (defaulting to 0). The loop attempts the request up to max_retries + 1 times:
for attempt in range(max_retries + 1):
try:
with requests.post(url, headers=headers, json=payload,
stream=True,
timeout=(connect_timeout, read_timeout),
proxies=proxies) as r:
if r.status_code >= 400:
if r.status_code in RETRYABLE and attempt < max_retries:
d = _delay(r, attempt)
print(f"[LLM Retry] HTTP {r.status_code}, retry in {d:.1f}s "
f"({attempt+1}/{max_retries+1})")
time.sleep(d)
continue
# Non-retryable error handling...
except (requests.Timeout, requests.ConnectionError) as e:
if attempt < max_retries:
# ... retry logic ...
continue
# Final failure: yield error block instead of crashing
err = f"Error: {type(e).__name__}: {e}"
yield err
return [{"type": "text", "text": err}]
If all retries are exhausted, the client yields a formatted error string and returns a text block. This graceful degradation ensures that agent_runner_loop receives a valid generator item rather than an unhandled exception, allowing the agent to decide how to proceed.
How the Agent Loop Handles Persistent Failures
The agent_runner_loop in agent_loop.py consumes the generator returned by client.chat():
response_gen = client.chat(messages=messages, tools=tools_schema)
if verbose:
response = yield from response_gen # streams text & errors
else:
response = exhaust(response_gen) # collect all at once
Because the LLM client guarantees that the generator always yields valid content (even if it is an error message), the loop treats the response as normal assistant text. If the response contains no tool calls, the loop continues to the next turn or exits based on the StepOutcome dataclass, which signals should_exit or provides a next_prompt.
The loop also enforces a max_turns parameter, providing an upper bound on the entire interaction to prevent infinite loops if persistent failures occur.
Summary
- Error recovery is layered: The LLM client (
llmcore.py) handles transient network and HTTP errors, while the agent loop (agent_loop.py) handles final failures as content. - Retry logic uses exponential backoff: The
_delayfunction inllmcore.pyimplements capped exponential backoff with1.5 * (2 ** attempt)and respectsRetry-Afterheaders. - Retryable status codes are explicitly defined:
RETRYABLE = {408, 409, 425, 429, 500, 502, 503, 504}covers transient server and rate-limit errors. - Graceful degradation prevents crashes: After exhausting
max_retries, the client yields an error string rather than raising an exception, allowingagent_runner_loopto continue. - Safety limits prevent infinite loops: The
max_turnsparameter in the agent loop caps total interaction length.
Frequently Asked Questions
How does GenericAgent handle rate limiting from OpenAI or Anthropic APIs?
GenericAgent handles rate limiting through the _openai_stream function in llmcore.py. When it receives an HTTP 429 status code (which is included in the RETRYABLE set), it checks for a Retry-After header. If present, it waits exactly that many seconds; otherwise, it falls back to exponential backoff with a base of 1.5 seconds. This ensures compliance with provider rate limits while minimizing unnecessary delays.
What happens when the LLM client exhausts all retry attempts?
When the maximum retry count is reached, the LLM client does not raise an exception that would crash the agent. Instead, llmcore.py yields a formatted error string (for example, "Error: requests.Timeout: ...") and returns a text block structure. The agent_runner_loop in agent_loop.py consumes this as regular assistant content, allowing the handler to inspect the failure and decide whether to terminate via StepOutcome.should_exit or attempt recovery with a next_prompt.
Can developers configure the retry behavior for specific HTTP errors?
Yes, developers can configure retry behavior through the max_retries parameter passed to the LLM client's chat method, which defaults to 0 (no retries). While the set of RETRYABLE status codes is hardcoded in llmcore.py as {408, 409, 425, 429, 500, 502, 503, 504}, the exponential backoff timing and respect for Retry-After headers apply automatically to all retryable errors. Developers can also wrap the client call in custom logic if they need to handle specific non-retryable errors differently.
How does the agent loop prevent infinite loops during persistent LLM failures?
The agent_runner_loop enforces a max_turns parameter that sets an absolute upper bound on the number of interaction turns. Even if the LLM client returns error blocks repeatedly (treating them as regular content), the loop increments its turn counter each iteration. Once max_turns is reached, the loop terminates regardless of the conversation state, preventing infinite resource consumption when persistent network or service failures occur.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →