# Error Handling and Retry Strategies for Failed ChatDev Workflow Nodes: A Complete Guide

> Master ChatDev workflow error handling and retry strategies. Learn how to configure Tenacity-based loops to automatically retry failed nodes based on status codes, exceptions, and messages.

- Repository: [OpenBMB/ChatDev](https://github.com/OpenBMB/ChatDev)
- Tags: how-to-guide
- Published: 2026-04-01

---

**ChatDev automatically retries failed model provider calls using a configurable Tenacity-based loop that inspects HTTP status codes, exception types, and error message substrings to determine whether to retry, with all policies definable per-node in YAML configuration.**

ChatDev is an open-source framework for orchestrating LLM-powered software development workflows. When agent nodes invoke external model providers like OpenAI or Gemini, transient failures such as rate limits, network timeouts, or server errors can interrupt execution. Understanding the error handling and retry strategies for failed ChatDev workflow nodes ensures your pipelines remain resilient without manual intervention.

## How ChatDev Implements Retry Logic at the Node Level

The retry mechanism is encapsulated in the `AgentNodeExecutor` class, which wraps every model provider call in a robust retry loop using the Tenacity library.

### The Core Retry Wrapper in AgentNodeExecutor

In [`runtime/node/executor/agent_executor.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/node/executor/agent_executor.py), the method `_execute_with_retry` (line 394) implements the retry loop:

```python
retrying = Retrying(
    stop=stop_after_attempt(retry_config.max_attempts),
    wait=wait_random_exponential(
        min=retry_config.min_wait_seconds,
        max=retry_config.max_wait_seconds
    ),
    retry=retry_if_exception(lambda exc: retry_config.should_retry(exc)),
    before_sleep=self._before_sleep,
    reraise=True
)

```

This wrapper delegates the retry decision to `retry_config.should_retry`, which inspects the exception to determine if it matches configurable retry criteria. The `_before_sleep` callback (lines 410-424) logs each failed attempt via the `log_manager` at warning level, capturing the node ID and exception details before the next retry attempt.

### Default Policy Resolution

When a node configuration omits retry settings, `_resolve_retry_policy` (line 434) automatically injects a default `AgentRetryConfig`. This guarantees every agent node has a retry policy even when users do not explicitly define one in the workflow YAML. The default configuration allows **5 maximum attempts** with exponential backoff between **1 and 6 seconds**, and includes a broad list of retryable HTTP status codes and exception types.

## Configuring Retry Policies in Workflow YAML Files

Each agent node can declare a `retry` block that maps directly to the `AgentRetryConfig` dataclass defined in [`entity/configs/node/agent.py`](https://github.com/OpenBMB/ChatDev/blob/main/entity/configs/node/agent.py) (line 106).

```yaml
type: agent
model:
  provider: openai
  name: gpt-4o
retry:
  enabled: true
  max_attempts: 4
  min_wait_seconds: 1.0
  max_wait_seconds: 8.0
  retry_on_status_codes: [429, 502, 503]
  retry_on_exception_types:
    - rate_limit_error
    - timeouterror
    - connectionerror
  non_retry_exception_types:
    - authenticationerror
  retry_on_error_substrings:
    - "temporarily unavailable"
    - "rate limit exceeded"
    - "connection reset"

```

### Key Configuration Parameters

- **enabled**: Boolean toggle to activate or deactivate the retry mechanism for this specific node.
- **max_attempts**: Total number of attempts including the initial call before the workflow treats the failure as terminal.
- **min_wait_seconds / max_wait_seconds**: Bounds for the `wait_random_exponential` function, which randomizes backoff time between attempts to prevent thundering herds.
- **retry_on_status_codes**: List of HTTP status codes that should trigger a retry (commonly 429, 502, 503, 504).
- **retry_on_exception_types**: Case-insensitive class names of exceptions that warrant retries, such as network timeouts or rate limit errors.
- **non_retry_exception_types**: Explicit exceptions that must fail immediately without retry, such as authentication errors.
- **retry_on_error_substrings**: Textual patterns within error messages that indicate transient failures, allowing retries on custom provider error messages.

## Understanding the Retry Decision Logic

The `should_retry` method in [`entity/configs/node/agent.py`](https://github.com/OpenBMB/ChatDev/blob/main/entity/configs/node/agent.py) (line 41) implements the decision logic that determines whether a specific exception warrants a retry attempt.

### Exception Chain Inspection

When a model call fails, `should_retry` examines the exception chain for three signals:

1. **HTTP Status Codes**: If the exception contains an HTTP response code, it checks against `retry_on_status_codes`.
2. **Exception Class Names**: It compares the exception type name (case-insensitive) against `retry_on_exception_types`.
3. **Message Substrings**: It searches the exception message for substrings defined in `retry_on_error_substrings`.

If any check matches and the exception is not explicitly listed in `non_retry_exception_types`, the retry loop continues. If the retry loop exhausts all attempts, the exception propagates to the outer `try/except` block in the `execute` method (lines 667-674), which logs the final error and returns a user-visible error message as a node response.

## Monitoring and Logging Failed Attempts

ChatDev provides comprehensive observability into retry behavior through the `log_manager` integration.

### Per-Attempt Logging

Before each retry, the `before_sleep` callback emits a warning log entry:

```python
self.log_manager.warning(
    f"[Node: {node.id}] Model call attempt {attempt} failed: {exc}",
    node_id=node.id,
    details=details,
)

```

This creates an audit trail showing exactly which node failed, which attempt number failed, and the specific exception encountered.

### Terminal Failure Handling

If all retry attempts are exhausted, the exception bubbles up to the outer exception handler in `AgentNodeExecutor.execute`:

```python
traceback.print_exc()
error_msg = f"[Node: {node.id}] Error calling model: {str(e)}"
self.log_manager.error(error_msg)

```

This logs the full stack trace and generates a user-facing error message that appears as a regular assistant message in the workflow output, preventing hard crashes while clearly communicating the failure.

## Extending Retry Strategies to Other Node Types

The retry configuration is not limited to agent nodes. The `PythonNodeExecutor` in [`runtime/node/executor/python_executor.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/node/executor/python_executor.py) reuses the same `AgentRetryConfig` dataclass for executing Python code nodes that may call external APIs. Because the retry logic lives in discrete utility methods like `_execute_with_retry`, extending resilient error handling to new node types requires only invoking the same helper with an appropriate configuration object.

## Summary

- **All agent nodes have automatic retry protection** via `AgentNodeExecutor._execute_with_retry`, using either user-defined YAML configuration or sensible defaults.
- **Retry policies inspect HTTP status codes, exception class names, and message substrings** through the `should_retry` method to distinguish transient from terminal failures.
- **Exponential backoff with jitter** prevents overwhelming upstream providers while maximizing the chance of recovery.
- **Comprehensive logging** captures every failed attempt via the `log_manager`, with terminal failures generating user-visible error messages rather than crashing the workflow.
- **The same configuration dataclass** supports multiple executor types, ensuring consistent retry semantics across agents and Python nodes.

## Frequently Asked Questions

### How do I disable retries for a specific ChatDev node?

Set `enabled: false` in the retry configuration block of your node YAML. When disabled, the executor bypasses the Tenacity retry loop and immediately propagates any exceptions to the outer error handler, causing the node to fail fast without retry attempts.

### What is the default retry behavior if I don't specify a retry block?

ChatDev automatically injects a default `AgentRetryConfig` via `_resolve_retry_policy` with **5 maximum attempts**, exponential backoff between **1 and 6 seconds**, and a broad list of retryable status codes (including 429, 502, 503, 504) and network-related exception types. This ensures basic resilience without requiring explicit configuration.

### Can I retry only on specific HTTP status codes like 429 but ignore 500 errors?

Yes. Configure `retry_on_status_codes: [429]` and leave `retry_on_exception_types` empty or undefined in your YAML. The `should_retry` method will only retry when the provider returns HTTP 429 (rate limit), while treating HTTP 500 and other errors as immediate failures that bypass the retry loop.

### Where can I see logs of retry attempts in ChatDev?

Retry attempts are logged at the **warning** level via the `log_manager` in the `before_sleep` callback (lines 410-424 of [`agent_executor.py`](https://github.com/OpenBMB/ChatDev/blob/main/agent_executor.py)), appearing in your configured log destination (console, file, or external pipeline) with the node ID and attempt number. Successful calls are logged at the **debug** level, while exhausted retries emit an **error** level log with the full stack trace.