# How to Use the Forge API Retry Mechanism for Reliable Batch Inference

> Learn to use the Forge API retry mechanism for reliable batch inference. Process thousands of sequences automatically adapting to API rate limits with tenacity and AIMD rate limiting.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: how-to-guide
- Published: 2026-05-30

---

**The Forge API retry mechanism combines a lightweight tenacity-based retry decorator for individual calls with a sophisticated batch executor that implements AIMD rate limiting, allowing you to process thousands of sequences reliably while automatically adapting to API rate limits.**

The Biohub/esm repository provides a production-ready SDK for the Forge/Biohub Platform that handles transient HTTP errors without manual intervention. When running large-scale batch inference, the `batch_executor` context manager in [`esm/sdk/__init__.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/__init__.py) coordinates a thread pool with an additive-increase/multiplicative-decrease (AIMD) algorithm to maximize throughput while respecting rate limits.

## Understanding the Dual-Layer Retry Architecture

The SDK implements retries at two levels: individual API calls use a decorator-based approach, while the batch executor orchestrates retries across concurrent workers. Understanding how these layers interact is crucial for configuring reliable inference pipelines.

### Per-Call Retry Decorator

Individual Forge clients (`ESM3ForgeInferenceClient`, `SequenceStructureForgeInferenceClient`, etc.) automatically retry transient HTTP errors through the `@retry_decorator` defined in [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py). This decorator uses tenacity to catch specific error codes—**429, 500, 502, and 504**—and applies exponential backoff between attempts.

However, when processing batches, you want to disable this per-call retry logic. The SDK exposes `skip_retries_var`, a `contextvars.ContextVar` defined in [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py), that signals the retry decorator to bypass its logic when set to `True`.

### Batch-Level Orchestration with AIMD

The `ForgeBatchExecutor` class in [`esm/utils/forge_context_manager.py`](https://github.com/Biohub/esm/blob/main/esm/utils/forge_context_manager.py) takes over retry management when `skip_retries_var` is enabled. It spawns a `ThreadPoolExecutor` and implements an **AIMD rate limiter** that:
- Halves concurrency when rate-limit errors (429) are observed
- Increments concurrency by `step_up` (default 1) when batches complete successfully
- Tracks per-task attempt counts up to `max_attempts` (default 10)

This approach prevents thundering-herd problems while maintaining optimal throughput as API conditions change.

## Implementing Reliable Batch Inference

The `batch_executor` factory function provides the primary interface for production batch workloads. It returns a configured `ForgeBatchExecutor` that handles queuing, thread management, and error classification automatically.

### Basic Batch Execution with Automatic Retries

The standard pattern involves creating a Forge client, preparing your input lists, and executing within the batch executor context:

```python
from esm.sdk import esmc_client, batch_executor
from esm.sdk.api import GenerationConfig

# Initialize client (reads ESM_API_KEY from environment)

client = esmc_client(model="esmc-600m-2024-12")

# Prepare equal-length lists of inputs and configurations

prompts = ["MKTAYIAKQRQISFVKSHFSRQDILDLW...", "MELKS..."]
configs = [GenerationConfig(num_steps=128), GenerationConfig(num_steps=128)]

# Execute with retry mechanism enabled

with batch_executor(max_attempts=12, show_progress=True) as executor:
    results = executor.execute_batch(client.batch_generate, prompts, configs)

# Handle results: either ESMProtein objects or ESMProteinError exceptions

for i, result in enumerate(results):
    if isinstance(result, Exception):
        print(f"Prompt {i} failed: {result}")
    else:
        print(f"Prompt {i} succeeded: {len(result.sequence)} residues")

```

When `batch_executor.__enter__` executes, it sets `skip_retries_var` to `True`, disabling the per-call retry logic in [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py). The executor then validates input lengths, initializes the `AIMDRateLimiter` with default concurrency of 32, and begins processing.

### Async Client Integration

If you prefer async APIs (`async_generate`, `async_fold`), wrap the coroutine with `asyncio.run` inside a blocking wrapper function:

```python
import asyncio
from esm.sdk import client as esm3_client, batch_executor
from esm.sdk.api import GenerationConfig

client = esm3_client(model="esm3-sm-open-v1")
prompts = ["MKT...", "MEK..."]
configs = [GenerationConfig(num_steps=64), GenerationConfig(num_steps=64)]

async def async_generate_one(prompt, config):
    return await client.async_generate(prompt, config)

with batch_executor(max_attempts=8) as executor:
    def blocking_wrapper(prompt, config):
        return asyncio.run(async_generate_one(prompt, config))
    
    results = executor.execute_batch(blocking_wrapper, prompts, configs)

```

The executor treats the wrapper as a standard blocking function while the underlying async client maintains connection pooling benefits.

### Configuration and Tuning

You can customize concurrency limits and retry behavior by modifying the executor's rate limiter directly:

```python
with batch_executor(max_attempts=5, show_progress=False) as executor:
    # Reduce maximum concurrent requests to 16

    executor.rate_limiter.max_concurrency = 16
    
    results = executor.execute_batch(
        client.fold, 
        sequences, 
        msa=None, 
        config=FoldingConfig()
    )

```

The `max_attempts` parameter controls per-item retry limits, while `show_progress` toggles the `tqdm` progress bar displaying success, failure, and retry counts.

## Key Implementation Details

Several components work together to provide the resilient batch processing behavior.

### The skip_retries_var Context Variable

Located in [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py), this `ContextVar` acts as a circuit breaker. When the batch executor sets it to `True`, the `retry_decorator` checks this value and immediately yields execution to the wrapped function, delegating all retry responsibility to the batch executor's AIMD-controlled loop.

### AIMDRateLimiter Mechanics

The `AIMDRateLimiter` class in [`esm/utils/forge_context_manager.py`](https://github.com/Biohub/esm/blob/main/esm/utils/forge_context_manager.py) maintains a `current_concurrency` value that dynamically adjusts based on observed error rates. When `execute_batch` detects an `ESMProteinError` with code 429, it calls `adjust_concurrency(error_seen=True)`, which multiplies the current limit by 0.5. Successful batches trigger `adjust_concurrency(error_seen=False)`, incrementing the limit by 1 until reaching the maximum.

### Error Classification

The `retry_if_specific_error` predicate in [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py) identifies retriable failures by checking the `error_code` attribute of `ESMProteinError` exceptions against the set `{429, 500, 502, 504}`. Only errors matching these codes trigger re-queueing; permanent errors (400, 401, 403) fail immediately after the first attempt.

## Summary

- **The Forge API retry mechanism** uses a dual-layer approach: automatic per-call retries for simple operations and `batch_executor` for coordinated batch processing.
- **The `batch_executor`** context manager disables per-call retries via `skip_retries_var` and manages a thread pool with AIMD rate limiting.
- **Transient error codes** (429, 500, 502, 504) trigger automatic retry with exponential backoff and concurrency adjustment.
- **Core files** include [`esm/utils/forge_context_manager.py`](https://github.com/Biohub/esm/blob/main/esm/utils/forge_context_manager.py) (executor implementation), [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py) (retry logic and context variables), and [`esm/sdk/__init__.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/__init__.py) (public API).
- **Configure reliability** using `max_attempts` (default 10) and dynamic `max_concurrency` adjustments based on API response patterns.

## Frequently Asked Questions

### What is the difference between per-call retries and the batch executor?

Individual API calls use the `@retry_decorator` in [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py) with exponential backoff for transient errors. The batch executor disables this per-call logic via `skip_retries_var` and implements its own retry loop with AIMD rate limiting, which is more efficient for concurrent batch processing because it adjusts global concurrency based on observed error rates rather than retrying independently.

### How does the AIMD rate limiter prevent API throttling?

The `AIMDRateLimiter` in [`esm/utils/forge_context_manager.py`](https://github.com/Biohub/esm/blob/main/esm/utils/forge_context_manager.py) starts with a default concurrency of 32 and applies multiplicative decrease (halving concurrency) when it encounters 429 errors, then additive increase (stepping up by 1) during successful batches. This quickly backs off under load and gradually probes for increased capacity when the API is healthy, preventing sustained rate limit violations.

### Which HTTP error codes trigger the retry mechanism?

According to the `retry_if_specific_error` predicate in [`esm/sdk/retry.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/retry.py), the SDK retries requests that fail with HTTP status codes **429** (Too Many Requests), **500** (Internal Server Error), **502** (Bad Gateway), and **504** (Gateway Timeout). Client errors like 400 (Bad Request) or 401 (Unauthorized) fail immediately without retry.

### How do I adjust the maximum number of retry attempts for batch jobs?

Pass the `max_attempts` parameter (default 10) to the `batch_executor` factory: `with batch_executor(max_attempts=12) as executor:`. The `execute_batch` method tracks attempts per task and surfaces an `ESMProteinError` for items that exceed this limit, allowing your application to handle permanent failures gracefully while successfully processing retriable ones.