How to Use the Forge API Retry Mechanism for Reliable Batch Inference

The Forge API retry mechanism combines a lightweight tenacity-based retry decorator for individual calls with a sophisticated batch executor that implements AIMD rate limiting, allowing you to process thousands of sequences reliably while automatically adapting to API rate limits.

The Biohub/esm repository provides a production-ready SDK for the Forge/Biohub Platform that handles transient HTTP errors without manual intervention. When running large-scale batch inference, the batch_executor context manager in esm/sdk/__init__.py coordinates a thread pool with an additive-increase/multiplicative-decrease (AIMD) algorithm to maximize throughput while respecting rate limits.

Understanding the Dual-Layer Retry Architecture

The SDK implements retries at two levels: individual API calls use a decorator-based approach, while the batch executor orchestrates retries across concurrent workers. Understanding how these layers interact is crucial for configuring reliable inference pipelines.

Per-Call Retry Decorator

Individual Forge clients (ESM3ForgeInferenceClient, SequenceStructureForgeInferenceClient, etc.) automatically retry transient HTTP errors through the @retry_decorator defined in esm/sdk/retry.py. This decorator uses tenacity to catch specific error codes—429, 500, 502, and 504—and applies exponential backoff between attempts.

However, when processing batches, you want to disable this per-call retry logic. The SDK exposes skip_retries_var, a contextvars.ContextVar defined in esm/sdk/retry.py, that signals the retry decorator to bypass its logic when set to True.

Batch-Level Orchestration with AIMD

The ForgeBatchExecutor class in esm/utils/forge_context_manager.py takes over retry management when skip_retries_var is enabled. It spawns a ThreadPoolExecutor and implements an AIMD rate limiter that:

  • Halves concurrency when rate-limit errors (429) are observed
  • Increments concurrency by step_up (default 1) when batches complete successfully
  • Tracks per-task attempt counts up to max_attempts (default 10)

This approach prevents thundering-herd problems while maintaining optimal throughput as API conditions change.

Implementing Reliable Batch Inference

The batch_executor factory function provides the primary interface for production batch workloads. It returns a configured ForgeBatchExecutor that handles queuing, thread management, and error classification automatically.

Basic Batch Execution with Automatic Retries

The standard pattern involves creating a Forge client, preparing your input lists, and executing within the batch executor context:

from esm.sdk import esmc_client, batch_executor
from esm.sdk.api import GenerationConfig

# Initialize client (reads ESM_API_KEY from environment)

client = esmc_client(model="esmc-600m-2024-12")

# Prepare equal-length lists of inputs and configurations

prompts = ["MKTAYIAKQRQISFVKSHFSRQDILDLW...", "MELKS..."]
configs = [GenerationConfig(num_steps=128), GenerationConfig(num_steps=128)]

# Execute with retry mechanism enabled

with batch_executor(max_attempts=12, show_progress=True) as executor:
    results = executor.execute_batch(client.batch_generate, prompts, configs)

# Handle results: either ESMProtein objects or ESMProteinError exceptions

for i, result in enumerate(results):
    if isinstance(result, Exception):
        print(f"Prompt {i} failed: {result}")
    else:
        print(f"Prompt {i} succeeded: {len(result.sequence)} residues")

When batch_executor.__enter__ executes, it sets skip_retries_var to True, disabling the per-call retry logic in esm/sdk/retry.py. The executor then validates input lengths, initializes the AIMDRateLimiter with default concurrency of 32, and begins processing.

Async Client Integration

If you prefer async APIs (async_generate, async_fold), wrap the coroutine with asyncio.run inside a blocking wrapper function:

import asyncio
from esm.sdk import client as esm3_client, batch_executor
from esm.sdk.api import GenerationConfig

client = esm3_client(model="esm3-sm-open-v1")
prompts = ["MKT...", "MEK..."]
configs = [GenerationConfig(num_steps=64), GenerationConfig(num_steps=64)]

async def async_generate_one(prompt, config):
    return await client.async_generate(prompt, config)

with batch_executor(max_attempts=8) as executor:
    def blocking_wrapper(prompt, config):
        return asyncio.run(async_generate_one(prompt, config))
    
    results = executor.execute_batch(blocking_wrapper, prompts, configs)

The executor treats the wrapper as a standard blocking function while the underlying async client maintains connection pooling benefits.

Configuration and Tuning

You can customize concurrency limits and retry behavior by modifying the executor's rate limiter directly:

with batch_executor(max_attempts=5, show_progress=False) as executor:
    # Reduce maximum concurrent requests to 16

    executor.rate_limiter.max_concurrency = 16
    
    results = executor.execute_batch(
        client.fold, 
        sequences, 
        msa=None, 
        config=FoldingConfig()
    )

The max_attempts parameter controls per-item retry limits, while show_progress toggles the tqdm progress bar displaying success, failure, and retry counts.

Key Implementation Details

Several components work together to provide the resilient batch processing behavior.

The skip_retries_var Context Variable

Located in esm/sdk/retry.py, this ContextVar acts as a circuit breaker. When the batch executor sets it to True, the retry_decorator checks this value and immediately yields execution to the wrapped function, delegating all retry responsibility to the batch executor's AIMD-controlled loop.

AIMDRateLimiter Mechanics

The AIMDRateLimiter class in esm/utils/forge_context_manager.py maintains a current_concurrency value that dynamically adjusts based on observed error rates. When execute_batch detects an ESMProteinError with code 429, it calls adjust_concurrency(error_seen=True), which multiplies the current limit by 0.5. Successful batches trigger adjust_concurrency(error_seen=False), incrementing the limit by 1 until reaching the maximum.

Error Classification

The retry_if_specific_error predicate in esm/sdk/retry.py identifies retriable failures by checking the error_code attribute of ESMProteinError exceptions against the set {429, 500, 502, 504}. Only errors matching these codes trigger re-queueing; permanent errors (400, 401, 403) fail immediately after the first attempt.

Summary

  • The Forge API retry mechanism uses a dual-layer approach: automatic per-call retries for simple operations and batch_executor for coordinated batch processing.
  • The batch_executor context manager disables per-call retries via skip_retries_var and manages a thread pool with AIMD rate limiting.
  • Transient error codes (429, 500, 502, 504) trigger automatic retry with exponential backoff and concurrency adjustment.
  • Core files include esm/utils/forge_context_manager.py (executor implementation), esm/sdk/retry.py (retry logic and context variables), and esm/sdk/__init__.py (public API).
  • Configure reliability using max_attempts (default 10) and dynamic max_concurrency adjustments based on API response patterns.

Frequently Asked Questions

What is the difference between per-call retries and the batch executor?

Individual API calls use the @retry_decorator in esm/sdk/retry.py with exponential backoff for transient errors. The batch executor disables this per-call logic via skip_retries_var and implements its own retry loop with AIMD rate limiting, which is more efficient for concurrent batch processing because it adjusts global concurrency based on observed error rates rather than retrying independently.

How does the AIMD rate limiter prevent API throttling?

The AIMDRateLimiter in esm/utils/forge_context_manager.py starts with a default concurrency of 32 and applies multiplicative decrease (halving concurrency) when it encounters 429 errors, then additive increase (stepping up by 1) during successful batches. This quickly backs off under load and gradually probes for increased capacity when the API is healthy, preventing sustained rate limit violations.

Which HTTP error codes trigger the retry mechanism?

According to the retry_if_specific_error predicate in esm/sdk/retry.py, the SDK retries requests that fail with HTTP status codes 429 (Too Many Requests), 500 (Internal Server Error), 502 (Bad Gateway), and 504 (Gateway Timeout). Client errors like 400 (Bad Request) or 401 (Unauthorized) fail immediately without retry.

How do I adjust the maximum number of retry attempts for batch jobs?

Pass the max_attempts parameter (default 10) to the batch_executor factory: with batch_executor(max_attempts=12) as executor:. The execute_batch method tracks attempts per task and surfaces an ESMProteinError for items that exceed this limit, allowing your application to handle permanent failures gracefully while successfully processing retriable ones.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →