# How Pathway Handles Rate Limiting and Retry Logic for LLM API Calls

> Learn how Pathway's LLM wrappers manage rate limiting and retry logic with configurable exponential backoff and fixed delay policies. Handle API errors and network failures seamlessly.

- Repository: [Pathway/llm-app](https://github.com/pathwaycom/llm-app)
- Tags: how-to-guide
- Published: 2026-03-07

---

**Pathway's LLM wrappers use a configurable `retry_strategy` parameter that implements exponential backoff or fixed delay policies to automatically handle HTTP 429 errors and transient network failures.**

The `pathwaycom/llm-app` repository provides production-ready templates for building LLM-powered data pipelines, with sophisticated **rate limiting and retry logic for LLM API calls** built directly into the core wrappers. Pathway abstracts resilience patterns into pluggable strategies that prevent application failures when encountering provider quotas or network instability, allowing developers to configure retry behavior through either Python code or declarative YAML files.

## Understanding Pathway's Retry Strategy Architecture

### The `retry_strategy` Parameter

Pathway's LLM wrappers—including `OpenAIChat` and `OpenAIEmbedder`—expose a unified `retry_strategy` parameter that accepts strategy objects from the `pw.udfs` or `pw.asynchronous` modules. This parameter dictates how the system classifies and responds to transient failures, specifically treating HTTP 429 "Too Many Requests" responses as recoverable errors that trigger the configured back-off sequence.

### Available Retry Strategies

Pathway implements two distinct retry strategies optimized for different rate-limiting scenarios:

- **ExponentialBackoffRetryStrategy**: Increases wait times exponentially between attempts (e.g., 1s, 2s, 4s) with random jitter to prevent thundering-herd effects when multiple instances simultaneously hit limits
- **FixedDelayRetryStrategy**: Maintains a constant pause duration between attempts, ideal for providers with documented, consistent rate-limit reset windows

## Implementing Exponential Backoff for Rate Limiting

The **exponential backoff** strategy serves as the default recommendation for handling aggressive rate limits from providers like OpenAI. In the Unstructured-to-SQL template, the implementation in [`templates/unstructured_to_sql_on_the_fly/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/unstructured_to_sql_on_the_fly/app.py) demonstrates proper configuration:

```python
model = OpenAIChat(
    api_key=api_key,
    model="gpt-4",
    temperature=0,
    max_tokens=256,
    retry_strategy=pw.udfs.ExponentialBackoffRetryStrategy(),
    cache_strategy=pw.udfs.DefaultCache(),
)

```

This configuration appears at lines 188-194 for the schema extraction model and lines 231-237 for the query generation model. The strategy defaults to **5 maximum retries** with exponentially increasing delays and random jitter to avoid synchronized retry storms when multiple pipeline components encounter rate limits simultaneously.

## Using Fixed Delay Retry Strategies

For scenarios requiring predictable retry intervals, Pathway offers the **FixedDelayRetryStrategy**. The Drive-Alert template in [`templates/drive_alert/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/drive_alert/app.py) demonstrates this approach for both embedding and chat operations:

```python
embedder = OpenAIEmbedder(
    api_key=os.environ["OPENAI_API_KEY"],
    model="text-embedding-3-small",
    retry_strategy=pw.asynchronous.FixedDelayRetryStrategy(),
)

chat = OpenAIChat(
    api_key=os.environ["OPENAI_API_KEY"],
    model="gpt-4o",
    retry_strategy=pw.asynchronous.FixedDelayRetryStrategy(),
)

```

This configuration appears at lines 150-155 for the embedder and lines 198-204 for the chat model. Use **fixed delay** when your LLM provider guarantees consistent rate-limit windows, as it maintains steady, predictable request spacing without the variable delays of exponential backoff.

## Configuring Retry Logic via YAML

Pathway supports **declarative configuration** of retry strategies through YAML files, enabling operations teams to adjust behavior without modifying application code. The Question-Answering RAG template in [`templates/question_answering_rag/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/question_answering_rag/app.yaml) demonstrates this pattern:

```yaml
$llm: !pw.xpacks.llm.llms.OpenAIChat
  model: "gpt-4.1-mini"
  retry_strategy: !pw.udfs.ExponentialBackoffRetryStrategy
    max_retries: 6
  cache_strategy: !pw.udfs.DefaultCache {}
  temperature: 0
  async_mode: "fully_async"

```

This configuration at lines 48-52 sets the maximum retries to **6** and enables fully asynchronous mode. YAML-based configuration allows teams to tune **rate limiting and retry logic for LLM API calls** independently of development cycles, facilitating rapid adjustments when provider quotas or latency requirements change.

## How the Retry Mechanism Works Under the Hood

Pathway's retry system operates through a four-stage pipeline when processing LLM API requests:

1. **Error Detection** – The wrapper captures HTTP errors from the provider, specifically classifying HTTP 429 "Too Many Requests", 5xx server errors, and connection timeouts as transient failures eligible for retry.

2. **Decision Logic** – If the error qualifies as transient, the attached `retry_strategy` calculates the next wait time based on the configured policy.

3. **Back-off Policy** – 
   - **Exponential back-off** multiplies the previous delay by a factor (typically 2) and adds random jitter to prevent thundering-herd effects when multiple instances retry simultaneously
   - **Fixed delay** maintains a constant pause duration between attempts

4. **Retry Limit** – Once `max_retries` is exhausted (default 5), the original exception propagates to the user-defined pipeline, enabling custom failure handling such as fallback responses or alerting mechanisms.

This architecture ensures **rate-limit compliance** by automatically pausing requests when encountering quota restrictions, while maintaining pipeline resilience against transient infrastructure issues.

## Summary

- Pathway's LLM wrappers accept a configurable `retry_strategy` parameter that automates handling of HTTP 429 errors and transient failures
- **ExponentialBackoffRetryStrategy** implements increasing delays with jitter (default 5 retries) to handle aggressive rate limits from providers like OpenAI
- **FixedDelayRetryStrategy** provides predictable retry intervals suitable for providers with consistent quota reset windows
- Both strategies can be configured via Python code or declarative YAML files, allowing operations teams to tune behavior without code changes
- The retry mechanism detects transient errors, applies the configured back-off policy, and propagates permanent failures after exhausting retry limits

## Frequently Asked Questions

### How does Pathway prevent thundering-herd effects during retries?

Pathway's `ExponentialBackoffRetryStrategy` includes **jitter**—random variations in the calculated delay times. This prevents synchronized retry storms when multiple pipeline instances simultaneously hit rate limits, distributing retry attempts across time rather than clustering them at fixed intervals that could overwhelm the provider.

### What is the default number of retries in Pathway's exponential backoff strategy?

The default `max_retries` value for `pw.udfs.ExponentialBackoffRetryStrategy` is **5**. You can override this by passing the `max_retries` parameter when instantiating the strategy in Python or setting it in your YAML configuration file, as demonstrated in the Question-Answering RAG template with `max_retries: 6`.

### How does Pathway handle HTTP 429 errors specifically?

Pathway treats HTTP 429 "Too Many Requests" as a transient error that triggers the configured retry strategy. When using `ExponentialBackoffRetryStrategy`, the system waits with exponentially increasing delays (plus jitter) before retrying, allowing the provider's rate limit window to reset without manual intervention or application failure.

### Can I use different retry strategies for different LLM models in the same application?

Yes. Each LLM wrapper instance accepts its own `retry_strategy` parameter independently. For example, you can configure `ExponentialBackoffRetryStrategy` for a high-priority GPT-4 model while using `FixedDelayRetryStrategy` for a background embedding model, optimizing retry behavior based on each model's specific rate limits and criticality to your pipeline.