# How to Implement Rate Limiting and Cost Controls for LLM API Calls in Cua

> Implement LLM API rate limiting and cost controls in Cua. Learn to use exponential backoff retries and BudgetManagerCallback for budget management and agent execution stops.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: how-to-guide
- Published: 2026-04-27

---

**Cua provides built-in rate limiting through exponential backoff retries and cost controls via the `BudgetManagerCallback` that automatically stops agent execution when a monetary budget is exceeded.**

Cua (Computer Use Agent) is an open-source framework for building AI agents that interact with computer interfaces. When deploying LLM-powered agents in production, controlling costs and handling API rate limits are critical concerns. The framework includes native mechanisms for both exponential backoff on transient errors and hard budget caps on LLM spending.

## Architecture Overview

The implementation spans the agent core and callback systems. `ComputerAgent` wraps every LLM call in retry logic while `BudgetManagerCallback` monitors cumulative spend.

### Transient Error Handling

In [`libs/python/agent/cua_agent/agent.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/agent.py), the `_predict_step_with_retry` method wraps each prediction step in an exponential backoff loop. The helper `_is_retryable_error` detects `RateLimitError`, `ServiceUnavailableError`, `APIConnectionError`, timeouts, and generic 429/5xx responses. By default, the system retries up to three times with increasing delays.

### Cost Tracking and Enforcement

The `BudgetManagerCallback` class in [`libs/python/agent/cua_agent/callbacks/budget_manager.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/callbacks/budget_manager.py) intercepts every LLM response. It extracts `_hidden_params.response_cost` (provided by litellm-compatible providers) and accumulates the total. When spending reaches the configured maximum, the callback raises `BudgetExceededError` or silently aborts the run.

## Configuring Rate Limits with Exponential Backoff

Cua handles rate limits automatically, but you can customize the retry behavior through constructor parameters.

### Setting Maximum Retry Attempts

Pass `max_retries` to `ComputerAgent` to control resilience against provider throttling:

```python
from cua.agent.cua_agent import ComputerAgent

agent = ComputerAgent(
    model="anthropic/claude-3-sonnet-20240229",
    max_retries=8,  # Total of 9 attempts (initial + 8 retries)

)

```

Under the hood, `_predict_step_with_retry` implements exponential backoff with jitter. The delay follows the pattern `base_delay * 2**attempt + random_jitter`, starting at 2 seconds. When a `RateLimitError` occurs, the agent logs the attempt and sleeps before retrying.

### Retryable Error Detection

The `_is_retryable_error` function in [`agent.py`](https://github.com/trycua/cua/blob/main/agent.py) identifies the following as transient:

- `RateLimitError` (HTTP 429)
- `ServiceUnavailableError` (HTTP 503)
- `APIConnectionError` and timeouts
- Generic 5xx server errors

If all retries fail, the original exception propagates to your calling code.

## Implementing Cost Controls

Cua provides granular budget management through the `max_trajectory_budget` parameter, which automatically registers a `BudgetManagerCallback`.

### Setting a Global Monetary Budget

Pass a float to set a USD limit for the entire agent trajectory:

```python
agent = ComputerAgent(
    model="openai/computer-use-preview",
    max_trajectory_budget=2.0,  # Stop after $2.00 USD

    max_retries=5,
    verbosity=10,  # Log cost updates to stdout

)

```

The callback initialization occurs in `ComputerAgent.__init__` (lines 290-304 in [`agent.py`](https://github.com/trycua/cua/blob/main/agent.py)). After each LLM request, the OpenAI loop extracts `response_cost` from `response._hidden_params` (lines 267-272 in [`libs/python/agent/cua_agent/loops/openai.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/loops/openai.py)) and passes it to the callback.

### Per-Model Budget Configuration

For multi-model trajectories, pass a dictionary mapping model names to budget limits:

```python
agent = ComputerAgent(
    model="omni+vertex_ai/gemini-pro",
    max_trajectory_budget={
        "openai/computer-use-preview": 1.5,
        "omni+vertex_ai/gemini-pro": 2.5,
    },
)

```

`BudgetManagerCallback` maintains separate cost trackers for each model key and enforces the appropriate ceiling when each respective model is invoked.

## Extending Rate Limiting to Custom Endpoints

For non-litellm API calls (direct HTTP requests to custom LLM endpoints), reuse the async token bucket implementation from the sandbox apps.

### Using the Token Bucket Pattern

The `_TokenBucket` class in [`libs/python/cua-sandbox-apps/cua_sandbox_apps/discovery/batch_enrich.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-sandbox-apps/cua_sandbox_apps/discovery/batch_enrich.py) provides async rate limiting:

```python
from cua_sandbox_apps.discovery.batch_enrich import _TokenBucket
import asyncio

# Limit to 10 requests per second

bucket = _TokenBucket(rate=10)

async def call_custom_llm(session, payload):
    await bucket.acquire()  # Blocks until token available

    async with session.post("https://api.custom-llm.com/v1/chat", json=payload) as resp:
        data = await resp.json()
        return data

```

### Integrating Custom Costs with BudgetManager

Feed custom endpoint costs back into the budget system by calling `on_usage` directly:

```python

# Assuming you have access to the callback instance

cost = data.get("cost", 0.0)
await budget_callback.on_usage({"response_cost": cost})

```

This allows the `BudgetManagerCallback` to track spending across both standard LLM calls and custom API endpoints.

## Complete Working Example

The following example demonstrates combined rate limiting and budget enforcement:

```python
import asyncio
from cua.agent.cua_agent import ComputerAgent

async def main():
    # Configure $1.00 budget with aggressive retry policy

    agent = ComputerAgent(
        model="openai/computer-use-preview",
        max_trajectory_budget=1.0,
        max_retries=5,
        verbosity=10,
    )

    # Execute computer use task

    result = await agent.run([
        {
            "type": "click",
            "instruction": "Click the 'Sign up' button on the page",
        }
    ])
    
    print("Task completed:", result)

if __name__ == "__main__":
    asyncio.run(main())

```

Running this script produces verbose logging showing each step's cost (e.g., `🛠️ step ... 💸 $0.03`). If OpenAI returns a 429 status, the agent automatically backs off and retries. Once cumulative spend reaches $1.00, the agent prints "Budget exceeded" and terminates execution.

## Summary

- **Rate limiting** is handled automatically by `_predict_step_with_retry` in [`agent.py`](https://github.com/trycua/cua/blob/main/agent.py), which implements exponential backoff with jitter for `RateLimitError` and other transient failures. Configure via `max_retries`.
- **Cost controls** use `BudgetManagerCallback` registered through `max_trajectory_budget` (float or dict) to enforce spending limits across single or multiple models.
- **Custom implementations** can leverage `_TokenBucket` for async rate limiting and feed costs back to the budget manager via `on_usage`.

## Frequently Asked Questions

### How does Cua detect rate limit errors from different providers?

Cua uses the `_is_retryable_error` helper in [`libs/python/agent/cua_agent/agent.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/agent.py) to catch standard litellm exceptions including `RateLimitError`, `ServiceUnavailableError`, and `APIConnectionError`, plus generic HTTP 429 and 5xx status codes. This covers OpenAI, Anthropic, and other litellm-compatible providers uniformly.

### Can I set different budgets for different models in the same agent run?

Yes. Pass a dictionary to `max_trajectory_budget` mapping model identifiers (e.g., `"openai/computer-use-preview"`) to specific USD limits. The `BudgetManagerCallback` tracks costs separately for each model key and enforces the respective limit when that model is called.

### What happens when the budget is exceeded during execution?

By default, the `BudgetManagerCallback` prints a "Budget exceeded" message and aborts the agent run. If you initialize the callback with `raise_error=True`, it raises a `BudgetExceededError` exception instead. In both cases, no further LLM calls are made once the limit is reached.

### Can I use Cua's rate limiting for non-LLM API calls like web scraping?

Yes. The `_TokenBucket` class in [`libs/python/cua-sandbox-apps/cua_sandbox_apps/discovery/batch_enrich.py`](https://github.com/trycua/cua/blob/main/libs/python/cua-sandbox-apps/cua_sandbox_apps/discovery/batch_enrich.py) implements an async token bucket algorithm suitable for any API rate limiting. Instantiate it with your desired rate and call `await bucket.acquire()` before making HTTP requests to enforce throughput limits.