How to Implement Rate Limiting and Cost Controls for LLM API Calls in Cua
Cua provides built-in rate limiting through exponential backoff retries and cost controls via the BudgetManagerCallback that automatically stops agent execution when a monetary budget is exceeded.
Cua (Computer Use Agent) is an open-source framework for building AI agents that interact with computer interfaces. When deploying LLM-powered agents in production, controlling costs and handling API rate limits are critical concerns. The framework includes native mechanisms for both exponential backoff on transient errors and hard budget caps on LLM spending.
Architecture Overview
The implementation spans the agent core and callback systems. ComputerAgent wraps every LLM call in retry logic while BudgetManagerCallback monitors cumulative spend.
Transient Error Handling
In libs/python/agent/cua_agent/agent.py, the _predict_step_with_retry method wraps each prediction step in an exponential backoff loop. The helper _is_retryable_error detects RateLimitError, ServiceUnavailableError, APIConnectionError, timeouts, and generic 429/5xx responses. By default, the system retries up to three times with increasing delays.
Cost Tracking and Enforcement
The BudgetManagerCallback class in libs/python/agent/cua_agent/callbacks/budget_manager.py intercepts every LLM response. It extracts _hidden_params.response_cost (provided by litellm-compatible providers) and accumulates the total. When spending reaches the configured maximum, the callback raises BudgetExceededError or silently aborts the run.
Configuring Rate Limits with Exponential Backoff
Cua handles rate limits automatically, but you can customize the retry behavior through constructor parameters.
Setting Maximum Retry Attempts
Pass max_retries to ComputerAgent to control resilience against provider throttling:
from cua.agent.cua_agent import ComputerAgent
agent = ComputerAgent(
model="anthropic/claude-3-sonnet-20240229",
max_retries=8, # Total of 9 attempts (initial + 8 retries)
)
Under the hood, _predict_step_with_retry implements exponential backoff with jitter. The delay follows the pattern base_delay * 2**attempt + random_jitter, starting at 2 seconds. When a RateLimitError occurs, the agent logs the attempt and sleeps before retrying.
Retryable Error Detection
The _is_retryable_error function in agent.py identifies the following as transient:
RateLimitError(HTTP 429)ServiceUnavailableError(HTTP 503)APIConnectionErrorand timeouts- Generic 5xx server errors
If all retries fail, the original exception propagates to your calling code.
Implementing Cost Controls
Cua provides granular budget management through the max_trajectory_budget parameter, which automatically registers a BudgetManagerCallback.
Setting a Global Monetary Budget
Pass a float to set a USD limit for the entire agent trajectory:
agent = ComputerAgent(
model="openai/computer-use-preview",
max_trajectory_budget=2.0, # Stop after $2.00 USD
max_retries=5,
verbosity=10, # Log cost updates to stdout
)
The callback initialization occurs in ComputerAgent.__init__ (lines 290-304 in agent.py). After each LLM request, the OpenAI loop extracts response_cost from response._hidden_params (lines 267-272 in libs/python/agent/cua_agent/loops/openai.py) and passes it to the callback.
Per-Model Budget Configuration
For multi-model trajectories, pass a dictionary mapping model names to budget limits:
agent = ComputerAgent(
model="omni+vertex_ai/gemini-pro",
max_trajectory_budget={
"openai/computer-use-preview": 1.5,
"omni+vertex_ai/gemini-pro": 2.5,
},
)
BudgetManagerCallback maintains separate cost trackers for each model key and enforces the appropriate ceiling when each respective model is invoked.
Extending Rate Limiting to Custom Endpoints
For non-litellm API calls (direct HTTP requests to custom LLM endpoints), reuse the async token bucket implementation from the sandbox apps.
Using the Token Bucket Pattern
The _TokenBucket class in libs/python/cua-sandbox-apps/cua_sandbox_apps/discovery/batch_enrich.py provides async rate limiting:
from cua_sandbox_apps.discovery.batch_enrich import _TokenBucket
import asyncio
# Limit to 10 requests per second
bucket = _TokenBucket(rate=10)
async def call_custom_llm(session, payload):
await bucket.acquire() # Blocks until token available
async with session.post("https://api.custom-llm.com/v1/chat", json=payload) as resp:
data = await resp.json()
return data
Integrating Custom Costs with BudgetManager
Feed custom endpoint costs back into the budget system by calling on_usage directly:
# Assuming you have access to the callback instance
cost = data.get("cost", 0.0)
await budget_callback.on_usage({"response_cost": cost})
This allows the BudgetManagerCallback to track spending across both standard LLM calls and custom API endpoints.
Complete Working Example
The following example demonstrates combined rate limiting and budget enforcement:
import asyncio
from cua.agent.cua_agent import ComputerAgent
async def main():
# Configure $1.00 budget with aggressive retry policy
agent = ComputerAgent(
model="openai/computer-use-preview",
max_trajectory_budget=1.0,
max_retries=5,
verbosity=10,
)
# Execute computer use task
result = await agent.run([
{
"type": "click",
"instruction": "Click the 'Sign up' button on the page",
}
])
print("Task completed:", result)
if __name__ == "__main__":
asyncio.run(main())
Running this script produces verbose logging showing each step's cost (e.g., 🛠️ step ... 💸 $0.03). If OpenAI returns a 429 status, the agent automatically backs off and retries. Once cumulative spend reaches $1.00, the agent prints "Budget exceeded" and terminates execution.
Summary
- Rate limiting is handled automatically by
_predict_step_with_retryinagent.py, which implements exponential backoff with jitter forRateLimitErrorand other transient failures. Configure viamax_retries. - Cost controls use
BudgetManagerCallbackregistered throughmax_trajectory_budget(float or dict) to enforce spending limits across single or multiple models. - Custom implementations can leverage
_TokenBucketfor async rate limiting and feed costs back to the budget manager viaon_usage.
Frequently Asked Questions
How does Cua detect rate limit errors from different providers?
Cua uses the _is_retryable_error helper in libs/python/agent/cua_agent/agent.py to catch standard litellm exceptions including RateLimitError, ServiceUnavailableError, and APIConnectionError, plus generic HTTP 429 and 5xx status codes. This covers OpenAI, Anthropic, and other litellm-compatible providers uniformly.
Can I set different budgets for different models in the same agent run?
Yes. Pass a dictionary to max_trajectory_budget mapping model identifiers (e.g., "openai/computer-use-preview") to specific USD limits. The BudgetManagerCallback tracks costs separately for each model key and enforces the respective limit when that model is called.
What happens when the budget is exceeded during execution?
By default, the BudgetManagerCallback prints a "Budget exceeded" message and aborts the agent run. If you initialize the callback with raise_error=True, it raises a BudgetExceededError exception instead. In both cases, no further LLM calls are made once the limit is reached.
Can I use Cua's rate limiting for non-LLM API calls like web scraping?
Yes. The _TokenBucket class in libs/python/cua-sandbox-apps/cua_sandbox_apps/discovery/batch_enrich.py implements an async token bucket algorithm suitable for any API rate limiting. Instantiate it with your desired rate and call await bucket.acquire() before making HTTP requests to enforce throughput limits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →