# How to Scale Cua Agents Across Multiple Cloud VMs: Async Fleet Management Guide

> Scale Cua agents across multiple cloud VMs efficiently. Learn how to use VMProviderType.CLOUD and asyncio.Semaphore for concurrent agent execution while managing API limits and resources.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: how-to-guide
- Published: 2026-04-27

---

**Use `VMProviderType.CLOUD` with `asyncio.Semaphore` to run hundreds of Cua agents concurrently across distributed cloud sandboxes while respecting API rate limits and VM resource caps.**

Scaling Cua agents beyond a single machine requires leveraging the SDK's provider-agnostic architecture. The Cua framework abstracts local and remote execution through a unified `Computer` interface, allowing you to deploy agent fleets across cloud VMs using the same API calls you use for local development. This guide demonstrates the production-ready pattern for provisioning cloud sandboxes and managing concurrent agent execution using the `CloudProvider` implementation and Python's `asyncio` concurrency primitives.

## Cloud Provider Architecture

The Cua SDK decouples agent logic from infrastructure through the `BaseVMProvider` abstraction defined in [`libs/python/computer/computer/base.py`](https://github.com/trycua/cua/blob/main/libs/python/computer/computer/base.py). When scaling across cloud VMs, you instantiate `Computer` objects with `provider_type=VMProviderType.CLOUD`, which routes all sandbox operations through the `CloudProvider` class.

### Host Resolution and Caching

In [`libs/python/computer/computer/providers/cloud/provider.py`](https://github.com/trycua/cua/blob/main/libs/python/computer/computer/providers/cloud/provider.py), the `CloudProvider` maintains an internal `_host_cache` dictionary that stores resolved VM hostnames. This implementation minimizes DNS lookups and connection latency by caching the resolved host for each sandbox ID. If a cached host becomes unreachable, the provider falls back to a legacy host format automatically, ensuring your fleet remains resilient without manual intervention.

### Async HTTP API Integration

The `CloudProvider` communicates with the Cua public API at `/v1/vms` endpoints. All VM lifecycle methods—including `run_vm()`, `stop_vm()`, and `list_vms()`—are implemented as async coroutines. This design eliminates blocking I/O during network operations, allowing your orchestration logic to manage thousands of pending VM operations without consuming thread resources.

## Implementing Fleet Scaling

The recommended scaling pattern involves three distinct phases: provisioning independent sandboxes, binding agents to those sandboxes with concurrency limits, and gracefully handling lifecycle cleanup.

### Step 1: Initialize Cloud Provider Instances

Each agent requires a dedicated `Computer` instance configured for cloud execution. Set the `provider_type` parameter to `VMProviderType.CLOUD` and provide a unique sandbox name. The SDK automatically resolves the connection string through the `CloudProvider` implementation.

```python
from computer import Computer, VMProviderType
import os

def make_computer(name: str) -> Computer:
    return Computer(
        os_type="linux",
        api_key=os.getenv("CUA_API_KEY"),
        name=name,
        provider_type=VMProviderType.CLOUD,
    )

```

### Step 2: Provision Sandbox Pools Programmatically

While you can manually provision VMs via the Cua CLI, fleet scaling requires programmatic control. Use `CloudProvider.run_vm()` to spawn sandboxes dynamically:

```python
import asyncio
from computer.computer.providers.cloud.provider import CloudProvider

async def provision_fleet(names: list[str], image: str = "cua/ubuntu:latest"):
    async with CloudProvider() as provider:
        tasks = [
            provider.run_vm(name, image=image) for name in names
        ]
        results = await asyncio.gather(*tasks)
        return results

# Provision 50 agents

sandbox_names = [f"agent-{i:02d}" for i in range(1, 51)]
asyncio.run(provision_fleet(sandbox_names))

```

This pattern, found in the core `CloudProvider` source code, executes VM creation calls concurrently while maintaining connection pooling for the underlying HTTP client.

### Step 3: Execute Agents with Bounded Concurrency

Raw parallelism will overwhelm API rate limits. The reference implementation in [`demo/1_fleet_throughput.py`](https://github.com/trycua/cua/blob/main/demo/1_fleet_throughput.py) demonstrates using `asyncio.Semaphore` to cap simultaneous operations. This prevents throttling while maximizing throughput.

```python
import asyncio
import logging
from cua_agent import ComputerAgent

async def run_fleet(sandbox_names: list[str], concurrency: int = 20):
    sem = asyncio.Semaphore(concurrency)
    
    async def run_one(name: str):
        async with sem:  # Limit active connections

            comp = make_computer(name)
            agent = ComputerAgent(
                model="openai/gpt-4o",
                tools=[comp],
                verbosity=logging.INFO,
            )
            
            async for result in agent.run("Open Chrome and navigate to https://cua.ai"):
                print(f"[{name}] {result.get('text')}")
            
            await comp.disconnect()  # Clean resource release

            return name
    
    return await asyncio.gather(*[run_one(n) for n in sandbox_names])

```

The semaphore value should match your API quota or VM resource constraints. For CPU-intensive agents, set concurrency equal to your cloud account's vCPU quota divided by the cores per sandbox.

## Error Handling and Resilience

The `CloudProvider` implementation catches HTTP errors (401 unauthorized, 404 not found, network timeouts) and converts them into structured status dictionaries with `{"status": "error", ...}` formatting. This allows your fleet orchestrator to implement circuit breakers or retry logic without crashing the entire agent pool.

**Retry Strategy**: Wrap `provider.run_vm()` calls in `tenacity` or similar retry libraries with exponential backoff for transient 5xx errors. For 401 errors, refresh the `CUA_API_KEY` immediately.

**Health Monitoring**: Collect per-sandbox metrics (throughput, latency, error rates) as demonstrated in the fleet demo's step latency aggregation. Track steps-per-hour to detect out-of-quota or flaky VMs before they degrade fleet performance.

## Elastic Autoscaled Sandbox Pools (Roadmap)

The Cua Cloud roadmap includes Elastic Autoscaled Sandbox Pools, which will allow you to define minimum and maximum pool sizes and let the service automatically scale container counts based on demand. Until this feature ships (as documented in the project blog), emulate elasticity by wrapping the provisioning and execution logic in a control loop that monitors queue depth and spawns new sandboxes when `concurrency < pending_tasks`.

## Summary

- **Use `VMProviderType.CLOUD`** to route `Computer` operations through the `CloudProvider` implementation rather than local QEMU or Docker backends.
- **Cache connections** via the `_host_cache` mechanism in [`libs/python/computer/computer/providers/cloud/provider.py`](https://github.com/trycua/cua/blob/main/libs/python/computer/computer/providers/cloud/provider.py) to minimize DNS overhead.
- **Bound concurrency** with `asyncio.Semaphore` (lines 90-102 in [`demo/1_fleet_throughput.py`](https://github.com/trycua/cua/blob/main/demo/1_fleet_throughput.py)) to prevent API throttling while maintaining high throughput.
- **Provision dynamically** using `CloudProvider.run_vm()` to create VM fleets programmatically without manual CLI steps.
- **Handle errors gracefully** by checking status dictionaries for error states and implementing retry logic for transient failures.

## Frequently Asked Questions

### How many concurrent Cua agents can I run on cloud VMs?

You can theoretically run hundreds of concurrent agents, but you must respect the `asyncio.Semaphore` limit to avoid overwhelming the Cua API. According to the fleet demo implementation, start with 20 concurrent connections and scale up based on your API quota and cloud account limits. The bottleneck is typically API rate limits or your cloud provider's vCPU quota, not the Cua SDK itself.

### What is the difference between local and cloud providers in Cua?

The `Computer` class uses the same API regardless of provider type. When you set `provider_type=VMProviderType.LOCAL`, the SDK communicates with local QEMU, Lume, or Docker instances. When you set `VMProviderType.CLOUD`, it routes through `CloudProvider` to remote VMs via the public API at `/v1/vms`. This abstraction allows you to develop locally and deploy to production cloud fleets without changing agent code—only the initialization parameter changes.

### How do I handle VM failures when scaling Cua agents across multiple cloud VMs?

The `CloudProvider` catches HTTP errors and network timeouts, returning status dictionaries with error states instead of raising unhandled exceptions. Check for `{"status": "error"}` in your orchestration loop and implement a circuit breaker: if a VM returns consistent 404 errors, mark it as unhealthy, call `await comp.disconnect()` to release resources, and spawn a replacement via `CloudProvider.run_vm()` with a new unique name.

### Is autoscaling supported for Cua cloud sandboxes?

Autoscaling is on the Cua Cloud roadmap under "Elastic Autoscaled Sandbox Pools," which will automatically adjust the number of containers based on demand. Until this feature is released, you must implement autoscaling manually by monitoring your task queue depth and programmatically invoking `CloudProvider.run_vm()` to scale up or `stop_vm()` to scale down based on your application's current load.