How to Scale Cua Agents Across Multiple Cloud VMs: Async Fleet Management Guide

Use VMProviderType.CLOUD with asyncio.Semaphore to run hundreds of Cua agents concurrently across distributed cloud sandboxes while respecting API rate limits and VM resource caps.

Scaling Cua agents beyond a single machine requires leveraging the SDK's provider-agnostic architecture. The Cua framework abstracts local and remote execution through a unified Computer interface, allowing you to deploy agent fleets across cloud VMs using the same API calls you use for local development. This guide demonstrates the production-ready pattern for provisioning cloud sandboxes and managing concurrent agent execution using the CloudProvider implementation and Python's asyncio concurrency primitives.

Cloud Provider Architecture

The Cua SDK decouples agent logic from infrastructure through the BaseVMProvider abstraction defined in libs/python/computer/computer/base.py. When scaling across cloud VMs, you instantiate Computer objects with provider_type=VMProviderType.CLOUD, which routes all sandbox operations through the CloudProvider class.

Host Resolution and Caching

In libs/python/computer/computer/providers/cloud/provider.py, the CloudProvider maintains an internal _host_cache dictionary that stores resolved VM hostnames. This implementation minimizes DNS lookups and connection latency by caching the resolved host for each sandbox ID. If a cached host becomes unreachable, the provider falls back to a legacy host format automatically, ensuring your fleet remains resilient without manual intervention.

Async HTTP API Integration

The CloudProvider communicates with the Cua public API at /v1/vms endpoints. All VM lifecycle methods—including run_vm(), stop_vm(), and list_vms()—are implemented as async coroutines. This design eliminates blocking I/O during network operations, allowing your orchestration logic to manage thousands of pending VM operations without consuming thread resources.

Implementing Fleet Scaling

The recommended scaling pattern involves three distinct phases: provisioning independent sandboxes, binding agents to those sandboxes with concurrency limits, and gracefully handling lifecycle cleanup.

Step 1: Initialize Cloud Provider Instances

Each agent requires a dedicated Computer instance configured for cloud execution. Set the provider_type parameter to VMProviderType.CLOUD and provide a unique sandbox name. The SDK automatically resolves the connection string through the CloudProvider implementation.

from computer import Computer, VMProviderType
import os

def make_computer(name: str) -> Computer:
    return Computer(
        os_type="linux",
        api_key=os.getenv("CUA_API_KEY"),
        name=name,
        provider_type=VMProviderType.CLOUD,
    )

Step 2: Provision Sandbox Pools Programmatically

While you can manually provision VMs via the Cua CLI, fleet scaling requires programmatic control. Use CloudProvider.run_vm() to spawn sandboxes dynamically:

import asyncio
from computer.computer.providers.cloud.provider import CloudProvider

async def provision_fleet(names: list[str], image: str = "cua/ubuntu:latest"):
    async with CloudProvider() as provider:
        tasks = [
            provider.run_vm(name, image=image) for name in names
        ]
        results = await asyncio.gather(*tasks)
        return results

# Provision 50 agents

sandbox_names = [f"agent-{i:02d}" for i in range(1, 51)]
asyncio.run(provision_fleet(sandbox_names))

This pattern, found in the core CloudProvider source code, executes VM creation calls concurrently while maintaining connection pooling for the underlying HTTP client.

Step 3: Execute Agents with Bounded Concurrency

Raw parallelism will overwhelm API rate limits. The reference implementation in demo/1_fleet_throughput.py demonstrates using asyncio.Semaphore to cap simultaneous operations. This prevents throttling while maximizing throughput.

import asyncio
import logging
from cua_agent import ComputerAgent

async def run_fleet(sandbox_names: list[str], concurrency: int = 20):
    sem = asyncio.Semaphore(concurrency)
    
    async def run_one(name: str):
        async with sem:  # Limit active connections

            comp = make_computer(name)
            agent = ComputerAgent(
                model="openai/gpt-4o",
                tools=[comp],
                verbosity=logging.INFO,
            )
            
            async for result in agent.run("Open Chrome and navigate to https://cua.ai"):
                print(f"[{name}] {result.get('text')}")
            
            await comp.disconnect()  # Clean resource release

            return name
    
    return await asyncio.gather(*[run_one(n) for n in sandbox_names])

The semaphore value should match your API quota or VM resource constraints. For CPU-intensive agents, set concurrency equal to your cloud account's vCPU quota divided by the cores per sandbox.

Error Handling and Resilience

The CloudProvider implementation catches HTTP errors (401 unauthorized, 404 not found, network timeouts) and converts them into structured status dictionaries with {"status": "error", ...} formatting. This allows your fleet orchestrator to implement circuit breakers or retry logic without crashing the entire agent pool.

Retry Strategy: Wrap provider.run_vm() calls in tenacity or similar retry libraries with exponential backoff for transient 5xx errors. For 401 errors, refresh the CUA_API_KEY immediately.

Health Monitoring: Collect per-sandbox metrics (throughput, latency, error rates) as demonstrated in the fleet demo's step latency aggregation. Track steps-per-hour to detect out-of-quota or flaky VMs before they degrade fleet performance.

Elastic Autoscaled Sandbox Pools (Roadmap)

The Cua Cloud roadmap includes Elastic Autoscaled Sandbox Pools, which will allow you to define minimum and maximum pool sizes and let the service automatically scale container counts based on demand. Until this feature ships (as documented in the project blog), emulate elasticity by wrapping the provisioning and execution logic in a control loop that monitors queue depth and spawns new sandboxes when concurrency < pending_tasks.

Summary

  • Use VMProviderType.CLOUD to route Computer operations through the CloudProvider implementation rather than local QEMU or Docker backends.
  • Cache connections via the _host_cache mechanism in libs/python/computer/computer/providers/cloud/provider.py to minimize DNS overhead.
  • Bound concurrency with asyncio.Semaphore (lines 90-102 in demo/1_fleet_throughput.py) to prevent API throttling while maintaining high throughput.
  • Provision dynamically using CloudProvider.run_vm() to create VM fleets programmatically without manual CLI steps.
  • Handle errors gracefully by checking status dictionaries for error states and implementing retry logic for transient failures.

Frequently Asked Questions

How many concurrent Cua agents can I run on cloud VMs?

You can theoretically run hundreds of concurrent agents, but you must respect the asyncio.Semaphore limit to avoid overwhelming the Cua API. According to the fleet demo implementation, start with 20 concurrent connections and scale up based on your API quota and cloud account limits. The bottleneck is typically API rate limits or your cloud provider's vCPU quota, not the Cua SDK itself.

What is the difference between local and cloud providers in Cua?

The Computer class uses the same API regardless of provider type. When you set provider_type=VMProviderType.LOCAL, the SDK communicates with local QEMU, Lume, or Docker instances. When you set VMProviderType.CLOUD, it routes through CloudProvider to remote VMs via the public API at /v1/vms. This abstraction allows you to develop locally and deploy to production cloud fleets without changing agent code—only the initialization parameter changes.

How do I handle VM failures when scaling Cua agents across multiple cloud VMs?

The CloudProvider catches HTTP errors and network timeouts, returning status dictionaries with error states instead of raising unhandled exceptions. Check for {"status": "error"} in your orchestration loop and implement a circuit breaker: if a VM returns consistent 404 errors, mark it as unhealthy, call await comp.disconnect() to release resources, and spawn a replacement via CloudProvider.run_vm() with a new unique name.

Is autoscaling supported for Cua cloud sandboxes?

Autoscaling is on the Cua Cloud roadmap under "Elastic Autoscaled Sandbox Pools," which will automatically adjust the number of containers based on demand. Until this feature is released, you must implement autoscaling manually by monitoring your task queue depth and programmatically invoking CloudProvider.run_vm() to scale up or stop_vm() to scale down based on your application's current load.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →