# Performance Overhead of Running Cua Agents in Sandboxes vs Native Execution

> Discover the performance overhead of running Cua agents in sandboxes versus native execution. Understand cold-start latency and serialization costs for optimized performance.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: performance
- Published: 2026-04-27

---

**Running Cua agents in sandboxes introduces cold‑start latency of 10–20 seconds plus per‑call serialization overhead measured in milliseconds, while native execution avoids these costs but sacrifices isolation.**

The `trycua/cua` repository provides a flexible execution model where agents can run either **natively** on the host or inside isolated sandboxes using the `@sandboxed` decorator. Understanding the performance trade‑offs between these modes is critical for optimizing agent throughput and responsiveness in production workloads.

## Understanding the Three Sources of Sandbox Overhead

Sandboxed execution in Cua introduces overhead through three distinct mechanisms. Each impacts different phases of the agent lifecycle, from initial provisioning to individual function calls.

### Cold‑Start Latency (Containers and VM Provisioning)

When invoking a sandboxed agent for the first time, the system must provision the isolation environment. According to the benchmark suite in [`tests/cold_start_benchmark.py`](https://github.com/trycua/cua/blob/main/tests/cold_start_benchmark.py), this involves creating the container or VM, loading the OS image, and performing health checks before any user code executes.

- **Linux containers**: 10–20 seconds
- **Android images**: Approximately 30 seconds

This cost is incurred once per session if you reuse the environment, or repeatedly if you spawn ephemeral sandboxes for each task.

### Serialization and Transport Costs

The `@sandboxed` decorator extracts function source code using `inspect.getsource()` located in [`libs/python/computer/helpers.py`](https://github.com/trycua/cua/blob/main/libs/python/computer/helpers.py), then marshals arguments and return values across process boundaries.

- Arguments and return values are JSON‑encoded
- Payloads travel over RPC transport (gRPC or HTTP)
- Overhead scales linearly with payload size

For small payloads, this adds only a few milliseconds per call. However, large data transfers or complex objects can significantly extend latency, as noted in the architecture documentation at [`blog/sandboxed-python-execution.md`](https://github.com/trycua/cua/blob/main/blog/sandboxed-python-execution.md).

### Container‑Level Isolation Tax

Once the sandbox is warm (already running), maintaining the isolated environment introduces a constant cost for namespace management, filesystem mounts, and networking limits. This typically adds **≤ 1 second** for a warm container, dropping to sub‑second durations after the first request completes.

## Benchmarks and Real‑World Measurements

The repository includes concrete benchmarking tools to quantify these overheads. The [`tests/cold_start_benchmark.py`](https://github.com/trycua/cua/blob/main/tests/cold_start_benchmark.py) script measures provisioning times, while micro‑benchmarks compare native versus sandboxed execution paths.

Consider this practical example from the codebase:

```python
import time
from computer.helpers import sandboxed, set_default_computer
from computer.computer import Computer

async def init():
    comp = Computer()
    await comp.run()
    set_default_computer(comp)

@sandboxed("demo_venv")
def heavy_compute(x: int) -> int:
    # All imports must be inside the function

    import math
    total = 0
    for i in range(x):
        total += math.sqrt(i)
    return total

async def benchmark():
    await init()
    
    # Native execution bypasses the decorator

    start = time.time()
    heavy_compute.__wrapped__(10_000)
    native_ms = (time.time() - start) * 1000
    
    # Sandboxed execution includes serialization + RPC

    start = time.time()
    await heavy_compute(10_000)
    sandbox_ms = (time.time() - start) * 1000
    
    print(f"Native exec: {native_ms:.2f} ms")
    print(f"Sandbox exec: {sandbox_ms:.2f} ms")
    # Typical warm container output:

    # Native exec: 12.34 ms

    # Sandbox exec: 28.71 ms

```

This demonstrates that **warm sandbox calls incur roughly 2× overhead** compared to native execution for compute‑bound workloads, exclusive of the initial cold‑start penalty.

## Architectural Details Driving Performance Costs

The performance characteristics stem from specific implementation choices in [`libs/python/computer/helpers.py`](https://github.com/trycua/cua/blob/main/libs/python/computer/helpers.py) and [`cua_sandbox/__init__.py`](https://github.com/trycua/cua/blob/main/cua_sandbox/__init__.py):

**Source Extraction**: The decorator uses `inspect.getsource()` to capture the decorated function's AST, transmitting the actual code text rather than pre‑compiled bytecode.

**Virtual Environment Isolation**: Each call executes within a named virtual environment inside the container, preventing dependency conflicts but requiring environment activation overhead.

**JSON Serialization**: All arguments and return values traverse the boundary as JSON strings. Exception handling preserves stack traces by serializing errors in the container and re‑raising them on the host.

**Transport Layer**: Communication uses either gRPC or HTTP protocols, adding network latency even for local containers due to the RPC indirection layer.

## When to Use Sandboxed vs Native Execution

Select your execution mode based on workload sensitivity and security requirements:

- **Latency‑critical loops** (high‑frequency trading, real‑time control): Use **native** execution or keep sandboxes warm with batched operations to amortize costs.
- **Security‑sensitive workloads** (untrusted user code, custom package installation): Accept the **sandboxed** overhead; the isolation benefit outweighs the 10–20 second cold‑start and per‑call latency.
- **Medium‑throughput pipelines** (data extraction, web scraping, model inference): **Sandboxed** execution is suitable; minimize payloads and reuse environments to maintain sub‑500 ms per‑call latency.

## Mitigating Overhead in Production

To minimize the performance overhead of running Cua agents in sandboxes:

1. **Maintain warm pools**: Keep containers running between requests to avoid cold‑start penalties.
2. **Batch operations**: Group multiple function calls into single `@sandboxed` invocations to amortize serialization costs.
3. **Optimize payloads**: Pass file references or identifiers rather than large data structures via JSON.
4. **Reuse environments**: Leverage named virtual environments across multiple agent tasks rather than spawning ephemeral sandboxes per the `Sandbox.ephemeral` pattern shown in [`examples/utils.py`](https://github.com/trycua/cua/blob/main/examples/utils.py).

## Summary

- **Cold‑start latency** dominates initial sandbox creation, costing 10–20 seconds for Linux and ~30 seconds for Android images according to [`tests/cold_start_benchmark.py`](https://github.com/trycua/cua/blob/main/tests/cold_start_benchmark.py).
- **Per‑call overhead** consists of source extraction, JSON serialization, and RPC transport, typically adding milliseconds to each invocation.
- **Warm containers** reduce ongoing isolation costs to ≤ 1 second after initialization.
- The `@sandboxed` decorator in [`libs/python/computer/helpers.py`](https://github.com/trycua/cua/blob/main/libs/python/computer/helpers.py) trades performance for security through container isolation and virtual environment separation.

## Frequently Asked Questions

### How long does a sandbox cold start take in Cua?

Cold‑start latency measures 10–20 seconds for Linux containers and approximately 30 seconds for Android images, as benchmarked in [`tests/cold_start_benchmark.py`](https://github.com/trycua/cua/blob/main/tests/cold_start_benchmark.py). This covers container creation, OS image loading, and health checks before code execution begins.

### What causes the per‑call latency in sandboxed Cua agents?

Per‑call overhead stems from three operations: extracting source code via `inspect.getsource()`, JSON‑encoding arguments and return values, and transmitting payloads over RPC (gRPC or HTTP). These steps add a few milliseconds for small payloads, scaling linearly with data size.

### Can I reduce sandbox overhead when running multiple Cua agent tasks?

Yes. Maintain warm container pools to eliminate cold‑start penalties, batch multiple operations into single function calls to amortize serialization costs, and reuse named virtual environments rather than creating ephemeral sandboxes for each task.

### Is native execution safe for untrusted code in Cua?

No. Native execution runs directly on the host without isolation, making it unsuitable for untrusted code. The sandboxed mode, despite its performance overhead, provides essential security boundaries through Linux namespaces and containerization that native execution cannot offer.