# How Cua's Multi-Provider VLM Architecture Supports OpenAI, Anthropic, Gemini, Qwen, and UITARS

> Discover how Cua's multi-provider VLM architecture seamlessly integrates OpenAI, Anthropic, Gemini, Qwen, and local models like UITARS using a unified OpenAI-compatible format. Switch providers effortlessly.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: architecture
- Published: 2026-04-27

---

**Cua leverages a plug-in adapter system built on LiteLLM's `CustomLLM` interface to normalize cloud APIs and local vision-language models into a unified OpenAI-compatible format, enabling seamless switching between providers without code changes.**

The trycua/cua repository implements a flexible multi-provider VLM architecture that abstracts away differences between remote inference services and on-device execution. This design allows computer-use agents to route vision-language tasks to OpenAI, Anthropic, Google Gemini, or local models like Qwen 2.5-VL and UITARS using a consistent message schema.

## The Adapter Abstraction Layer

All VLM adapters in Cua inherit from **`litellm.llms.custom_llm.CustomLLM`** and reside under `libs/python/agent/cua_agent/adapters/`. The central registry in **[`adapters/__init__.py`](https://github.com/trycua/cua/blob/main/adapters/__init__.py)** exports these implementations, providing a consistent factory interface for the agent core.

The key adapters include:

- **`MLXVLMAdapter`** – Handles local inference on Apple Silicon via the *mlx-vlm* library. It converts OpenAI-style message lists (containing base-64 images) into MLX-specific tensors and calls the model's `generate` method.
- **`HuggingFaceLocalAdapter`** – A generic wrapper for Hugging Face transformers with vision-language capabilities. Concrete handlers live in `adapters/models/`:
  - **`Qwen2_5_VLModel`** manages Qwen 2.5-VL variants (e.g., `Qwen/Qwen2.5-VL-7B-Instruct`)
  - **`InternVLModel`** manages InternVL variants
- **`CUAAdapter`** – Routes requests to Cua's remote inference service for models prefixed with `cua/`, injecting bearer tokens and version headers.

The HuggingFace model handlers share a factory in **[`adapters/models/__init__.py`](https://github.com/trycua/cua/blob/main/adapters/models/__init__.py)** that inspects the Hugging Face config class name to instantiate the correct processor.

## Cloud Provider Integration

For **OpenAI**, **Anthropic**, and **Gemini**, Cua utilizes provider-specific loops that construct OpenAI-compatible tool schemas. These reside in [`libs/python/agent/cua_agent/loops/openai.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/loops/openai.py), [`loops/anthropic.py`](https://github.com/trycua/cua/blob/main/loops/anthropic.py), and [`loops/gemini.py`](https://github.com/trycua/cua/blob/main/loops/gemini.py).

Each loop implements **`_prepare_tools_for_openai`** (defined around lines 47-65 in [`loops/openai.py`](https://github.com/trycua/cua/blob/main/loops/openai.py)) to generate the `"computer"` function description required for vision tasks. When processing messages:

- **OpenAI** models (e.g., `gpt-4o`) receive base-64 image URLs embedded in the content list, serialized automatically by LiteLLM.
- **Anthropic** Claude 3.5 Sonnet and **Gemini** 1.5-pro receive analogous handling through their respective loops, reusing the same tool preparation helper to ensure schema consistency.

Because these providers speak the OpenAI message format natively, the agent routes to them via LiteLLM without requiring custom adapter translation.

## Local VLM Support

### MLX-VLM and UITARS on Apple Silicon

UITARS and other MLX-compatible models run locally through **`MLXVLMAdapter`** ([`libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py)). This adapter implements two critical transformations (lines 81-87):

1. **Message conversion** – The `_convert_openai_to_mlx` method extracts `image` objects from the OpenAI-style message list and converts them into `torch` tensors suitable for the MLX runtime.
2. **Generation** – Inference runs under `torch.no_grad()` context, returning the first decoded string from the model's output.

To activate this adapter, prefix the model name with `mlx/` when constructing the agent.

### HuggingFace Local Models

The **`HuggingFaceLocalAdapter`** supports Qwen 2.5-VL and InternVL through specialized model classes in `adapters/models/`. When loading a model, the factory in [`adapters/models/__init__.py`](https://github.com/trycua/cua/blob/main/adapters/models/__init__.py) examines the Hugging Face configuration to return either `Qwen2_5_VLModel` or `InternVLModel`, handling tokenizer and processor initialization automatically.

These models accept the same OpenAI-style message format, with the adapter managing the conversion to Hugging Face `chat_template` expectations internally.

## Provider Selection in the Agent Core

The central agent ([`libs/python/agent/cua_agent/agent.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/agent.py)) implements dispatch logic that instantiates the appropriate adapter based on model prefixes:

```python
if model.startswith("mlx/"):
    vlm_adapter = MLXVLMAdapter()
elif model.startswith(("huggingface/", "qwen2_5_vl", "internvl")):
    vlm_adapter = HuggingFaceLocalAdapter()
elif model.startswith("cua/"):
    vlm_adapter = CUAAdapter()

# Fallback to OpenAI/Anthropic/Gemini via LiteLLm standard routes

```

This prefix-based routing (importing `MLXVLMAdapter` at line 35 in the agent file) ensures that cloud providers use optimized API loops while local models load the correct inference backend.

## Practical Implementation Examples

```python

# OpenAI GPT-4o (cloud)

from cua.agent import ComputerAgent

openai_agent = ComputerAgent(
    model="openai/gpt-4o",
    temperature=0.2,
)

# Anthropic Claude 3.5 Sonnet (cloud)

anthropic_agent = ComputerAgent(
    model="anthropic/claude-3-5-sonnet",
    temperature=0.0,
)

# Google Gemini 1.5-pro (cloud)

gemini_agent = ComputerAgent(
    model="gemini/gemini-1.5-pro",
)

# Qwen 2.5-VL (local HuggingFace)

from cua.agent.adapters.models import load_model

qwen = load_model(
    model_name="Qwen/Qwen2.5-VL-7B-Instruct",
    device="auto",
    trust_remote_code=True,
)

# UITARS via MLX (local Apple Silicon)

from cua.agent.adapters.mlxvlm_adapter import MLXVLMAdapter

mlx_adapter = MLXVLMAdapter()
response = mlx_adapter.generate(messages)  # OpenAI-style message format

```

## Summary

- **Unified Interface**: All adapters inherit from `litellm.llms.custom_llm.CustomLLM`, providing a consistent `generate` method signature across providers.
- **Prefix-Based Routing**: The agent selects backends using model name prefixes (`mlx/`, `huggingface/`, `openai/`, etc.) defined in [`agent.py`](https://github.com/trycua/cua/blob/main/agent.py).
- **Cloud Optimization**: OpenAI, Anthropic, and Gemini use dedicated loops in `libs/python/agent/cua_agent/loops/` that leverage LiteLLM's native serialization.
- **Local Flexibility**: MLX-VLM support enables UITARS inference on Apple Silicon, while the HuggingFace factory pattern supports Qwen and InternVL variants.
- **Format Consistency**: All providers receive OpenAI-compatible message lists, with adapters handling translation only for local models that require it.

## Frequently Asked Questions

### How does Cua handle different message formats between cloud and local VLMs?

Cua standardizes on the OpenAI message format (a list of `{role, content}` dictionaries where content may contain image URLs). **Cloud providers** receive this format directly via LiteLLM. **Local adapters** like `MLXVLMAdapter` and `HuggingFaceLocalAdapter` implement private conversion methods (e.g., `_convert_openai_to_mlx`) to transform these messages into backend-specific tensors or chat templates before inference.

### What hardware is required to run UITARS locally?

UITARS runs through the `MLXVLMAdapter`, which requires **Apple Silicon** (M-series chips) and the *mlx-vlm* library. The adapter utilizes `torch.no_grad()` optimization and MLX runtime acceleration to execute vision-language inference entirely on-device without cloud dependencies.

### How do I switch between OpenAI and a local Qwen model in the same codebase?

Change the `model` parameter when initializing `ComputerAgent`:

```python

# Cloud

agent = ComputerAgent(model="openai/gpt-4o")

# Local

agent = ComputerAgent(model="huggingface/Qwen/Qwen2.5-VL-7B-Instruct")

```

The agent constructor in [`libs/python/agent/cua_agent/agent.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/agent.py) automatically instantiates the correct adapter based on the prefix, requiring no other code changes.

### Why does Cua use LiteLLM's CustomLLM as the base class for adapters?

Inheriting from **`litellm.llms.custom_llm.CustomLLM`** provides a standard interface that LiteLLM recognizes, allowing Cua to integrate local models into existing LiteLLM routing logic while maintaining compatibility with LiteLLM's retry, logging, and error-handling mechanisms. This base class enforces the `generate` method signature that the agent expects across all providers.