How Cua's Multi-Provider VLM Architecture Supports OpenAI, Anthropic, Gemini, Qwen, and UITARS
Cua leverages a plug-in adapter system built on LiteLLM's CustomLLM interface to normalize cloud APIs and local vision-language models into a unified OpenAI-compatible format, enabling seamless switching between providers without code changes.
The trycua/cua repository implements a flexible multi-provider VLM architecture that abstracts away differences between remote inference services and on-device execution. This design allows computer-use agents to route vision-language tasks to OpenAI, Anthropic, Google Gemini, or local models like Qwen 2.5-VL and UITARS using a consistent message schema.
The Adapter Abstraction Layer
All VLM adapters in Cua inherit from litellm.llms.custom_llm.CustomLLM and reside under libs/python/agent/cua_agent/adapters/. The central registry in adapters/__init__.py exports these implementations, providing a consistent factory interface for the agent core.
The key adapters include:
MLXVLMAdapter– Handles local inference on Apple Silicon via the mlx-vlm library. It converts OpenAI-style message lists (containing base-64 images) into MLX-specific tensors and calls the model'sgeneratemethod.HuggingFaceLocalAdapter– A generic wrapper for Hugging Face transformers with vision-language capabilities. Concrete handlers live inadapters/models/:Qwen2_5_VLModelmanages Qwen 2.5-VL variants (e.g.,Qwen/Qwen2.5-VL-7B-Instruct)InternVLModelmanages InternVL variants
CUAAdapter– Routes requests to Cua's remote inference service for models prefixed withcua/, injecting bearer tokens and version headers.
The HuggingFace model handlers share a factory in adapters/models/__init__.py that inspects the Hugging Face config class name to instantiate the correct processor.
Cloud Provider Integration
For OpenAI, Anthropic, and Gemini, Cua utilizes provider-specific loops that construct OpenAI-compatible tool schemas. These reside in libs/python/agent/cua_agent/loops/openai.py, loops/anthropic.py, and loops/gemini.py.
Each loop implements _prepare_tools_for_openai (defined around lines 47-65 in loops/openai.py) to generate the "computer" function description required for vision tasks. When processing messages:
- OpenAI models (e.g.,
gpt-4o) receive base-64 image URLs embedded in the content list, serialized automatically by LiteLLM. - Anthropic Claude 3.5 Sonnet and Gemini 1.5-pro receive analogous handling through their respective loops, reusing the same tool preparation helper to ensure schema consistency.
Because these providers speak the OpenAI message format natively, the agent routes to them via LiteLLM without requiring custom adapter translation.
Local VLM Support
MLX-VLM and UITARS on Apple Silicon
UITARS and other MLX-compatible models run locally through MLXVLMAdapter (libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py). This adapter implements two critical transformations (lines 81-87):
- Message conversion – The
_convert_openai_to_mlxmethod extractsimageobjects from the OpenAI-style message list and converts them intotorchtensors suitable for the MLX runtime. - Generation – Inference runs under
torch.no_grad()context, returning the first decoded string from the model's output.
To activate this adapter, prefix the model name with mlx/ when constructing the agent.
HuggingFace Local Models
The HuggingFaceLocalAdapter supports Qwen 2.5-VL and InternVL through specialized model classes in adapters/models/. When loading a model, the factory in adapters/models/__init__.py examines the Hugging Face configuration to return either Qwen2_5_VLModel or InternVLModel, handling tokenizer and processor initialization automatically.
These models accept the same OpenAI-style message format, with the adapter managing the conversion to Hugging Face chat_template expectations internally.
Provider Selection in the Agent Core
The central agent (libs/python/agent/cua_agent/agent.py) implements dispatch logic that instantiates the appropriate adapter based on model prefixes:
if model.startswith("mlx/"):
vlm_adapter = MLXVLMAdapter()
elif model.startswith(("huggingface/", "qwen2_5_vl", "internvl")):
vlm_adapter = HuggingFaceLocalAdapter()
elif model.startswith("cua/"):
vlm_adapter = CUAAdapter()
# Fallback to OpenAI/Anthropic/Gemini via LiteLLm standard routes
This prefix-based routing (importing MLXVLMAdapter at line 35 in the agent file) ensures that cloud providers use optimized API loops while local models load the correct inference backend.
Practical Implementation Examples
# OpenAI GPT-4o (cloud)
from cua.agent import ComputerAgent
openai_agent = ComputerAgent(
model="openai/gpt-4o",
temperature=0.2,
)
# Anthropic Claude 3.5 Sonnet (cloud)
anthropic_agent = ComputerAgent(
model="anthropic/claude-3-5-sonnet",
temperature=0.0,
)
# Google Gemini 1.5-pro (cloud)
gemini_agent = ComputerAgent(
model="gemini/gemini-1.5-pro",
)
# Qwen 2.5-VL (local HuggingFace)
from cua.agent.adapters.models import load_model
qwen = load_model(
model_name="Qwen/Qwen2.5-VL-7B-Instruct",
device="auto",
trust_remote_code=True,
)
# UITARS via MLX (local Apple Silicon)
from cua.agent.adapters.mlxvlm_adapter import MLXVLMAdapter
mlx_adapter = MLXVLMAdapter()
response = mlx_adapter.generate(messages) # OpenAI-style message format
Summary
- Unified Interface: All adapters inherit from
litellm.llms.custom_llm.CustomLLM, providing a consistentgeneratemethod signature across providers. - Prefix-Based Routing: The agent selects backends using model name prefixes (
mlx/,huggingface/,openai/, etc.) defined inagent.py. - Cloud Optimization: OpenAI, Anthropic, and Gemini use dedicated loops in
libs/python/agent/cua_agent/loops/that leverage LiteLLM's native serialization. - Local Flexibility: MLX-VLM support enables UITARS inference on Apple Silicon, while the HuggingFace factory pattern supports Qwen and InternVL variants.
- Format Consistency: All providers receive OpenAI-compatible message lists, with adapters handling translation only for local models that require it.
Frequently Asked Questions
How does Cua handle different message formats between cloud and local VLMs?
Cua standardizes on the OpenAI message format (a list of {role, content} dictionaries where content may contain image URLs). Cloud providers receive this format directly via LiteLLM. Local adapters like MLXVLMAdapter and HuggingFaceLocalAdapter implement private conversion methods (e.g., _convert_openai_to_mlx) to transform these messages into backend-specific tensors or chat templates before inference.
What hardware is required to run UITARS locally?
UITARS runs through the MLXVLMAdapter, which requires Apple Silicon (M-series chips) and the mlx-vlm library. The adapter utilizes torch.no_grad() optimization and MLX runtime acceleration to execute vision-language inference entirely on-device without cloud dependencies.
How do I switch between OpenAI and a local Qwen model in the same codebase?
Change the model parameter when initializing ComputerAgent:
# Cloud
agent = ComputerAgent(model="openai/gpt-4o")
# Local
agent = ComputerAgent(model="huggingface/Qwen/Qwen2.5-VL-7B-Instruct")
The agent constructor in libs/python/agent/cua_agent/agent.py automatically instantiates the correct adapter based on the prefix, requiring no other code changes.
Why does Cua use LiteLLM's CustomLLM as the base class for adapters?
Inheriting from litellm.llms.custom_llm.CustomLLM provides a standard interface that LiteLLM recognizes, allowing Cua to integrate local models into existing LiteLLM routing logic while maintaining compatibility with LiteLLM's retry, logging, and error-handling mechanisms. This base class enforces the generate method signature that the agent expects across all providers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →