# Omni vs UI‑TARS: Differences Between VLM Model Options (omni, uitars, uitars‑mlx, uitars‑hf) in CUA

> Explore the differences between CUA's VLM model options including OmniParser and UI-TARS variants (uitars, uitars-mlx, uitars-hf). Choose the best fit for your hardware & privacy needs.

- Repository: [Cua/cua](https://github.com/trycua/cua)
- Tags: deep-dive
- Published: 2026-04-27

---

**The `omni` option pipes screenshots through Microsoft’s OmniParser for UI element detection before sending them to an external LLM, while the three `uitars` variants run the ByteDance UI‑TARS vision‑language model locally—either via standard PyTorch (`uitars`/`uitars‑hf`), Apple‑silicon‑optimized MLX (`uitars‑mlx`), or a generic OpenAI‑compatible endpoint—depending on your hardware and privacy requirements.**

The CUA (Computer Use Agent) framework from `trycua/cua` supports multiple vision‑language model (VLM) backends for GUI automation. Understanding the architectural differences between **omni**, **uitars**, **uitars‑mlx**, and **uitars‑hf** ensures you pick the right balance of accuracy, latency, and hardware utilization for your automated workflows.

## How Each VLM Option Works

### OmniParser (omni)

The **`omni`** option implements a two‑stage pipeline: screenshots are first processed by Microsoft’s **OmniParser** (located in [`libs/python/som/som/detect.py`](https://github.com/trycua/cua/blob/main/libs/python/som/som/detect.py)) to detect interactive UI elements, and the annotated image plus element IDs are forwarded to your choice of LLM (OpenAI, Anthropic, Ollama, etc.). This loop is implemented in [`libs/python/agent/cua_agent/loops/omniparser.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/loops/omniparser.py), which registers models matching the regex `omni\+.*`.

This approach decouples visual grounding from language generation, letting you use GPT‑4o, Claude 3.7 Sonnet, or local LLMs while still receiving precise pixel‑level element coordinates.

### UI‑TARS Generic (uitars)

The **`uitars`** option treats UI‑TARS as a unified VLM that performs both element detection and response generation in a single forward pass. It routes requests through the generic adapter at [`libs/python/agent/cua_agent/adapters/models/generic.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/adapters/models/generic.py), which can target a custom OpenAI‑compatible endpoint or a local HuggingFace model. Unlike `omni`, there is no separate detection step—UI‑TARS understands screen layout natively.

### UI‑TARS MLX (uitars‑mlx)

**`uitars‑mlx`** is optimized specifically for Apple Silicon (M‑series chips). It uses the `MLXVLMAdapter` class defined in [`libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py) to load `mlx‑community/UI‑TARS‑1.5‑7B‑4bit` via the Apple‑optimized MLX‑VLM library. This provides low‑memory, high‑throughput inference without requiring CUDA or heavy PyTorch installations.

### UI‑TARS HuggingFace (uitars‑hf)

**`uitars‑hf`** runs the same ByteDance UI‑TARS model (e.g., `ByteDance-Seed/UI‑TARS‑1.5‑7B`) through standard HuggingFace Transformers. The implementation lives in [`libs/python/agent/cua_agent/adapters/huggingfacelocal_adapter.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/adapters/huggingfacelocal_adapter.py), which checks for `transformers` and `accelerate` dependencies (and optionally `bitsandbytes` for 4‑bit quantization). This backend works on any platform that supports PyTorch—Linux, Windows, or macOS—regardless of chip architecture.

## Source Code Architecture

The model selection dropdown in the Gradio UI maps these strings to specific adapters in [`libs/python/agent/cua_agent/ui/gradio/ui_components.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/ui/gradio/ui_components.py) (lines 70‑81). When you instantiate a `ComputerAgent`, the prefix in your model string determines which code path executes:

- **`omni+`** → [`omniparser.py`](https://github.com/trycua/cua/blob/main/omniparser.py) → `OmniParser` class ([`som/detect.py`](https://github.com/trycua/cua/blob/main/som/detect.py))
- **`uitars+`** → [`generic.py`](https://github.com/trycua/cua/blob/main/generic.py) → HuggingFace or custom endpoint
- **`uitars‑mlx+`** → [`mlxvlm_adapter.py`](https://github.com/trycua/cua/blob/main/mlxvlm_adapter.py) → MLX‑VLM runtime
- **`uitars‑hf+`** → [`huggingfacelocal_adapter.py`](https://github.com/trycua/cua/blob/main/huggingfacelocal_adapter.py) → Transformers pipeline

Each adapter validates dependencies on initialization. For example, [`huggingfacelocal_adapter.py`](https://github.com/trycua/cua/blob/main/huggingfacelocal_adapter.py) raises an informative error prompting you to install `cua‑agent[uitars‑hf]` if `transformers` is missing, while [`mlxvlm_adapter.py`](https://github.com/trycua/cua/blob/main/mlxvlm_adapter.py) requires the `cua‑agent[uitars‑mlx]` extra.

## Installation Requirements

Install the corresponding extras based on your chosen backend:

```bash

# For omni (requires the SOM package + your LLM provider SDK)

pip install cua-agent[omni] openai anthropic

# For uitars-hf (PyTorch/Transformers backbone)

pip install cua-agent[uitars-hf]

# For uitars-mlx (Apple Silicon only)

pip install cua-agent[uitars-mlx]

```

The `omni` option additionally requires the `cua-som` package, which bundles the Microsoft OmniParser weights and inference code.

## Code Examples

### OmniParser with GPT‑4o

Use this when you want state‑of‑the‑art LLM reasoning combined with pixel‑perfect UI detection:

```python
from cua import ComputerAgent, Sandbox, Image

async with Sandbox.ephemeral(Image.linux(), local=True) as sb:
    agent = ComputerAgent(
        model="omni+openai/gpt-4o",  # OmniParser first, then GPT-4o

        tools=[sb],
    )
    async for r in agent.run([{"role": "user", "content": "Open Calculator"}]):
        print(r["output"])

```

### UI‑TARS via HuggingFace (uitars‑hf)

Best for local GPU inference without external API calls:

```python
from cua import ComputerAgent, Sandbox, Image

async with Sandbox.ephemeral(Image.linux(), local=True) as sb:
    agent = ComputerAgent(
        model="uitars+ByteDance-Seed/UI-TARS-1.5-7B",
        tools=[sb],
    )
    async for r in agent.run([{"role": "user", "content": "Click the first button"}]):
        print(r["output"])

```

### UI‑TARS via MLX (uitars‑mlx)

Optimized for MacBook Pro and Mac Studio deployments:

```python
from cua import ComputerAgent, Sandbox, Image

async with Sandbox.ephemeral(Image.macos(), local=True) as sb:
    agent = ComputerAgent(
        model="uitars-mlx+mlx-community/UI-TARS-1.5-7B-4bit",
        tools=[sb],
    )
    async for r in agent.run([{"role": "user", "content": "Resize the window"}]):
        print(r["output"])

```

## Choosing the Right VLM

- **Choose `omni`** when you need the reasoning capabilities of GPT‑4o or Claude 3.7 Sonnet but still require accurate UI element grounding via the `OmniParser` detector.
- **Choose `uitars‑hf`** when you want a self‑hosted, single‑model solution that runs on any CUDA‑capable machine or CPU without relying on external APIs.
- **Choose `uitars‑mlx`** exclusively for Apple Silicon Macs where MLX delivers faster token generation and lower RAM usage than PyTorch equivalents.
- **Choose `uitars`** (generic) when connecting to a custom OpenAI‑compatible endpoint (e.g., vLLM, TGI) that hosts UI‑TARS or a fine‑tuned variant.

## Summary

- **`omni`** splits work between Microsoft OmniParser (visual grounding) and an external LLM (reasoning), configured in [`libs/python/agent/cua_agent/loops/omniparser.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/loops/omniparser.py).
- **`uitars`** routes to a unified UI‑TARS endpoint via [`libs/python/agent/cua_agent/adapters/models/generic.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/adapters/models/generic.py), supporting both local HuggingFace and remote OpenAI‑compatible servers.
- **`uitars‑mlx`** targets Apple Silicon using [`libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py) for MLX‑optimized inference.
- **`uitars‑hf`** runs standard PyTorch/Transformers inference through [`libs/python/agent/cua_agent/adapters/huggingfacelocal_adapter.py`](https://github.com/trycua/cua/blob/main/libs/python/agent/cua_agent/adapters/huggingfacelocal_adapter.py), requiring the `cua‑agent[uitars‑hf]` dependency.

## Frequently Asked Questions

### Can I use uitars‑mlx on an Intel Mac or Linux machine?

No. The `uitars‑mlx` option depends on the MLX framework, which is exclusively designed for Apple Silicon (M1, M2, M3, M4 chips). Intel Macs and Linux systems should use `uitars‑hf` with standard PyTorch/Transformers instead.

### Does the omni option work offline?

Only partially. While the OmniParser detection stage in [`libs/python/som/som/detect.py`](https://github.com/trycua/cua/blob/main/libs/python/som/som/detect.py) runs locally, the LLM stage (`openai/gpt-4o`, `anthropic/claude-3-7-sonnet`, etc.) requires an active internet connection to the respective API provider. For fully offline operation, use `uitars‑hf` or `uitars‑mlx`.

### Which option provides the lowest latency on consumer hardware?

**`uitars‑mlx`** typically offers the lowest latency on Apple Silicon Macs due to MLX’s memory efficiency and unified memory architecture. On CUDA‑equipped Linux machines, `uitars‑hf` with 4‑bit quantization (`bitsandbytes`) provides the fastest local inference, while `omni` introduces additional latency from the two‑stage detection process.

### What is the difference between uitars and uitars‑hf?

Technically, both use the same underlying model (ByteDance UI‑TARS). The distinction lies in the installation extra and adapter path: `uitars` is the generic entry point that can resolve to either a remote endpoint or a local model, whereas `uitars‑hf` specifically loads the checkpoint via [`huggingfacelocal_adapter.py`](https://github.com/trycua/cua/blob/main/huggingfacelocal_adapter.py) and enforces the `transformers` dependency. In practice, model strings prefixed with `uitars+` usually map to the same HuggingFace loader unless you specify a custom base URL.