Omni vs UI‑TARS: Differences Between VLM Model Options (omni, uitars, uitars‑mlx, uitars‑hf) in CUA

The omni option pipes screenshots through Microsoft’s OmniParser for UI element detection before sending them to an external LLM, while the three uitars variants run the ByteDance UI‑TARS vision‑language model locally—either via standard PyTorch (uitars/uitars‑hf), Apple‑silicon‑optimized MLX (uitars‑mlx), or a generic OpenAI‑compatible endpoint—depending on your hardware and privacy requirements.

The CUA (Computer Use Agent) framework from trycua/cua supports multiple vision‑language model (VLM) backends for GUI automation. Understanding the architectural differences between omni, uitars, uitars‑mlx, and uitars‑hf ensures you pick the right balance of accuracy, latency, and hardware utilization for your automated workflows.

How Each VLM Option Works

OmniParser (omni)

The omni option implements a two‑stage pipeline: screenshots are first processed by Microsoft’s OmniParser (located in libs/python/som/som/detect.py) to detect interactive UI elements, and the annotated image plus element IDs are forwarded to your choice of LLM (OpenAI, Anthropic, Ollama, etc.). This loop is implemented in libs/python/agent/cua_agent/loops/omniparser.py, which registers models matching the regex omni\+.*.

This approach decouples visual grounding from language generation, letting you use GPT‑4o, Claude 3.7 Sonnet, or local LLMs while still receiving precise pixel‑level element coordinates.

UI‑TARS Generic (uitars)

The uitars option treats UI‑TARS as a unified VLM that performs both element detection and response generation in a single forward pass. It routes requests through the generic adapter at libs/python/agent/cua_agent/adapters/models/generic.py, which can target a custom OpenAI‑compatible endpoint or a local HuggingFace model. Unlike omni, there is no separate detection step—UI‑TARS understands screen layout natively.

UI‑TARS MLX (uitars‑mlx)

uitars‑mlx is optimized specifically for Apple Silicon (M‑series chips). It uses the MLXVLMAdapter class defined in libs/python/agent/cua_agent/adapters/mlxvlm_adapter.py to load mlx‑community/UI‑TARS‑1.5‑7B‑4bit via the Apple‑optimized MLX‑VLM library. This provides low‑memory, high‑throughput inference without requiring CUDA or heavy PyTorch installations.

UI‑TARS HuggingFace (uitars‑hf)

uitars‑hf runs the same ByteDance UI‑TARS model (e.g., ByteDance-Seed/UI‑TARS‑1.5‑7B) through standard HuggingFace Transformers. The implementation lives in libs/python/agent/cua_agent/adapters/huggingfacelocal_adapter.py, which checks for transformers and accelerate dependencies (and optionally bitsandbytes for 4‑bit quantization). This backend works on any platform that supports PyTorch—Linux, Windows, or macOS—regardless of chip architecture.

Source Code Architecture

The model selection dropdown in the Gradio UI maps these strings to specific adapters in libs/python/agent/cua_agent/ui/gradio/ui_components.py (lines 70‑81). When you instantiate a ComputerAgent, the prefix in your model string determines which code path executes:

Each adapter validates dependencies on initialization. For example, huggingfacelocal_adapter.py raises an informative error prompting you to install cua‑agent[uitars‑hf] if transformers is missing, while mlxvlm_adapter.py requires the cua‑agent[uitars‑mlx] extra.

Installation Requirements

Install the corresponding extras based on your chosen backend:


# For omni (requires the SOM package + your LLM provider SDK)

pip install cua-agent[omni] openai anthropic

# For uitars-hf (PyTorch/Transformers backbone)

pip install cua-agent[uitars-hf]

# For uitars-mlx (Apple Silicon only)

pip install cua-agent[uitars-mlx]

The omni option additionally requires the cua-som package, which bundles the Microsoft OmniParser weights and inference code.

Code Examples

OmniParser with GPT‑4o

Use this when you want state‑of‑the‑art LLM reasoning combined with pixel‑perfect UI detection:

from cua import ComputerAgent, Sandbox, Image

async with Sandbox.ephemeral(Image.linux(), local=True) as sb:
    agent = ComputerAgent(
        model="omni+openai/gpt-4o",  # OmniParser first, then GPT-4o

        tools=[sb],
    )
    async for r in agent.run([{"role": "user", "content": "Open Calculator"}]):
        print(r["output"])

UI‑TARS via HuggingFace (uitars‑hf)

Best for local GPU inference without external API calls:

from cua import ComputerAgent, Sandbox, Image

async with Sandbox.ephemeral(Image.linux(), local=True) as sb:
    agent = ComputerAgent(
        model="uitars+ByteDance-Seed/UI-TARS-1.5-7B",
        tools=[sb],
    )
    async for r in agent.run([{"role": "user", "content": "Click the first button"}]):
        print(r["output"])

UI‑TARS via MLX (uitars‑mlx)

Optimized for MacBook Pro and Mac Studio deployments:

from cua import ComputerAgent, Sandbox, Image

async with Sandbox.ephemeral(Image.macos(), local=True) as sb:
    agent = ComputerAgent(
        model="uitars-mlx+mlx-community/UI-TARS-1.5-7B-4bit",
        tools=[sb],
    )
    async for r in agent.run([{"role": "user", "content": "Resize the window"}]):
        print(r["output"])

Choosing the Right VLM

  • Choose omni when you need the reasoning capabilities of GPT‑4o or Claude 3.7 Sonnet but still require accurate UI element grounding via the OmniParser detector.
  • Choose uitars‑hf when you want a self‑hosted, single‑model solution that runs on any CUDA‑capable machine or CPU without relying on external APIs.
  • Choose uitars‑mlx exclusively for Apple Silicon Macs where MLX delivers faster token generation and lower RAM usage than PyTorch equivalents.
  • Choose uitars (generic) when connecting to a custom OpenAI‑compatible endpoint (e.g., vLLM, TGI) that hosts UI‑TARS or a fine‑tuned variant.

Summary

Frequently Asked Questions

Can I use uitars‑mlx on an Intel Mac or Linux machine?

No. The uitars‑mlx option depends on the MLX framework, which is exclusively designed for Apple Silicon (M1, M2, M3, M4 chips). Intel Macs and Linux systems should use uitars‑hf with standard PyTorch/Transformers instead.

Does the omni option work offline?

Only partially. While the OmniParser detection stage in libs/python/som/som/detect.py runs locally, the LLM stage (openai/gpt-4o, anthropic/claude-3-7-sonnet, etc.) requires an active internet connection to the respective API provider. For fully offline operation, use uitars‑hf or uitars‑mlx.

Which option provides the lowest latency on consumer hardware?

uitars‑mlx typically offers the lowest latency on Apple Silicon Macs due to MLX’s memory efficiency and unified memory architecture. On CUDA‑equipped Linux machines, uitars‑hf with 4‑bit quantization (bitsandbytes) provides the fastest local inference, while omni introduces additional latency from the two‑stage detection process.

What is the difference between uitars and uitars‑hf?

Technically, both use the same underlying model (ByteDance UI‑TARS). The distinction lies in the installation extra and adapter path: uitars is the generic entry point that can resolve to either a remote endpoint or a local model, whereas uitars‑hf specifically loads the checkpoint via huggingfacelocal_adapter.py and enforces the transformers dependency. In practice, model strings prefixed with uitars+ usually map to the same HuggingFace loader unless you specify a custom base URL.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →