Apache Maka Computer Use Capabilities: A Complete Guide to the maka_computer Tool

Apache Maka provides semantic desktop automation through the maka_computer tool, enabling LLMs to observe UI elements and perform structured actions like clicking, typing, and scrolling while enforcing strict safety policies.

Apache Maka’s computer use capabilities allow large language models to interact with desktop applications through a provider-neutral model-loop. The system exposes these features via the maka_computer tool, which implements an accessibility-first approach that captures UI context and executes semantic actions rather than raw pixel coordinates. This architecture ensures that automated interactions remain deterministic, auditable, and compliant with platform permission models like macOS TCC.

Core Computer Use Capabilities

The maka_computer tool supports twelve distinct operations defined in the JSON schema at packages/runtime/src/computer-use-tools.ts. These capabilities enable end-to-end desktop automation while maintaining strict safety boundaries.

Application Management Capabilities

list_apps returns a filtered or unfiltered list of running applications, allowing the model to identify automation targets by bundle identifier or process name. launch_app starts new applications by specifying the target bundle ID or display name.

Observation and Sensing

observe captures the Accessibility (AX) tree for a specified application or window, optionally including screenshots, menu structures, or conditional wait states. This operation generates a unique observation_id that subsequent actions must reference. screenshot captures bitmap images of specific windows, typically invoked internally by the observe action.

Element Interaction Actions

click_element performs primary clicks on UI elements using the combination of observation_id and element_id from a prior observation. set_value populates text fields with specified values, while select_text highlights specific text ranges within editable elements. secondary_action triggers context-menu operations on elements, and scroll_element navigates scrollable containers in specified directions with configurable amounts.

Workflow and Sequencing

element_sequence executes atomic action chains—such as click-then-wait-then-type—in a single API call, reducing latency for multi-step operations. window_action manipulates window geometry through move, resize, or minimize operations. wait pauses execution until specified text appears or disappears from the last observed window state, enabling synchronization with asynchronous UI updates.

How the maka_computer Tool Works

Maka implements a strict observe-then-act workflow that prevents blind automation. According to the design documents in docs/computer-use-model-loop-foundation.md, the model-loop requires a fresh observation ID for every action, ensuring the LLM operates on current UI state rather than stale coordinates.

Semantic Action Model

All capabilities are semantic only—they operate on accessibility metadata and element identifiers rather than screen coordinates. Raw pixel-level actions are deliberately disabled and fail-closed, preventing brittle automation that breaks across display resolutions or UI themes. This approach is enforced by the policies defined in apps/desktop/src/main/computer-use-real-model-policy.ts, which validate every action against TCC permissions and provenance requirements.

Safety and Budget Controls

Each action execution consumes budget from a per-operation allowance tracked by the runtime. The computer-use-real-model-policy.ts module enforces these limits alongside permission checks, ensuring that automated sessions cannot exceed authorized scope or duration.

Implementation Examples

CLI Usage with Real Model Execution

The following example demonstrates loading the computer-use toolset and executing a simple calculator workflow from the command line:


# Load the computer-use toolset first

npm run load-tools -- --group computer_use

# Perform a simple "open Calculator, click the '7' button"

npm run computer-use -- real-model \
  --scenario simple-calculator \
  --steps '
  - action: launch_app
    app: Calculator
  - action: observe
    app: Calculator
  - action: click_element
    observation_id: <from-previous-observe>
    element_id: 7
  - action: wait
    wait_for_text: "7"
'

This pattern appears in the Product Boundary section of docs/computer-use-model-loop-foundation.md, illustrating the typical maka_computer flow for real-model testing.

JavaScript Runtime API Integration

Developers can invoke computer use capabilities programmatically through the @maka/runtime package:

import { runTool } from '@maka/runtime';

// 1. Load the tool group
await runTool('load_tools', { group: 'computer_use' });

// 2. Observe the target app
const obs = await runTool('maka_computer', {
  action: 'observe',
  app: 'TextEdit',
  include_screenshot: false,
});

// 3. Click a button from the observation
await runTool('maka_computer', {
  action: 'click_element',
  observation_id: obs.id,
  element_id: obs.elements.find(e => e.label === 'Bold')!.id,
});

The runTool implementation for each action resides in packages/runtime/src/computer-use-tools.ts, handling schema validation and backend communication.

Skill Catalog Configuration

To expose computer use capabilities to the LLM, define a skill entry in the catalog:

name: Computer Use
description: |
  Use when the user asks to inspect or operate a local desktop application UI.
allowed-tools:
  - load_tools
  - maka_computer
required-tools:
  - maka_computer

When processing a request like "Please open Calculator and press the equals button," the model generates:

- tool: load_tools
  args: { group: "computer_use" }
- tool: maka_computer
  args:
    action: launch_app
    app: Calculator
- tool: maka_computer
  args:
    action: observe
    app: Calculator
- tool: maka_computer
  args:
    action: click_element
    observation_id: <id>
    element_id: "="

This skill definition is bundled in packages/runtime/src/bundled-skill-catalog.generated.ts.

Key Implementation Files

Understanding Maka's computer use capabilities requires familiarity with several critical source files:

Summary

  • Apache Maka's computer use capabilities are exposed exclusively through the maka_computer tool, which provides twelve semantic operations for desktop automation.

  • The architecture enforces an observe-then-act workflow where every action requires a valid observation_id from a prior accessibility tree capture.

  • Raw coordinate-based actions are disabled by design; all interactions use semantic element identifiers to ensure cross-platform stability and safety.

  • Implementation spans packages/runtime/src/computer-use-tools.ts for logic, packages/runtime/src/computer-use-types.ts for type safety, and apps/desktop/src/main/computer-use-real-model-policy.ts for permission enforcement.

  • Integration options include CLI execution, JavaScript runtime APIs, and declarative skill catalog configurations.

Frequently Asked Questions

What is the difference between observe and screenshot in Maka's computer use tools?

The observe action captures the full Accessibility (AX) tree structure of a target window or application, generating an observation_id and optional screenshot metadata. The screenshot action captures only the bitmap image of a window without accessibility context. While screenshot is available as a standalone capability, it is typically invoked internally by observe when the include_screenshot parameter is enabled, as defined in packages/runtime/src/computer-use-tools.ts.

Can Maka's computer use tools click on specific screen coordinates?

No. Maka deliberately disables raw coordinate or pixel-level actions as part of its safety policy outlined in docs/computer-use-model-loop-foundation.md. All click operations must use click_element with a valid element_id derived from a previous observe call. This semantic approach prevents automation failures across different screen resolutions and ensures compliance with accessibility standards.

How does Maka enforce safety and permissions during computer use sessions?

Safety enforcement occurs in apps/desktop/src/main/computer-use-real-model-policy.ts, which validates every action against platform TCC (Transparency, Consent, and Control) permissions, per-action budgets, and provenance requirements. The system maintains a strict "fail-closed" posture where unauthorized actions or attempts to bypass the observe-then-act workflow are rejected before execution.

What file contains the type definitions for computer use observations?

The TypeScript interfaces for observations, elements, and action parameters are defined in packages/runtime/src/computer-use-types.ts. This file establishes the contracts between the accessibility backend and the model-loop runtime, ensuring type safety for operations like set_value, select_text, and element_sequence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →