How to Automate Browser Interactions with cmux In-App Browser API

The cmux in-app browser API exposes a JSON-RPC command surface that lets scripts and AI agents programmatically control WKWebView panels via CLI commands like browser.open_split, browser.click, and browser.snapshot for native headless automation within the terminal multiplexer.

cmux is a terminal-centric macOS window manager that embeds a full-featured browser engine directly into its workspace panes. The cmux in-app browser API enables you to automate web interactions—navigation, form filling, screenshots, and JavaScript execution—using stable surface handles and JSON-RPC commands routed through the daemon's V2 protocol in Sources/TerminalController.swift.

Understanding the Browser Architecture

The browser automation system follows a three-tier architecture: CLI parsing, V2 protocol dispatch, and WebView execution. All commands operate on surface IDs (or short refs like surface:3) that remain stable even when panes are moved or reorganized.

CLI to V2 Protocol Pipeline

When you run cmux browser … commands, the CLI entry point in CLI/cmux.swift parses your input and constructs a JSON-RPC payload. This payload routes to TerminalController.handleV2Method, which dispatches to specific browser handlers like v2BrowserOpenSplit, v2BrowserClick, or v2BrowserEval defined in Sources/TerminalController.swift.

Surface Resolution and Handle Semantics

Every browser command resolves targets through v2ResolveSurfaceId (called via v2ResolveTabManager → v2ResolveWorkspace). The system returns handle refs (window_ref, workspace_ref, pane_ref, surface_ref) that act as persistent identifiers. According to the agent-browser specification in docs/agent-browser-port-spec.md, these UUID-based handles never change during surface moves, ensuring reliable automation scripts.

The actual WebView interaction occurs in Sources/Panels/BrowserPanel.swift, which hosts the native WKWebView and provides methods for navigation, snapshots, and proxy management that the V2 methods consume.

Core Browser Automation Commands

The browser API implements the agent-browser command surface documented in docs/agent-browser-port-spec.md. Each command accepts a --surface parameter and returns JSON containing updated handle refs.

Surface Identification and Navigation

Before automating, obtain the target surface ID using the identity system:

cmux identify --json

This returns the surface_id of the currently focused browser pane. Use this ID to open new splits or navigate:


# Open a URL in a new split (or reuse the right-most pane)

cmux browser --surface surface:b8d2 open_split url=https://developer.apple.com

# Navigate an existing surface to a new URL

cmux browser --surface surface:4f12 navigate url=https://github.com/manaflow-ai/cmux

The v2BrowserOpenSplit function (lines 7369+ in TerminalController.swift) handles placement logic, creating splits when necessary or routing to external browsers based on the "external open" branch logic.

Element Interaction (Click, Type, Fill)

Interact with DOM elements using CSS selectors. The v2BrowserClick implementation (line 8157) injects JavaScript to perform the click event:


# Click an element by CSS selector

cmux browser --surface surface:4f12 click selector="a[href$='README.md']"

# Type text into form fields

cmux browser --surface surface:4f12 type selector="#search" text="cmux automation"

# Fill forms with specific values

cmux browser --surface surface:4f12 fill selector="#username" text="admin"

These commands generate JavaScript snippets executed via BrowserPanel.executeJavaScript, providing synchronous feedback on success or selector-not-found errors.

Waiting and Screenshot Capture

Automation scripts often need to pause for page loads or element appearance. The v2BrowserWait command polls for selector presence:


# Wait up to 5000ms for an element to appear

cmux browser --surface surface:4f12 wait selector="#results" timeout=5000

# Capture the visible viewport as PNG

cmux browser --surface surface:4f12 screenshot output=page.png

The v2BrowserSnapshot function (line 7781) interfaces with BrowserPanel to rasterize the WKWebView content, returning image data through the JSON-RPC response.

JavaScript Execution

For complex automation beyond built-in commands, execute arbitrary JavaScript and retrieve results:

cmux browser --surface surface:4f12 eval script="document.title"

Returns:

{ "result": "cmux – a terminal‑centric macOS window manager" }

This uses v2BrowserEval to pipe scripts through the WebView's JavaScript engine and serialize return values back to the CLI.

Complete Automation Script Example

The following bash script demonstrates a full workflow: identifying the current surface, opening a search query in a split, interacting with forms, and capturing results:

#!/usr/bin/env bash
set -euo pipefail

# Identify the currently focused browser surface

SURF=$(cmux identify --json | jq -r '.focused.surface_id')
if [[ -z "$SURF" ]]; then
  echo "No focused browser surface." >&2; exit 1
fi

# Open DuckDuckGo in a new split

cmux browser --surface "$SURF" open_split \
  url="https://duckduckgo.com/?q=cmux+automation"

# Get the new surface ID from the focused pane

NEW_SURF=$(cmux identify --json | jq -r '.focused.surface_id')

# Wait for the search box to render

cmux browser --surface "$NEW_SURF" wait \
  selector="#search_form_input_homepage" timeout=8000

# Type query and submit

cmux browser --surface "$NEW_SURF" type \
  selector="#search_form_input_homepage" text="cmux screenshots"
cmux browser --surface "$NEW_SURF" press key="Enter"

# Wait for results and capture screenshot

cmux browser --surface "$NEW_SURF" wait \
  selector=".result" timeout=5000
cmux browser --surface "$NEW_SURF" screenshot output="cmux-results.png"

# Close the browser surface when done

cmux surface close "$NEW_SURF"

This script leverages the stable handle semantics—surface_id remains valid throughout the session even if the pane is repositioned within the workspace.

Key Implementation Files

Understanding the source structure helps when debugging automation failures or extending functionality:

Summary

  • The cmux in-app browser API provides native WKWebView automation through a JSON-RPC interface exposed via the CLI.
  • All commands flow through TerminalController.swift and operate on stable surface IDs that persist across workspace reorganizations.
  • Available automation primitives include navigation, element interaction (click/type/fill), explicit waiting, screenshots, and arbitrary JavaScript execution.
  • The system supports headless-style automation within a visible or backgrounded macOS terminal multiplexer, making it ideal for LLM-driven workflows that require web data extraction.

Frequently Asked Questions

How do I obtain the surface ID for browser automation?

Run cmux identify --json to retrieve the currently focused surface. The output includes surface_id (full UUID) and short refs (e.g., surface:3). Use either format with the --surface flag in subsequent browser commands. The ID remains stable even if you move the pane to a different workspace or split configuration.

What is the difference between browser.open_split and browser.navigate?

browser.open_split (implemented in v2BrowserOpenSplit) creates a new pane or reuses the right-most pane to open a URL, returning a new surface_id. browser.navigate (via v2BrowserNavigate) changes the URL of an existing browser surface without creating new panes. Use open_split when you need side-by-side browser instances; use navigate for sequential page loads within the same pane.

Can I automate cmux browser interactions from Python or Node.js?

Yes. The CLI interface wraps a JSON-RPC protocol over Unix domain sockets. Any language that can spawn subprocesses or open sockets can send payloads matching the schema in docs/agent-browser-port-spec.md. The Python test suite in tests_v2/test_browser_api_comprehensive.py demonstrates the exact JSON structure expected by TerminalController.swift.

Why does my screenshot command return blank images?

Screenshots require the browser surface to have completed rendering. Ensure you call browser.wait with a relevant CSS selector or sufficient timeout before browser.screenshot to guarantee the WKWebView has finished layout and paint operations. The v2BrowserSnapshot function captures the current composited layer; if the page is still loading, you may capture an incomplete or transparent buffer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →