How to Automate Browser Interactions with cmux In-App Browser API
The cmux in-app browser API exposes a JSON-RPC command surface that lets scripts and AI agents programmatically control WKWebView panels via CLI commands like browser.open_split, browser.click, and browser.snapshot for native headless automation within the terminal multiplexer.
cmux is a terminal-centric macOS window manager that embeds a full-featured browser engine directly into its workspace panes. The cmux in-app browser API enables you to automate web interactions—navigation, form filling, screenshots, and JavaScript execution—using stable surface handles and JSON-RPC commands routed through the daemon's V2 protocol in Sources/TerminalController.swift.
Understanding the Browser Architecture
The browser automation system follows a three-tier architecture: CLI parsing, V2 protocol dispatch, and WebView execution. All commands operate on surface IDs (or short refs like surface:3) that remain stable even when panes are moved or reorganized.
CLI to V2 Protocol Pipeline
When you run cmux browser … commands, the CLI entry point in CLI/cmux.swift parses your input and constructs a JSON-RPC payload. This payload routes to TerminalController.handleV2Method, which dispatches to specific browser handlers like v2BrowserOpenSplit, v2BrowserClick, or v2BrowserEval defined in Sources/TerminalController.swift.
Surface Resolution and Handle Semantics
Every browser command resolves targets through v2ResolveSurfaceId (called via v2ResolveTabManager → v2ResolveWorkspace). The system returns handle refs (window_ref, workspace_ref, pane_ref, surface_ref) that act as persistent identifiers. According to the agent-browser specification in docs/agent-browser-port-spec.md, these UUID-based handles never change during surface moves, ensuring reliable automation scripts.
The actual WebView interaction occurs in Sources/Panels/BrowserPanel.swift, which hosts the native WKWebView and provides methods for navigation, snapshots, and proxy management that the V2 methods consume.
Core Browser Automation Commands
The browser API implements the agent-browser command surface documented in docs/agent-browser-port-spec.md. Each command accepts a --surface parameter and returns JSON containing updated handle refs.
Surface Identification and Navigation
Before automating, obtain the target surface ID using the identity system:
cmux identify --json
This returns the surface_id of the currently focused browser pane. Use this ID to open new splits or navigate:
# Open a URL in a new split (or reuse the right-most pane)
cmux browser --surface surface:b8d2 open_split url=https://developer.apple.com
# Navigate an existing surface to a new URL
cmux browser --surface surface:4f12 navigate url=https://github.com/manaflow-ai/cmux
The v2BrowserOpenSplit function (lines 7369+ in TerminalController.swift) handles placement logic, creating splits when necessary or routing to external browsers based on the "external open" branch logic.
Element Interaction (Click, Type, Fill)
Interact with DOM elements using CSS selectors. The v2BrowserClick implementation (line 8157) injects JavaScript to perform the click event:
# Click an element by CSS selector
cmux browser --surface surface:4f12 click selector="a[href$='README.md']"
# Type text into form fields
cmux browser --surface surface:4f12 type selector="#search" text="cmux automation"
# Fill forms with specific values
cmux browser --surface surface:4f12 fill selector="#username" text="admin"
These commands generate JavaScript snippets executed via BrowserPanel.executeJavaScript, providing synchronous feedback on success or selector-not-found errors.
Waiting and Screenshot Capture
Automation scripts often need to pause for page loads or element appearance. The v2BrowserWait command polls for selector presence:
# Wait up to 5000ms for an element to appear
cmux browser --surface surface:4f12 wait selector="#results" timeout=5000
# Capture the visible viewport as PNG
cmux browser --surface surface:4f12 screenshot output=page.png
The v2BrowserSnapshot function (line 7781) interfaces with BrowserPanel to rasterize the WKWebView content, returning image data through the JSON-RPC response.
JavaScript Execution
For complex automation beyond built-in commands, execute arbitrary JavaScript and retrieve results:
cmux browser --surface surface:4f12 eval script="document.title"
Returns:
{ "result": "cmux – a terminal‑centric macOS window manager" }
This uses v2BrowserEval to pipe scripts through the WebView's JavaScript engine and serialize return values back to the CLI.
Complete Automation Script Example
The following bash script demonstrates a full workflow: identifying the current surface, opening a search query in a split, interacting with forms, and capturing results:
#!/usr/bin/env bash
set -euo pipefail
# Identify the currently focused browser surface
SURF=$(cmux identify --json | jq -r '.focused.surface_id')
if [[ -z "$SURF" ]]; then
echo "No focused browser surface." >&2; exit 1
fi
# Open DuckDuckGo in a new split
cmux browser --surface "$SURF" open_split \
url="https://duckduckgo.com/?q=cmux+automation"
# Get the new surface ID from the focused pane
NEW_SURF=$(cmux identify --json | jq -r '.focused.surface_id')
# Wait for the search box to render
cmux browser --surface "$NEW_SURF" wait \
selector="#search_form_input_homepage" timeout=8000
# Type query and submit
cmux browser --surface "$NEW_SURF" type \
selector="#search_form_input_homepage" text="cmux screenshots"
cmux browser --surface "$NEW_SURF" press key="Enter"
# Wait for results and capture screenshot
cmux browser --surface "$NEW_SURF" wait \
selector=".result" timeout=5000
cmux browser --surface "$NEW_SURF" screenshot output="cmux-results.png"
# Close the browser surface when done
cmux surface close "$NEW_SURF"
This script leverages the stable handle semantics—surface_id remains valid throughout the session even if the pane is repositioned within the workspace.
Key Implementation Files
Understanding the source structure helps when debugging automation failures or extending functionality:
Sources/TerminalController.swift— Contains V2 protocol dispatch and all browser method implementations (v2BrowserOpenSplit,v2BrowserClick,v2BrowserEval,v2BrowserSnapshot, etc.)Sources/Panels/BrowserPanel.swift— UI container hosting theWKWebView, handling navigation delegates, snapshots, and proxy configurationCLI/cmux.swift— CLI entry point that parsescmux browser …arguments and forwards JSON-RPC payloads to the daemondocs/agent-browser-port-spec.md— Formal specification of the command surface, method families, and handle semanticsdocs/v2-api-migration.md— Migration guide mapping legacy v1 commands (e.g.,open_browser) to currentbrowser.*syntaxtests_v2/test_browser_api_comprehensive.py— End-to-end test suite serving as reference implementations for correct JSON-RPC payloads
Summary
- The cmux in-app browser API provides native WKWebView automation through a JSON-RPC interface exposed via the CLI.
- All commands flow through
TerminalController.swiftand operate on stable surface IDs that persist across workspace reorganizations. - Available automation primitives include navigation, element interaction (click/type/fill), explicit waiting, screenshots, and arbitrary JavaScript execution.
- The system supports headless-style automation within a visible or backgrounded macOS terminal multiplexer, making it ideal for LLM-driven workflows that require web data extraction.
Frequently Asked Questions
How do I obtain the surface ID for browser automation?
Run cmux identify --json to retrieve the currently focused surface. The output includes surface_id (full UUID) and short refs (e.g., surface:3). Use either format with the --surface flag in subsequent browser commands. The ID remains stable even if you move the pane to a different workspace or split configuration.
What is the difference between browser.open_split and browser.navigate?
browser.open_split (implemented in v2BrowserOpenSplit) creates a new pane or reuses the right-most pane to open a URL, returning a new surface_id. browser.navigate (via v2BrowserNavigate) changes the URL of an existing browser surface without creating new panes. Use open_split when you need side-by-side browser instances; use navigate for sequential page loads within the same pane.
Can I automate cmux browser interactions from Python or Node.js?
Yes. The CLI interface wraps a JSON-RPC protocol over Unix domain sockets. Any language that can spawn subprocesses or open sockets can send payloads matching the schema in docs/agent-browser-port-spec.md. The Python test suite in tests_v2/test_browser_api_comprehensive.py demonstrates the exact JSON structure expected by TerminalController.swift.
Why does my screenshot command return blank images?
Screenshots require the browser surface to have completed rendering. Ensure you call browser.wait with a relevant CSS selector or sufficient timeout before browser.screenshot to guarantee the WKWebView has finished layout and paint operations. The v2BrowserSnapshot function captures the current composited layer; if the page is still loading, you may capture an incomplete or transparent buffer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →