How pi-computer-use Integrates Browser Pages via Chrome DevTools Protocol (CDP) as UI Roots
pi-computer-use treats Chromium browser tabs as first-class UI roots by connecting through the Chrome DevTools Protocol (CDP) when the PI_COMPUTER_USE_CDP_PORT environment variable is set, enabling automated interaction without platform-specific accessibility APIs.
The pi-computer-use framework extends its UI automation capabilities beyond native macOS and Windows applications to include Chromium-family browsers. By leveraging the Chrome DevTools Protocol (CDP), the system treats browser pages as interactive roots within the same architectural model used for platform accessibility trees, allowing seamless automation of web applications alongside native desktop software.
Enabling CDP Mode
CDP integration is optional and activates when the environment variable PI_COMPUTER_USE_CDP_PORT specifies a valid port number. The function cdpEnabled() in src/cdp.ts validates this configuration and checks for WebSocket API availability. If these conditions are not met, the CDP module remains inert and the system falls back to native accessibility backends.
To enable CDP support, launch Chromium with the --remote-debugging-port flag and set the environment variable:
# Launch Chrome with remote debugging
/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222
# In your pi-computer-use environment
export PI_COMPUTER_USE_CDP_PORT=9222
Discovering Browser Pages
Once enabled, the system discovers available tabs through the standard CDP HTTP endpoint. The cdpPages() function queries http://127.0.0.1:<port>/json/list to retrieve all debuggable targets, filtering for local pages that expose a webSocketDebuggerUrl.
The listCdpPageContexts() function then transforms these targets into CdpPageContext objects, assigning each a contextId in the format browser:<targetId>. This identifier allows the rest of the system to reference specific browser tabs using the same addressing scheme as native UI elements.
Connecting to Target Tabs
To establish control over a specific browser window, cdpTabForWindow(windowTitle, frame) manages tab selection through the connectedTabs cache. The selection logic follows a specific disambiguation order implemented in pickTab():
- Exact title match – Tab title equals the target window title
- Prefix match – Tab title starts with the window title
- Frame matching – Optional screen coordinate intersection checks
- Visibility heuristics – Selects the visible tab (
document.visibilityState === "visible")
Once identified, CdpTab.connect(wsUrl, targetId, title) opens a WebSocket connection to the tab's debugger URL and enables the Runtime and Page CDP domains. This connection persists for the duration of the automation session, buffering console messages and exceptions for later retrieval.
Capturing Page State and Accessibility
When the system requests a UI snapshot via cdpSnapshotForContext(contextId), the CdpTab instance executes a multi-step capture process:
- Text extraction – Evaluates
document.body.innerTextto capture readable content - Accessibility tree retrieval – Calls
Accessibility.getFullAXTreeto obtain the complete browser accessibility structure - Target generation – Processes AX nodes through
cdpSnapshotOutline()to createCdpSnapshotTargetobjects for actionable elements like buttons, links, and text inputs
The resulting snapshot merges with the system's internal model, exposing browser elements through the same interface as native macOS Accessibility or Windows UI Automation nodes.
Handling Browser Console and Errors
The CdpTab.handleMessage() method listens for Runtime.consoleAPICalled and Runtime.exceptionThrown events, maintaining a circular buffer of recent messages limited by CONSOLE_BUFFER_LIMIT. This buffer provides diagnostic parity with native backends through the drainConsole() method, which flushes accumulated logs and errors for inclusion in tool results.
Executing Interactions via CDP
Interaction functions translate high-level commands into CDP Input domain methods. The system implements specific handlers such as cdpClickForContext, cdpTypeForContext, cdpScrollForContext, and cdpKeypressForContext, which invoke:
Input.dispatchMouseEventfor clicks and scrollsInput.insertTextandInput.dispatchKeyEventfor typingRuntime.evaluatefor JavaScript execution
These functions target specific backend node IDs or perform actions on the currently focused element, maintaining consistency with the native automation API surface.
Unified Root Architecture
The integration bridges CDP-derived browser trees with native accessibility backends through src/runtime.ts, which orchestrates platform-specific handlers. When CDP is enabled, browser pages appear alongside macOS and Windows UI elements as additional roots. The src/bridge.ts module routes actions to CDP functions when encountering contexts with source equal to "browser_ax", while src/platform/macos/browser.ts handles macOS-specific delegation logic for Chrome and Edge instances.
Code Examples
Enable CDP and connect to a specific browser tab:
// Enable CDP by setting the env var to the browser's remote debugging port
process.env.PI_COMPUTER_USE_CDP_PORT = "9222";
// Get a CDP tab that matches the active window title "My Web App"
const tab = await cdpTabForWindow("My Web App");
if (tab) {
// Navigate to a URL
await tab.navigate("https://example.com");
// Click a button identified by its backend node id
await tab.clickBackendNode(12345);
// Type into a focused element
await tab.typeIntoFocused("Hello world");
// Retrieve recent console messages
const logs = tab.drainConsole();
console.log(logs);
}
Capture a snapshot of the current page context:
// Obtain a snapshot of the current page as a pi-computer-use context
const ctxId = "browser:abcdef1234"; // contextId of a previously discovered tab
const snapshot = await cdpSnapshotForContext(ctxId);
if (snapshot) {
console.log("Page title:", snapshot.title);
console.log("Text content length:", snapshot.text.length);
console.log("Interactive targets:", snapshot.targets.length);
}
Summary
- Optional activation – CDP integration requires setting
PI_COMPUTER_USE_CDP_PORTand launching Chromium with--remote-debugging-port - Page discovery –
src/cdp.tsquerieshttp://127.0.0.1:<port>/json/listand maps targets tobrowser:<targetId>contexts - Smart tab selection –
pickTab()uses title matching, frame intersection, and visibility heuristics to identify the correct tab - Unified model – Browser accessibility trees merge with native OS trees through
CdpSnapshotTargetobjects and"browser_ax"source tagging - Full interaction suite – Click, type, scroll, and keypress actions map to CDP
Inputdomain commands - Diagnostic parity – Console and exception buffering via
drainConsole()provides equivalent debugging to native backends
Frequently Asked Questions
How does pi-computer-use identify the correct browser tab among multiple open pages?
The system uses cdpTabForWindow() to match the target window title against available tabs, falling back to pickTab() which applies a priority order: exact title matches, prefix matches, optional frame-based coordinate intersection, and finally visibility state (document.visibilityState === "visible"). This ensures reliable selection even when multiple similar pages are open.
What CDP domains does pi-computer-use require to function?
The implementation utilizes several Chrome DevTools Protocol domains: Runtime for JavaScript evaluation and console monitoring, Page for navigation events, Accessibility for retrieving the AX tree via getFullAXTree, and Input for dispatching mouse events, text insertion, and keyboard actions. The CdpTab.connect() method automatically enables the necessary domains upon connection.
Can pi-computer-use access browser console output when using CDP mode?
Yes. The CdpTab class maintains an internal buffer of console messages and exceptions captured through Runtime.consoleAPICalled and Runtime.exceptionThrown events. Calling drainConsole() retrieves and clears this buffer, attaching diagnostic output to tool results in the same manner as native platform backends.
How does the CDP integration differ from native macOS or Windows browser automation?
Unlike platform-specific accessibility APIs that rely on OS-level hooks, the CDP approach communicates directly with the browser via WebSocket, eliminating dependencies on AppleScript or Windows UI Automation for Chromium-based browsers. This allows pi-computer-use to treat browser pages as roots within the same abstract model used for native applications, as coordinated through src/bridge.ts and src/runtime.ts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →