How to Debug Browser Automation Failures Using CDP Firehose Tracing: A Complete Guide
You can debug intermittent browser automation failures by attaching a secondary read-only CDP client that captures a full firehose of Chrome DevTools Protocol events, which the browser-trace skill then bisects into searchable per-page and per-domain buckets.
Intermittent failures in Playwright, Selenium, or browse CLI scripts often stem from hidden timing or network issues that disappear when you add traditional breakpoints. The browserbase/skills repository provides a browser-trace skill that solves this by recording a complete CDP firehose during automation execution, giving you forensic-level visibility without interfering with your primary test flow.
Why CDP Firehose Tracing Is Essential for Debugging
Traditional debugging methods break timing-sensitive bugs. The browser-trace skill implements a dual CDP client pattern where your primary automation client drives the page while a second read-only client, implemented in skills/browser-trace/scripts/start-capture.mjs, observes only. This secondary client enables observation domains including Network, Console, Runtime, Log, Page, and optionally DOM, but never sends action commands, ensuring zero interference with your automation timing.
Architecture of the Browser-Trace Skill
The Dual CDP Client Pattern
The architecture separates execution from observation. Your main automation script connects to Chrome DevTools Protocol (CDP) to perform actions—clicking, typing, navigating—while the tracer opens a parallel websocket connection to the same target. According to the source code in skills/browser-trace/SKILL.md, this design prevents the tracer from altering page state or consuming events that your automation needs.
The Three-Component Pipeline
The skill processes telemetry through three distinct stages:
- Firehose: The
browse cdp <target>command streams every CDP event as newline-delimited JSON (NDJSON) to.o11y/<run-id>/cdp/raw.ndjson. This captures the raw, unfiltered protocol traffic. - Sampler: A polling loop in
scripts/start-capture.mjscaptures visual state every 2 seconds (default) by callingbrowse --ws <target> screenshotandbrowse --ws <target> get html body, writing PNGs toscreenshots/and HTML snapshots todom/. - Bisector: After the run completes,
scripts/bisect-cdp.mjswalksraw.ndjsononce to create organized buckets—per-domain files likecdp/network/requests.jsonland per-page slices undercdp/pages/<pid>/that group events between successivePage.frameNavigatedevents. The bisector is idempotent; rerunning it simply overwrites the bucket files.
Filesystem Layout and Key Artifacts
All artifacts live under .o11y/<run-id>/. The most important entry points for debugging are:
cdp/summary.json– High-level session totals and an array of page summaries.cdp/pages/<pid>/summary.json– Per-page breakdown including network requests, console logs, and navigation timings.screenshots/anddom/– Visual snapshots that reveal whether the DOM actually changed during hangs or failures.
Setup and Usage Patterns
Local Chrome Debugging
For local Chrome instances, launch the browser with --remote-debugging-port and execute the capture workflow:
# 1. Launch Chrome with debugging port
"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
--remote-debugging-port=9222 \
--user-data-dir=/tmp/chrome-o11y \
about:blank &
# 2. Start the tracer (creates run directory .o11y/my-run)
node skills/browser-trace/scripts/start-capture.mjs 9222 my-run
# 3. Run your automation against the same port
browse env local 9222
browse open https://example.com
# ... your automation script ...
# 4. Stop and process the trace
node skills/browser-trace/scripts/stop-capture.mjs my-run
node skills/browser-trace/scripts/bisect-cdp.mjs my-run
Remote Browserbase Debugging
For remote Browserbase sessions, use the keep-alive wrapper scripts to prevent the session from terminating when your automation disconnects:
export BROWSERBASE_API_KEY=YOUR_KEY
# 1. Create keep-alive session and start tracing
node skills/browser-trace/scripts/bb-capture.mjs --new my-run
# 2. Attach your automation (session ID stored in manifest)
SID=$(jq -r .browserbase.session_id .o11y/my-run/manifest.json)
browse --connect "$SID" open https://example.com
# 3. Finalize and optionally release the session
node skills/browser-trace/scripts/stop-capture.mjs my-run
node skills/browser-trace/scripts/bisect-cdp.mjs my-run
node skills/browser-trace/scripts/bb-finalize.mjs my-run --release
Critical Note: You must start the tracer before the automation client connects, otherwise the Browserbase session may close as soon as the only CDP client disconnects.
Querying the Trace Data
After bisection, use scripts/query.mjs to explore the structured data without memorizing paths:
# List all captured pages
node skills/browser-trace/scripts/query.mjs my-run list
# Show summary for page ID 2
node skills/browser-trace/scripts/query.mjs my-run page 2
# Find failed network requests on specific page
node skills/browser-trace/scripts/query.mjs my-run page 2 network/failed
# Get all console errors across the run
node skills/browser-trace/scripts/query.mjs my-run errors
For ad-hoc analysis, target the JSON Lines files directly with jq:
# All 4xx/5xx responses
jq -c 'select(.params.response.status >= 400) |
{status: .params.response.status, url: .params.response.url}' \
.o11y/my-run/cdp/network/responses.jsonl
# Console errors only
jq -c 'select(.params.type == "error")' \
.o11y/my-run/cdp/console/logs.jsonl
Diagnostic Workflows for Common Failures
Intermittent form submission failures: Capture the firehose to correlate the exact timestamp of your submit action with network requests and console errors. Check cdp/pages/<pid>/network/requests.jsonl for 4xx/5xx status codes that indicate server-side rejection or CORS issues.
Page hangs after navigation: Compare the sampler's screenshots in screenshots/ and DOM dumps in dom/ to determine if the browser rendered the expected content or stalled on a loading spinner. The Page.frameNavigated events in the per-page buckets reveal navigation lifecycle timing.
Per-page performance regression: The bisected page/ buckets contain navigation timestamps and lifecycle events (domContentEventFired, loadEventFired) that let you pinpoint exactly which page in a multi-page flow introduced latency.
Summary
- Dual CDP architecture isolates the tracer from your automation logic using read-only observation domains in a second client.
- Three-stage pipeline (Firehose, Sampler, Bisector) turns raw NDJSON protocol traffic into organized, queryable artifacts under
.o11y/<run-id>/. - Idempotent bisection via
scripts/bisect-cdp.mjscreates per-domain and per-page JSON Lines files safe for repeated analysis. - Flexible deployment supports both local Chrome (
--remote-debugging-port) and remote Browserbase sessions with keep-alive management viabb-capture.mjs. - Query interfaces including
scripts/query.mjsand directjqfiltering enable rapid post-mortem debugging of network failures, console errors, and timing issues.
Frequently Asked Questions
Will the tracer interfere with my automation timing?
No. The tracer uses a second CDP client that only enables observation domains (Network, Console, Runtime, Log, Page, optionally DOM) and never sends action commands. According to the implementation in skills/browser-trace/SKILL.md, this read-only design ensures the tracer does not consume events or alter state that your primary automation client depends on.
How do I find which page failed in a multi-page session?
Use node skills/browser-trace/scripts/query.mjs <run-id> list to enumerate all captured pages with their IDs. Then drill into a specific page with query.mjs <run-id> page <pid> to view navigation timestamps, lifecycle events, and aggregated metrics that identify where the flow broke.
Can I run this against an existing Browserbase session?
Yes. Use node skills/browser-trace/scripts/bb-capture.mjs with the existing session ID to attach the tracer to a running session. However, you must attach the tracer before your automation client connects, because Browserbase sessions terminate when the last CDP client disconnects. The bb-capture.mjs script creates a keep-alive mechanism to prevent premature closure.
What is the storage and performance impact of the firehose?
The firehose streams events to disk as NDJSON in real-time, which has minimal memory impact but writes to .o11y/<run-id>/cdp/raw.ndjson continuously. The optional sampler adds overhead by capturing screenshots and DOM dumps every 2 seconds (configurable). For long-running sessions, ensure adequate disk space for the raw trace file, which can grow to several hundred megabytes for complex single-page applications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →