# How to Debug Browser Automation Failures Using CDP Firehose Tracing: A Complete Guide

> Debug browser automation failures with CDP firehose tracing. Capture full DevTools Protocol events and bisect them into searchable per-page or per-domain buckets for faster issue resolution.

- Repository: [browserbase/skills](https://github.com/browserbase/skills)
- Tags: how-to-guide
- Published: 2026-05-01

---

**You can debug intermittent browser automation failures by attaching a secondary read-only CDP client that captures a full firehose of Chrome DevTools Protocol events, which the browser-trace skill then bisects into searchable per-page and per-domain buckets.**

Intermittent failures in Playwright, Selenium, or `browse` CLI scripts often stem from hidden timing or network issues that disappear when you add traditional breakpoints. The **browserbase/skills** repository provides a **browser-trace** skill that solves this by recording a complete CDP firehose during automation execution, giving you forensic-level visibility without interfering with your primary test flow.

## Why CDP Firehose Tracing Is Essential for Debugging

Traditional debugging methods break timing-sensitive bugs. The **browser-trace** skill implements a **dual CDP client** pattern where your primary automation client drives the page while a second read-only client, implemented in `skills/browser-trace/scripts/start-capture.mjs`, observes only. This secondary client enables observation domains including `Network`, `Console`, `Runtime`, `Log`, `Page`, and optionally `DOM`, but never sends action commands, ensuring zero interference with your automation timing.

## Architecture of the Browser-Trace Skill

### The Dual CDP Client Pattern

The architecture separates execution from observation. Your main automation script connects to Chrome DevTools Protocol (CDP) to perform actions—clicking, typing, navigating—while the tracer opens a parallel websocket connection to the same target. According to the source code in [`skills/browser-trace/SKILL.md`](https://github.com/browserbase/skills/blob/main/skills/browser-trace/SKILL.md), this design prevents the tracer from altering page state or consuming events that your automation needs.

### The Three-Component Pipeline

The skill processes telemetry through three distinct stages:

- **Firehose**: The `browse cdp <target>` command streams every CDP event as newline-delimited JSON (NDJSON) to `.o11y/<run-id>/cdp/raw.ndjson`. This captures the raw, unfiltered protocol traffic.
- **Sampler**: A polling loop in `scripts/start-capture.mjs` captures visual state every 2 seconds (default) by calling `browse --ws <target> screenshot` and `browse --ws <target> get html body`, writing PNGs to `screenshots/` and HTML snapshots to `dom/`.
- **Bisector**: After the run completes, `scripts/bisect-cdp.mjs` walks `raw.ndjson` once to create organized buckets—per-domain files like `cdp/network/requests.jsonl` and per-page slices under `cdp/pages/<pid>/` that group events between successive `Page.frameNavigated` events. The bisector is idempotent; rerunning it simply overwrites the bucket files.

### Filesystem Layout and Key Artifacts

All artifacts live under `.o11y/<run-id>/`. The most important entry points for debugging are:

- [`cdp/summary.json`](https://github.com/browserbase/skills/blob/main/cdp/summary.json) – High-level session totals and an array of page summaries.
- `cdp/pages/<pid>/summary.json` – Per-page breakdown including network requests, console logs, and navigation timings.
- `screenshots/` and `dom/` – Visual snapshots that reveal whether the DOM actually changed during hangs or failures.

## Setup and Usage Patterns

### Local Chrome Debugging

For local Chrome instances, launch the browser with `--remote-debugging-port` and execute the capture workflow:

```bash

# 1. Launch Chrome with debugging port

"/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
  --remote-debugging-port=9222 \
  --user-data-dir=/tmp/chrome-o11y \
  about:blank &

# 2. Start the tracer (creates run directory .o11y/my-run)

node skills/browser-trace/scripts/start-capture.mjs 9222 my-run

# 3. Run your automation against the same port

browse env local 9222
browse open https://example.com

# ... your automation script ...

# 4. Stop and process the trace

node skills/browser-trace/scripts/stop-capture.mjs my-run
node skills/browser-trace/scripts/bisect-cdp.mjs my-run

```

### Remote Browserbase Debugging

For remote Browserbase sessions, use the keep-alive wrapper scripts to prevent the session from terminating when your automation disconnects:

```bash
export BROWSERBASE_API_KEY=YOUR_KEY

# 1. Create keep-alive session and start tracing

node skills/browser-trace/scripts/bb-capture.mjs --new my-run

# 2. Attach your automation (session ID stored in manifest)

SID=$(jq -r .browserbase.session_id .o11y/my-run/manifest.json)
browse --connect "$SID" open https://example.com

# 3. Finalize and optionally release the session

node skills/browser-trace/scripts/stop-capture.mjs my-run
node skills/browser-trace/scripts/bisect-cdp.mjs my-run
node skills/browser-trace/scripts/bb-finalize.mjs my-run --release

```

**Critical Note**: You must start the tracer **before** the automation client connects, otherwise the Browserbase session may close as soon as the only CDP client disconnects.

### Querying the Trace Data

After bisection, use `scripts/query.mjs` to explore the structured data without memorizing paths:

```bash

# List all captured pages

node skills/browser-trace/scripts/query.mjs my-run list

# Show summary for page ID 2

node skills/browser-trace/scripts/query.mjs my-run page 2

# Find failed network requests on specific page

node skills/browser-trace/scripts/query.mjs my-run page 2 network/failed

# Get all console errors across the run

node skills/browser-trace/scripts/query.mjs my-run errors

```

For ad-hoc analysis, target the JSON Lines files directly with `jq`:

```bash

# All 4xx/5xx responses

jq -c 'select(.params.response.status >= 400) |
      {status: .params.response.status, url: .params.response.url}' \
    .o11y/my-run/cdp/network/responses.jsonl

# Console errors only

jq -c 'select(.params.type == "error")' \
    .o11y/my-run/cdp/console/logs.jsonl

```

## Diagnostic Workflows for Common Failures

**Intermittent form submission failures**: Capture the firehose to correlate the exact timestamp of your submit action with network requests and console errors. Check `cdp/pages/<pid>/network/requests.jsonl` for 4xx/5xx status codes that indicate server-side rejection or CORS issues.

**Page hangs after navigation**: Compare the sampler's screenshots in `screenshots/` and DOM dumps in `dom/` to determine if the browser rendered the expected content or stalled on a loading spinner. The `Page.frameNavigated` events in the per-page buckets reveal navigation lifecycle timing.

**Per-page performance regression**: The bisected `page/` buckets contain navigation timestamps and lifecycle events (`domContentEventFired`, `loadEventFired`) that let you pinpoint exactly which page in a multi-page flow introduced latency.

## Summary

- **Dual CDP architecture** isolates the tracer from your automation logic using read-only observation domains in a second client.
- **Three-stage pipeline** (Firehose, Sampler, Bisector) turns raw NDJSON protocol traffic into organized, queryable artifacts under `.o11y/<run-id>/`.
- **Idempotent bisection** via `scripts/bisect-cdp.mjs` creates per-domain and per-page JSON Lines files safe for repeated analysis.
- **Flexible deployment** supports both local Chrome (`--remote-debugging-port`) and remote Browserbase sessions with keep-alive management via `bb-capture.mjs`.
- **Query interfaces** including `scripts/query.mjs` and direct `jq` filtering enable rapid post-mortem debugging of network failures, console errors, and timing issues.

## Frequently Asked Questions

### Will the tracer interfere with my automation timing?

No. The tracer uses a second CDP client that only enables observation domains (`Network`, `Console`, `Runtime`, `Log`, `Page`, optionally `DOM`) and never sends action commands. According to the implementation in [`skills/browser-trace/SKILL.md`](https://github.com/browserbase/skills/blob/main/skills/browser-trace/SKILL.md), this read-only design ensures the tracer does not consume events or alter state that your primary automation client depends on.

### How do I find which page failed in a multi-page session?

Use `node skills/browser-trace/scripts/query.mjs <run-id> list` to enumerate all captured pages with their IDs. Then drill into a specific page with `query.mjs <run-id> page <pid>` to view navigation timestamps, lifecycle events, and aggregated metrics that identify where the flow broke.

### Can I run this against an existing Browserbase session?

Yes. Use `node skills/browser-trace/scripts/bb-capture.mjs` with the existing session ID to attach the tracer to a running session. However, you must attach the tracer before your automation client connects, because Browserbase sessions terminate when the last CDP client disconnects. The `bb-capture.mjs` script creates a keep-alive mechanism to prevent premature closure.

### What is the storage and performance impact of the firehose?

The firehose streams events to disk as NDJSON in real-time, which has minimal memory impact but writes to `.o11y/<run-id>/cdp/raw.ndjson` continuously. The optional sampler adds overhead by capturing screenshots and DOM dumps every 2 seconds (configurable). For long-running sessions, ensure adequate disk space for the raw trace file, which can grow to several hundred megabytes for complex single-page applications.