How to Bisect a CDP Firehose into Per-Page Searchable Buckets

Use the bisect-cdp.mjs script from the browserbase/skills repository to split monolithic Chrome DevTools Protocol (CDP) recordings into hierarchical, per-page buckets that support grep, jq, and SQL-style queries.

The browser-trace skill in the browserbase/skills repository solves the problem of navigating massive CDP event dumps. By using the bisect-cdp.mjs script to bisect the CDP firehose into discrete page contexts, you transform an unreadable stream of network, console, and DOM events into a structured dataset where every request, log, and mutation is tied to a specific navigation boundary.

Understanding the CDP Firehose Input

The recorder outputs every CDP event to <run-id>/cdp/raw.ndjson as newline-delimited JSON. These events span multiple domains—Network, Page, Runtime, Debugger—and include two timestamp formats: a monotonic clock (seconds since browser start) and a wall-clock epoch (milliseconds since 1970). Without bisection, querying this file requires complex temporal joins and frame-aware filtering.

The Bisecting Pipeline

The core logic resides in skills/browser-trace/scripts/bisect-cdp.mjs. It processes the raw NDJSON through four deterministic stages to produce the bucketed output.

Run Identification and Directory Resolution

The script accepts a <run-id> argument and resolves the run directory via runDir() from skills/browser-trace/scripts/lib.mjs. This establishes the base path for reading raw.ndjson and writing the resulting hierarchy.

Timestamp Anchoring

To handle mixed clock formats, the script records the first monotonic timestamp as anchorCdp. The toMs() helper converts all subsequent timestamps to wall-clock milliseconds relative to this anchor. This stabilization ensures chronological accuracy even when the CDP stream alternates between clock types.

Page Detection and PID Assignment

Navigation boundaries are detected using isTopNav() (defined in lib.mjs), which identifies top-level Page.frameNavigated events. A running counter (pid) increments on each top-level navigation, and every event receives a _pid stamp. Events occurring before the first navigation are forced into page 0, guaranteeing a predictable output shape regardless of when recording started.

Session-Wide Buckets and Per-Page Slices

The BUCKETS constant in lib.mjs maps bucket names to predicates testing event.method. The script performs two passes:

  • Session-wide buckets: Filters the full event list against each bucket predicate and writes matching events to <run>/cdp/<bucket>.jsonl. Empty buckets are materialized to preserve legacy layout expectations.
  • Per-page slices: Groups events by _pid. For each page, it writes:
    • url.txt: The first top-level navigation URL or "(initial)" if none exists.
    • raw.jsonl: The page's raw events with the _pid field stripped.
    • Bucket-specific JSONL files (e.g., network/requests.jsonl), written only if skipEmpty is true and events exist.
    • summary.json: A roll-up generated by computePageSummary().

Finally, a top-level summary.json combines session metadata (ID, timestamps, event counts) with the array of per-page summaries.

Bucket Definitions and Filtering

The BUCKETS object in lib.mjs defines categories like network/requests, console/logs, runtime/events, and page/lifecycle. Each bucket uses a predicate function to test the event.method string. This declarative approach makes the output compatible with standard Unix tools—grep, jq, or SQLite loaders—without requiring CDP-specific parsing logic in your query layer.

Running the Bisection

Execute the script with Node.js 14 or higher. The utility uses only standard library modules (fs, path, child_process), ensuring zero runtime dependencies.


# Bisect a specific run into searchable buckets

node ./skills/browser-trace/scripts/bisect-cdp.mjs my-run-123

The resulting directory structure:

my-run-123/
└─ cdp/
   ├─ summary.json                # Session overview with metadata

   ├─ network/requests.jsonl      # Session-wide network traffic

   └─ pages/
      ├─ 000/
      │  ├─ url.txt              # Navigation URL

      │  ├─ raw.jsonl            # Raw events for this page

      │  ├─ network/requests.jsonl
      │  ├─ console/logs.jsonl
      │  └─ summary.json         # Page-specific metrics

      └─ 001/
         └─ ...

Querying the Bucketed Output

The JSONL format enables stream-processing without loading entire sessions into memory.

List all network requests for page 1:

jq -r '.params.request.url' my-run-123/cdp/pages/001/network/requests.jsonl

Aggregate console errors across the entire session:

grep '"method":"Runtime.consoleAPICalled"' my-run-123/cdp/pages/*/console/logs.jsonl | \
jq -r 'select(.params.type=="error") | .params.args[].value'

Summary

  • The browser-trace skill in browserbase/skills provides a deterministic bisection pipeline via bisect-cdp.mjs.
  • It relies on top-level navigations (isTopNav()) to partition events into page IDs (_pid), avoiding iframe noise.
  • Timestamp anchoring (anchorCdp, toMs()) harmonizes mixed monotonic and wall-clock timestamps.
  • Bucket definitions (BUCKETS in lib.mjs) separate concerns like network, console, and runtime into queryable JSONL files.
  • The output structure supports zero-dependency querying with standard Unix tools and requires only Node.js ≥ 14.

Frequently Asked Questions

What CDP event triggers a new page bucket?

A new page bucket is created when isTopNav() detects a top-level Page.frameNavigated event. The script increments an internal pid counter and stamps all subsequent events with that _pid until the next top-level navigation occurs. Events before the first navigation are assigned to page 0.

Why are events stamped with _pid instead of using the CDP session ID?

The _pid (page ID) provides a deterministic, sequential index based on navigation order rather than the CDP session or frame ID. This prevents iframe navigations from fragmenting the output and ensures that all activity during a single browser tab’s lifecycle is grouped together until a genuine top-level navigation resets the context.

Can I customize which events go into specific buckets?

Yes. Modify the BUCKETS object exported from skills/browser-trace/scripts/lib.mjs. Each bucket is defined by a name and a predicate function that receives the event object. Return true to include the event in that bucket.

Does the script handle large NDJSON files efficiently?

The current implementation reads the entire raw.ndjson into memory for processing. For extremely large traces (several gigabytes), ensure your Node.js runtime has access to sufficient heap memory, or consider streaming modifications to the bisect-cdp.mjs logic. The output format uses JSONL specifically to facilitate memory-efficient downstream querying regardless of input size.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →