# How to Bisect a CDP Firehose into Per-Page Searchable Buckets

> Bisect CDP firehose recordings into per-page searchable buckets using the browserbase/skills script. Enable powerful grep, jq, and SQL queries for your CDP data.

- Repository: [browserbase/skills](https://github.com/browserbase/skills)
- Tags: how-to-guide
- Published: 2026-05-01

---

**Use the `bisect-cdp.mjs` script from the browserbase/skills repository to split monolithic Chrome DevTools Protocol (CDP) recordings into hierarchical, per-page buckets that support grep, jq, and SQL-style queries.**

The **browser-trace** skill in the [browserbase/skills](https://github.com/browserbase/skills) repository solves the problem of navigating massive CDP event dumps. By using the `bisect-cdp.mjs` script to bisect the CDP firehose into discrete page contexts, you transform an unreadable stream of network, console, and DOM events into a structured dataset where every request, log, and mutation is tied to a specific navigation boundary.

## Understanding the CDP Firehose Input

The recorder outputs every CDP event to `<run-id>/cdp/raw.ndjson` as newline-delimited JSON. These events span multiple domains—Network, Page, Runtime, Debugger—and include two timestamp formats: a monotonic clock (seconds since browser start) and a wall-clock epoch (milliseconds since 1970). Without bisection, querying this file requires complex temporal joins and frame-aware filtering.

## The Bisecting Pipeline

The core logic resides in `skills/browser-trace/scripts/bisect-cdp.mjs`. It processes the raw NDJSON through four deterministic stages to produce the bucketed output.

### Run Identification and Directory Resolution

The script accepts a `<run-id>` argument and resolves the run directory via `runDir()` from `skills/browser-trace/scripts/lib.mjs`. This establishes the base path for reading `raw.ndjson` and writing the resulting hierarchy.

### Timestamp Anchoring

To handle mixed clock formats, the script records the first monotonic timestamp as `anchorCdp`. The `toMs()` helper converts all subsequent timestamps to wall-clock milliseconds relative to this anchor. This stabilization ensures chronological accuracy even when the CDP stream alternates between clock types.

### Page Detection and PID Assignment

Navigation boundaries are detected using `isTopNav()` (defined in *lib.mjs*), which identifies top-level `Page.frameNavigated` events. A running counter (`pid`) increments on each top-level navigation, and every event receives a `_pid` stamp. Events occurring before the first navigation are forced into page `0`, guaranteeing a predictable output shape regardless of when recording started.

### Session-Wide Buckets and Per-Page Slices

The `BUCKETS` constant in *lib.mjs* maps bucket names to predicates testing `event.method`. The script performs two passes:

- **Session-wide buckets**: Filters the full event list against each bucket predicate and writes matching events to `<run>/cdp/<bucket>.jsonl`. Empty buckets are materialized to preserve legacy layout expectations.
- **Per-page slices**: Groups events by `_pid`. For each page, it writes:
  - [`url.txt`](https://github.com/browserbase/skills/blob/main/url.txt): The first top-level navigation URL or "(initial)" if none exists.
  - `raw.jsonl`: The page's raw events with the `_pid` field stripped.
  - Bucket-specific JSONL files (e.g., `network/requests.jsonl`), written only if `skipEmpty` is true and events exist.
  - [`summary.json`](https://github.com/browserbase/skills/blob/main/summary.json): A roll-up generated by `computePageSummary()`.

Finally, a top-level [`summary.json`](https://github.com/browserbase/skills/blob/main/summary.json) combines session metadata (ID, timestamps, event counts) with the array of per-page summaries.

## Bucket Definitions and Filtering

The `BUCKETS` object in `lib.mjs` defines categories like `network/requests`, `console/logs`, `runtime/events`, and `page/lifecycle`. Each bucket uses a predicate function to test the `event.method` string. This declarative approach makes the output compatible with standard Unix tools—`grep`, `jq`, or SQLite loaders—without requiring CDP-specific parsing logic in your query layer.

## Running the Bisection

Execute the script with Node.js 14 or higher. The utility uses only standard library modules (`fs`, `path`, `child_process`), ensuring zero runtime dependencies.

```bash

# Bisect a specific run into searchable buckets

node ./skills/browser-trace/scripts/bisect-cdp.mjs my-run-123

```

The resulting directory structure:

```text
my-run-123/
└─ cdp/
   ├─ summary.json                # Session overview with metadata

   ├─ network/requests.jsonl      # Session-wide network traffic

   └─ pages/
      ├─ 000/
      │  ├─ url.txt              # Navigation URL

      │  ├─ raw.jsonl            # Raw events for this page

      │  ├─ network/requests.jsonl
      │  ├─ console/logs.jsonl
      │  └─ summary.json         # Page-specific metrics

      └─ 001/
         └─ ...

```

## Querying the Bucketed Output

The JSONL format enables stream-processing without loading entire sessions into memory.

List all network requests for page 1:

```bash
jq -r '.params.request.url' my-run-123/cdp/pages/001/network/requests.jsonl

```

Aggregate console errors across the entire session:

```bash
grep '"method":"Runtime.consoleAPICalled"' my-run-123/cdp/pages/*/console/logs.jsonl | \
jq -r 'select(.params.type=="error") | .params.args[].value'

```

## Summary

- The **browser-trace** skill in `browserbase/skills` provides a **deterministic bisection** pipeline via `bisect-cdp.mjs`.
- It relies on **top-level navigations** (`isTopNav()`) to partition events into page IDs (`_pid`), avoiding iframe noise.
- **Timestamp anchoring** (`anchorCdp`, `toMs()`) harmonizes mixed monotonic and wall-clock timestamps.
- **Bucket definitions** (`BUCKETS` in *lib.mjs*) separate concerns like network, console, and runtime into queryable JSONL files.
- The output structure supports **zero-dependency querying** with standard Unix tools and requires only Node.js ≥ 14.

## Frequently Asked Questions

### What CDP event triggers a new page bucket?

A new page bucket is created when `isTopNav()` detects a top-level `Page.frameNavigated` event. The script increments an internal `pid` counter and stamps all subsequent events with that `_pid` until the next top-level navigation occurs. Events before the first navigation are assigned to page `0`.

### Why are events stamped with `_pid` instead of using the CDP session ID?

The `_pid` (page ID) provides a **deterministic, sequential index** based on navigation order rather than the CDP session or frame ID. This prevents iframe navigations from fragmenting the output and ensures that all activity during a single browser tab’s lifecycle is grouped together until a genuine top-level navigation resets the context.

### Can I customize which events go into specific buckets?

Yes. Modify the `BUCKETS` object exported from `skills/browser-trace/scripts/lib.mjs`. Each bucket is defined by a name and a predicate function that receives the event object. Return `true` to include the event in that bucket.

### Does the script handle large NDJSON files efficiently?

The current implementation reads the entire `raw.ndjson` into memory for processing. For extremely large traces (several gigabytes), ensure your Node.js runtime has access to sufficient heap memory, or consider streaming modifications to the `bisect-cdp.mjs` logic. The output format uses JSONL specifically to facilitate memory-efficient downstream querying regardless of input size.