How camofox-browser Handles Large Pages with Offset-Based Pagination
camofox-browser handles large pages by exposing an offset query parameter on HTTP endpoints and window-slicing data server-side, returning paginated chunks of accessibility snapshots and link lists while preserving navigation references.
The jo-inc/camofox-browser repository provides a browser automation API designed to process oversized web pages without overwhelming client applications. Through offset-based pagination, the server segments both ARIA accessibility snapshots and hyperlink collections, allowing clients to retrieve massive documents incrementally using standard query parameters.
Snapshot Pagination with Window-Slicing
Creating and Caching the Full Snapshot
When a client requests a tab snapshot via GET /tabs/:tabId/snapshot, the server generates the complete ARIA tree YAML once and stores it in tabState.lastSnapshot. This single-pass approach ensures expensive DOM traversal occurs only when the page content changes, not during subsequent pagination requests.
Processing Offset Requests
If a request includes offset > 0 and a cached snapshot exists, the server invokes windowSnapshot(tabState.lastSnapshot, offset) rather than rebuilding the entire tree. This logic is implemented in server.js lines 1860‑1885. The windowSnapshot function, defined in lib/snapshot.js, clamps the requested offset, extracts a budgeted chunk of characters from the full YAML, and appends a tail segment of approximately 5 KB containing pagination navigation links.
Pagination Metadata and Response
The API response contains the sliced snapshot text alongside metadata fields: truncated (boolean), totalChars (integer), hasMore (boolean), and nextOffset (integer or null). The nextOffset value indicates the character position for the subsequent chunk. Clients repeatedly call the endpoint with the returned nextOffset until hasMore becomes false, retrieving the complete page piece-by-piece. See the implementation in lib/snapshot.js lines 10‑38.
Link Pagination with Offset and Limit
Extracting Hyperlinks
The GET /tabs/:tabId/links endpoint gathers all <a> elements from the current page, as implemented in server.js lines 789‑803. The server maintains the full link array in memory while exposing pagination controls to limit the response size.
Slicing the Link Array
The endpoint accepts offset and limit query parameters, defaulting to 0 and 50 respectively. It slices the complete link array using allLinks.slice(offset, offset + limit), returning only the requested window. The response includes a pagination object containing total, offset, limit, and hasMore fields. This implementation appears in server.js lines 2380‑2412.
Practical Code Examples
Fetch the first chunk of a large accessibility snapshot:
GET /tabs/abc123/snapshot?userId=agent1&offset=0 HTTP/1.1
Host: localhost:9377
The response includes pagination metadata:
{
"url": "https://example.com",
"snapshot": "…first‑chunk…\n[... truncated at char 80000 of 215432. Call snapshot with offset=80000 to see more. Pagination links below. ...]\n…tail‑with‑links…",
"truncated": true,
"totalChars": 215432,
"hasMore": true,
"nextOffset": 80000
}
Request the next chunk using the returned nextOffset:
GET /tabs/abc123/snapshot?userId=agent1&offset=80000 HTTP/1.1
Host: localhost:9377
Paginate through a large link list:
GET /tabs/abc123/links?userId=agent1&limit=20&offset=0 HTTP/1.1
Host: localhost:9377
Response format:
{
"links": [
{"url":"https://example.com/page1","text":"Page 1"}
],
"pagination": {
"total": 1234,
"offset": 0,
"limit": 20,
"hasMore": true
}
}
Advance through the dataset by incrementing offset until hasMore becomes false.
Key Implementation Files
| File | Purpose | Lines |
|---|---|---|
lib/snapshot.js |
Implements windowSnapshot for slicing YAML snapshots and adding pagination metadata |
10‑38 |
server.js |
Handles snapshot requests with offset logic | 1860‑1885 |
server.js |
Implements link pagination with offset and limit | 2380‑2412 |
server.js |
Link extraction logic | 789‑803 |
tests/e2e/snapshot-truncation.test.js |
Validates offset-based snapshot pagination | - |
tests/e2e/snapshotLinks.test.js |
Validates link pagination using offset | - |
Summary
- camofox-browser implements offset-based pagination for both accessibility snapshots and link lists to handle large pages efficiently.
- The snapshot endpoint caches the full ARIA tree, then uses
windowSnapshotinlib/snapshot.jsto return budgeted chunks with a 5 KB tail segment containing navigation links. - The links endpoint accepts
offsetandlimitparameters (defaulting to 0 and 50) to slice the full hyperlink array server-side. - Both methods return
hasMoreandnextOffsetmetadata, enabling clients to iterate through massive datasets without memory pressure.
Frequently Asked Questions
How does camofox-browser determine the chunk size for snapshot pagination?
The windowSnapshot function in lib/snapshot.js extracts a budgeted chunk of characters based on the requested offset and appends a fixed tail segment of approximately 5 KB. This tail ensures pagination links remain available in every response chunk, while the main budget prevents individual responses from exceeding manageable sizes.
What are the default values for link pagination in camofox-browser?
According to the source code in server.js lines 2380‑2412, the offset parameter defaults to 0 and the limit parameter defaults to 50. Clients can request larger or smaller windows by explicitly setting these query parameters.
Why does camofox-browser cache the full snapshot instead of regenerating it for each offset?
The server stores the complete ARIA tree YAML in tabState.lastSnapshot to avoid expensive DOM traversal on every pagination request. This cache persists until the page content changes, allowing rapid offset-based slicing via windowSnapshot without rebuilding the accessibility tree.
What happens when I reach the end of a paginated resource?
When the final chunk is returned, the response includes "hasMore": false and "nextOffset": null (for snapshots) or equivalent pagination metadata indicating no further data exists. The client should check these fields to terminate the pagination loop.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →