# How camofox-browser Extracts DOM Images as Inline Data URLs

> Learn how camofox-browser extracts DOM images as inline data URLs. Discover its JavaScript injection, base64 conversion, and REST API functionality for efficient image capture.

- Repository: [jo/camofox-browser](https://github.com/jo-inc/camofox-browser)
- Tags: how-to-guide
- Published: 2026-04-15

---

**Camofox-browser extracts visible DOM images by injecting JavaScript into the Playwright page context, converting them to base64 data URLs when they fall under a configurable byte limit, and returns the results via a REST API endpoint.**

The **camofox-browser** open-source project provides a specialized mechanism for **DOM image extraction with inline data URLs** that operates entirely within the browser context. This architecture avoids server-side network calls that could trigger security scanner false-positives while delivering base64-encoded image data through a clean REST interface.

## The Two-Step Extraction Pipeline

The image extraction feature is implemented as a coordinated pipeline between the Node.js server and the Playwright-controlled page.

### Step 1: In-Page Collection with page.evaluate

The core logic resides in [`lib/images.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/images.js) within the exported `extractPageImages` function. When invoked, this utility injects code into the Playwright page context using `page.evaluate`. The in-page script queries the DOM with `document.querySelectorAll('img')` to locate all visible image elements.

It builds a candidate list (defaulting to ≤ 8 entries, hard-capped at 20) containing unique sources, alt text, and natural dimensions. For each candidate, the script creates a result entry capturing `src`, `alt`, `width`, and `height` properties directly from the HTMLImageElement.

### Step 2: Optional Data-URL Inclusion

When the caller sets `includeData=true`, the in-page script attempts to embed raw image data directly in the response. The handling differs by source type:

- **Data-URL sources** — If `src` already begins with `data:`, the MIME type is extracted and the payload size is estimated. The full data URL is returned only if the estimated byte count remains ≤ `maxBytes`.
- **Network sources** — The script performs a `fetch(src, { credentials: 'include' })` request. Upon success, it captures the response blob's MIME type and byte length. If the blob size satisfies the limit, a `FileReader` helper function `toDataUrl` converts the binary data to a base64-encoded data URL.

Errors—including fetch failures, size limit violations, or `FileReader` issues—are recorded as `fetchError` or `dataSkipped` flags in the result entry rather than throwing exceptions.

## Configuration Parameters and Limits

The extraction behavior is controlled through three primary parameters passed to the `GET /tabs/:tabId/images` endpoint:

- **includeData** — When set to `true`, the pipeline attempts to inline image data as base64 data URLs.
- **maxBytes** — Defines the ceiling for inlining (defaults to `MAX_DOWNLOAD_INLINE_BYTES`, defined in [`lib/downloads.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/downloads.js) at lines 13-15 as **20 MiB**).
- **limit** — Controls the maximum number of distinct images returned (capped at 20, default 8).

## REST API Endpoint and Usage

The server-side implementation in [`server.js`](https://github.com/jo-inc/camofox-browser/blob/main/server.js) (lines 47-65) exposes the extraction functionality via the `GET /tabs/:tabId/images` endpoint. The Node process coordinates the operation, enforces limits, and formats the final JSON payload, while all heavy I/O occurs within the browser context.

Example HTTP request using curl:

```bash
curl "http://localhost:9377/tabs/abc123/images?userId=agent1&includeData=true&maxBytes=524288&limit=5"

```

Example using a JavaScript SDK wrapper:

```javascript
const res = await client.get(
  `/tabs/${tabId}/images`,
  {
    params: {
      userId: 'agent1',
      includeData: true,
      maxBytes: 1 * 1024 * 1024, // 1 MiB
      limit: 10
    }
  }
);

console.log('Extracted images:', res.data.images);

```

Example direct call to the internal helper for custom tooling:

```javascript
import { extractPageImages } from './lib/images.js';
import { createPage } from './lib/launcher.js';

(async () => {
  const page = await createPage('https://example.com');
  const images = await extractPageImages(page, {
    includeData: true,
    maxBytes: 500_000,
    limit: 4
  });
  console.log(images);
})();

```

## Response Format and Structure

A successful response returns a JSON object containing the `tabId` and an `images` array. When `includeData` is enabled and size constraints are met, each object includes a `dataUrl` field with the base64-encoded content.

Standard response structure:

```json
{
  "tabId": "abc123",
  "images": [
    {
      "src": "https://example.com/foo.png",
      "alt": "Foo",
      "width": 640,
      "height": 480,
      "mimeType": "image/png",
      "bytes": 34212,
      "dataUrl": "data:image/png;base64,iVBORw0KGgoAAA..."
    },
    {
      "src": "data:image/svg+xml;base64,PHN2ZyB4bWxucz0...",
      "alt": "",
      "width": 100,
      "height": 100,
      "mimeType": "image/svg+xml",
      "bytes": 1234,
      "dataUrl": "data:image/svg+xml;base64,PHN2ZyB4bWxucz0..."
    }
  ]
}

```

If `includeData` is omitted or the image exceeds `maxBytes`, the response contains only metadata with a `dataSkipped` flag indicating the limit was hit.

## Summary

- Camofox-browser performs **DOM image extraction** entirely within the Playwright page context to avoid triggering network-level security scanners like OpenClaw.
- The `extractPageImages` function in [`lib/images.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/images.js) collects visible `<img>` elements and optionally converts them to inline data URLs using `FileReader`.
- The `GET /tabs/:tabId/images` endpoint accepts `includeData`, `maxBytes`, and `limit` parameters to control output size and content.
- Images exceeding the `maxBytes` threshold (default 20 MiB from [`lib/downloads.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/downloads.js)) return metadata only with a `dataSkipped` flag rather than base64 data.
- All authenticated fetching uses `credentials: 'include'` to support session-based image access without exposing cookies to the Node process.

## Frequently Asked Questions

### Why does camofox-browser extract images inside the page context instead of using Node.js network requests?

Running the extraction inside the browser context prevents OpenClaw scanner false-positives that could occur if the Node server initiated direct network calls to remote image URLs. By using `fetch` within the Playwright page with `credentials: 'include'`, the browser handles cookies and authentication naturally while keeping the server-side footprint minimal.

### What happens if an image exceeds the maxBytes limit?

When an image blob or data URL exceeds the specified `maxBytes` threshold, the entry includes a `dataSkipped` flag set to `true` and omits the `dataUrl` field. The metadata—such as `src`, `alt`, dimensions, and MIME type—remains available, but the inline binary data is excluded to prevent oversized responses.

### Does camofox-browser support authenticated image fetching?

Yes. The in-page fetch operation in [`lib/images.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/images.js) uses `credentials: 'include'`, allowing the browser to send cookies and authentication headers associated with the target domain. This enables extraction of images that require session-based authentication without exposing credentials to the Node process.

### How many images can I extract per request?

By default, the system returns up to 8 distinct images, but you can request a maximum of 20 by setting the `limit` parameter. The hard cap at 20 prevents excessive memory usage and processing time during the DOM scan and base64 conversion operations.