How to Capture and Retrieve Browser Downloads via the camofox-browser API

camofox-browser automatically intercepts every file downloaded by the headless Chromium instance through Playwright's download event, storing metadata and temporary files that you can retrieve via the GET /tabs/:tabId/downloads endpoint with optional base64 data inclusion.

The camofox-browser API provides a robust mechanism for capturing browser downloads in headless environments. By leveraging Playwright's native download events, the server stores downloaded files temporarily and exposes them through a dedicated HTTP endpoint. This guide explains how the download capture system works internally and how to retrieve those files programmatically using the jo-inc/camofox-browser repository.

How Download Capture Works

Attaching the Event Listener

The download interception starts when a new tab is created. In lib/downloads.js, the attachDownloadListener function (lines 63-71) registers a Playwright download event listener exactly once per tab. This function is invoked from setupTab in server.js during tab initialization, ensuring every browsing context can capture files without duplicate listeners.

Processing Download Events

When the browser initiates a download, the listener executes a specific sequence defined in lib/downloads.js (lines 67-99):

  1. Generates a UUID to uniquely identify the download.
  2. Sanitizes the filename using sanitizeFilename to remove unsafe characters.
  3. Saves the file to a temporary location inside the OS temp directory.
  4. Records metadata in the tab's tabState.downloads array, including id, tabId, url, suggestedFilename, mimeType, bytes, createdAt, and any failure messages.
  5. Enforces retention limits by trimming the array to the newest 20 entries (MAX_DOWNLOAD_RECORDS_PER_TAB), as implemented in lines 39-44.

Retrieving Captured Downloads via the API

The Downloads Endpoint

To access stored downloads, send a request to the dedicated endpoint defined in server.js (lines 2418-2444):

GET /tabs/:tabId/downloads?userId=<uid>&includeData=true&consume=true&maxBytes=5242880

The endpoint performs the following actions:

  • Validates the tab exists and belongs to the user.
  • Increments the tab's toolCalls counter.
  • Invokes getDownloadsList(tabState, {includeData, maxBytes}) from lib/downloads.js (lines 118-148) to construct the response.

Including File Data in Responses

When you set includeData=true, the getDownloadsList function processes each download record:

  • Removes the internal filePath field for security.
  • Reads the temporary file and returns a base64-encoded string in the dataBase64 field.
  • Respects the maxBytes parameter (default 20 MiB); if exceeded, returns dataSkipped: "max_bytes_exceeded" instead of the data.
  • Reports any filesystem errors as readError in the response object.

Consuming Download Records

Set consume=true to delete the download records after retrieval. This triggers clearTabDownloads (lines 46-61 in lib/downloads.js), which:

  • Removes the in-memory references from tabState.downloads.
  • Deletes the underlying temporary files from the OS temp directory.

For session-wide cleanup, clearSessionDownloads iterates through all tabs and invokes the clear function for each.

Practical Code Examples

The following examples use the bundled test client from tests/helpers/client.js to demonstrate the complete workflow.

1. Create a Tab and Trigger a Download

import { createClient } from './tests/helpers/client.js';

const client = createClient('http://localhost:9377');   // camofox-browser default port
const { tabId } = await client.createTab('http://localhost:9377/download-page');

// The test site serves a link with id="downloadLink"
await client.click(tabId, { selector: '#downloadLink' });

The click action causes Playwright to fire the download event, which the server captures automatically using the mechanism in lib/downloads.js.

2. Retrieve Download Metadata

// Poll until the server reports at least one download
let downloads = [];
while (downloads.length === 0) {
  const result = await client.getDownloads(tabId);
  downloads = result.downloads;
}

// Inspect the first entry
console.log('File name:', downloads[0].suggestedFilename);
console.log('MIME type:', downloads[0].mimeType);
console.log('Size (bytes):', downloads[0].bytes);

3. Retrieve Binary Content as Base64

const { downloads } = await client.getDownloads(tabId, {
  includeData: true,           // ask server to embed file data
  maxBytes: 5 * 1024 * 1024, // raise limit if needed (default 20 MiB)
});

// `dataBase64` contains the file contents
const fileBuffer = Buffer.from(downloads[0].dataBase64, 'base64');
require('fs').writeFileSync('downloaded-file', fileBuffer);

4. Consume (Delete) Stored Records

await client.getDownloads(tabId, { consume: true });

After this call, the tab's downloads array is cleared and the temporary files are removed via clearTabDownloads.

5. Full Session Cleanup

await client.cleanup();   // closes all tabs and the session, removing any lingering temp files

Summary

  • Automatic interception: The attachDownloadListener function in lib/downloads.js captures all Playwright download events once per tab.
  • Secure storage: Files are saved to temporary OS directories with sanitized filenames, and metadata is limited to 20 entries per tab (MAX_DOWNLOAD_RECORDS_PER_TAB).
  • Data retrieval: Use GET /tabs/:tabId/downloads with includeData=true to receive base64-encoded file contents, subject to maxBytes limits.
  • Resource cleanup: Set consume=true to delete records and temp files immediately, or use clearTabDownloads for manual cleanup.

Frequently Asked Questions

How many downloads does camofox-browser store per tab?

The system retains a maximum of 20 download records per tab, defined by the constant MAX_DOWNLOAD_RECORDS_PER_TAB in lib/downloads.js (lines 39-44). When this limit is exceeded, the oldest entries are automatically removed from the tabState.downloads array while keeping the newest 20.

Can I retrieve the actual file contents or just metadata?

You can retrieve both. By default, the API returns metadata only (filename, MIME type, size, URL). To receive the actual file data, add includeData=true to your query parameters. The server will then return a base64-encoded string in the dataBase64 field for each download, provided the file size does not exceed the maxBytes limit (default 20 MiB).

What happens to temporary files after I consume the download?

When you request downloads with consume=true, the server invokes clearTabDownloads (lines 46-61), which removes the in-memory download records from tabState.downloads and deletes the underlying temporary files from the OS temp directory. This ensures no orphaned files remain on the server.

Is there a size limit for retrieved file data?

Yes. The maxBytes query parameter controls how much data the server will read and encode. The default limit is 20,971,520 bytes (20 MiB). If a file exceeds this limit, the API returns dataSkipped: "max_bytes_exceeded" instead of the base64 content, allowing you to handle large files differently or request them in smaller chunks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →