How to Capture and Retrieve Browser Downloads via the camofox-browser API
camofox-browser automatically intercepts every file downloaded by the headless Chromium instance through Playwright's download event, storing metadata and temporary files that you can retrieve via the GET /tabs/:tabId/downloads endpoint with optional base64 data inclusion.
The camofox-browser API provides a robust mechanism for capturing browser downloads in headless environments. By leveraging Playwright's native download events, the server stores downloaded files temporarily and exposes them through a dedicated HTTP endpoint. This guide explains how the download capture system works internally and how to retrieve those files programmatically using the jo-inc/camofox-browser repository.
How Download Capture Works
Attaching the Event Listener
The download interception starts when a new tab is created. In lib/downloads.js, the attachDownloadListener function (lines 63-71) registers a Playwright download event listener exactly once per tab. This function is invoked from setupTab in server.js during tab initialization, ensuring every browsing context can capture files without duplicate listeners.
Processing Download Events
When the browser initiates a download, the listener executes a specific sequence defined in lib/downloads.js (lines 67-99):
- Generates a UUID to uniquely identify the download.
- Sanitizes the filename using
sanitizeFilenameto remove unsafe characters. - Saves the file to a temporary location inside the OS temp directory.
- Records metadata in the tab's
tabState.downloadsarray, includingid,tabId,url,suggestedFilename,mimeType,bytes,createdAt, and anyfailuremessages. - Enforces retention limits by trimming the array to the newest 20 entries (
MAX_DOWNLOAD_RECORDS_PER_TAB), as implemented in lines 39-44.
Retrieving Captured Downloads via the API
The Downloads Endpoint
To access stored downloads, send a request to the dedicated endpoint defined in server.js (lines 2418-2444):
GET /tabs/:tabId/downloads?userId=<uid>&includeData=true&consume=true&maxBytes=5242880
The endpoint performs the following actions:
- Validates the tab exists and belongs to the user.
- Increments the tab's
toolCallscounter. - Invokes
getDownloadsList(tabState, {includeData, maxBytes})fromlib/downloads.js(lines 118-148) to construct the response.
Including File Data in Responses
When you set includeData=true, the getDownloadsList function processes each download record:
- Removes the internal
filePathfield for security. - Reads the temporary file and returns a base64-encoded string in the
dataBase64field. - Respects the
maxBytesparameter (default 20 MiB); if exceeded, returnsdataSkipped: "max_bytes_exceeded"instead of the data. - Reports any filesystem errors as
readErrorin the response object.
Consuming Download Records
Set consume=true to delete the download records after retrieval. This triggers clearTabDownloads (lines 46-61 in lib/downloads.js), which:
- Removes the in-memory references from
tabState.downloads. - Deletes the underlying temporary files from the OS temp directory.
For session-wide cleanup, clearSessionDownloads iterates through all tabs and invokes the clear function for each.
Practical Code Examples
The following examples use the bundled test client from tests/helpers/client.js to demonstrate the complete workflow.
1. Create a Tab and Trigger a Download
import { createClient } from './tests/helpers/client.js';
const client = createClient('http://localhost:9377'); // camofox-browser default port
const { tabId } = await client.createTab('http://localhost:9377/download-page');
// The test site serves a link with id="downloadLink"
await client.click(tabId, { selector: '#downloadLink' });
The click action causes Playwright to fire the download event, which the server captures automatically using the mechanism in lib/downloads.js.
2. Retrieve Download Metadata
// Poll until the server reports at least one download
let downloads = [];
while (downloads.length === 0) {
const result = await client.getDownloads(tabId);
downloads = result.downloads;
}
// Inspect the first entry
console.log('File name:', downloads[0].suggestedFilename);
console.log('MIME type:', downloads[0].mimeType);
console.log('Size (bytes):', downloads[0].bytes);
3. Retrieve Binary Content as Base64
const { downloads } = await client.getDownloads(tabId, {
includeData: true, // ask server to embed file data
maxBytes: 5 * 1024 * 1024, // raise limit if needed (default 20 MiB)
});
// `dataBase64` contains the file contents
const fileBuffer = Buffer.from(downloads[0].dataBase64, 'base64');
require('fs').writeFileSync('downloaded-file', fileBuffer);
4. Consume (Delete) Stored Records
await client.getDownloads(tabId, { consume: true });
After this call, the tab's downloads array is cleared and the temporary files are removed via clearTabDownloads.
5. Full Session Cleanup
await client.cleanup(); // closes all tabs and the session, removing any lingering temp files
Summary
- Automatic interception: The
attachDownloadListenerfunction inlib/downloads.jscaptures all Playwright download events once per tab. - Secure storage: Files are saved to temporary OS directories with sanitized filenames, and metadata is limited to 20 entries per tab (
MAX_DOWNLOAD_RECORDS_PER_TAB). - Data retrieval: Use
GET /tabs/:tabId/downloadswithincludeData=trueto receive base64-encoded file contents, subject tomaxByteslimits. - Resource cleanup: Set
consume=trueto delete records and temp files immediately, or useclearTabDownloadsfor manual cleanup.
Frequently Asked Questions
How many downloads does camofox-browser store per tab?
The system retains a maximum of 20 download records per tab, defined by the constant MAX_DOWNLOAD_RECORDS_PER_TAB in lib/downloads.js (lines 39-44). When this limit is exceeded, the oldest entries are automatically removed from the tabState.downloads array while keeping the newest 20.
Can I retrieve the actual file contents or just metadata?
You can retrieve both. By default, the API returns metadata only (filename, MIME type, size, URL). To receive the actual file data, add includeData=true to your query parameters. The server will then return a base64-encoded string in the dataBase64 field for each download, provided the file size does not exceed the maxBytes limit (default 20 MiB).
What happens to temporary files after I consume the download?
When you request downloads with consume=true, the server invokes clearTabDownloads (lines 46-61), which removes the in-memory download records from tabState.downloads and deletes the underlying temporary files from the OS temp directory. This ensures no orphaned files remain on the server.
Is there a size limit for retrieved file data?
Yes. The maxBytes query parameter controls how much data the server will read and encode. The default limit is 20,971,520 bytes (20 MiB). If a file exceeds this limit, the API returns dataSkipped: "max_bytes_exceeded" instead of the base64 content, allowing you to handle large files differently or request them in smaller chunks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →