How Camofox-Browser's Accessibility Snapshot Reduces Token Usage Compared to Raw HTML
Camofox-Browser's accessibility snapshot reduces token usage by returning a compact JSON representation of the page's semantic structure instead of raw HTML, typically shrinking payloads by 10-20x by stripping markup, scripts, and styles while preserving only the text content and element references an LLM needs.
Camofox-Browser (from the jo-inc/camofox-browser repository) eliminates the cost penalty of sending full HTML to large language models by implementing an accessibility-tree snapshot system. Instead of transmitting thousands of characters of nested tags, inline CSS, and JavaScript, the browser extracts only the semantic elements and visible text content that an AI agent requires for decision-making. This approach fundamentally changes how automation tools consume web content, shifting from document parsing to intent-based interaction.
What the Accessibility Snapshot Contains
The accessibility snapshot is not the HTML source. It is a purified data structure generated from Firefox's accessibility API that captures only the interactive and readable surface of a web page.
The snapshot payload includes:
- Semantic element types – Identifies headings, paragraphs, links, buttons, and other interactive roles without DOM tag noise.
- Plain-text content – The actual human-readable strings displayed on the page, stripped of HTML entities and formatting tags.
- Stable element references – Short identifiers like
e1,e2, etc., that allow the agent to target specific elements for clicks or text input without transmitting the full DOM hierarchy. - Hierarchical depth – Parent-child relationships that provide navigation context without the cost of nested tag structures.
Because the snapshot is essentially a compact, plain-text outline, its JSON payload avoids tokenizing the thousands of irrelevant characters found in standard HTML.
How the Accessibility Snapshot Reduces Token Usage (10-20x Compression)
Raw HTML contains massive amounts of data that never contribute to an LLM's reasoning but still consume tokens: attribute declarations, inline styles, script blocks, and deeply nested <div> structures. The accessibility snapshot eliminates this waste entirely.
| Component | Raw HTML Impact | Snapshot Replacement | Token Savings |
|---|---|---|---|
| Markup syntax | Every <, >, and / is tokenized |
Removed; only key-value pairs remain | Eliminates structural noise |
| CSS & scripts | Inline and external references add hundreds of tokens | Excluded entirely | Zero styling or script tokens |
| DOM depth | Nested tags multiply with each level | Flattened hierarchy with depth integers | Reduces structural overhead |
| Element targeting | Complex XPath or CSS selectors | Short refs like e7 |
Minimal reference cost |
The result is a 10-20× smaller payload compared to the raw HTML string. When tokenized, the snapshot's count is driven almost entirely by visible text content rather than markup characters, directly lowering API costs and speeding up inference.
Technical Implementation: Building Snapshots in lib/snapshot.js
The snapshot generation logic resides in lib/snapshot.js. This module interfaces with Firefox's native accessibility API to construct the tree:
- Tree Walking – Parses the live page and traverses the accessibility tree (not the DOM tree).
- Field Extraction – Pulls only the role, name, and hierarchical position of each accessible element.
- Serialization – Condenses the data into a concise JSON structure.
- Endpoint Delivery – The HTTP endpoint
/tabs/:tabId/snapshotreturns this JSON to callers.
The implementation deliberately avoids document.documentElement.outerHTML or similar full-source methods. Instead, it queries Firefox's accessibility engine for the semantic relationships that matter to automation agents, discarding presentation-layer data that would bloat the token count.
Working with the Snapshot API: Code Examples
Fetching a Snapshot
Use the REST API to retrieve the accessibility snapshot for a specific tab:
import fetch from 'node-fetch';
const tabId = 'abc123';
const userId = 'agent1';
const base = 'http://localhost:9377';
const resp = await fetch(
`${base}/tabs/${tabId}/snapshot?userId=${userId}`
);
const snapshot = await resp.json();
console.log('Snapshot size (bytes):', JSON.stringify(snapshot).length);
console.log('First few elements:', snapshot.elements.slice(0, 5));
Interacting via Stable References
The snapshot returns stable references (e.g., e7) that you use for subsequent interactions without resending the full page context:
// Find the element by its semantic properties
const linkRef = snapshot.elements.find(
e => e.role === 'link' && /more info/i.test(e.name)
).ref;
// Click using the lightweight reference identifier
await fetch(`${base}/tabs/${tabId}/click`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ userId, ref: linkRef })
});
Measuring Token Savings
Compare token counts between raw HTML and the snapshot using a tokenizer:
import tiktoken
def tokens(text):
return len(tiktoken.encoding_for_model("gpt-4").encode(text))
raw_html = "<html>…very long …</html>" # Full page source
snapshot_json = '{"elements":[{"ref":"e1","role":"heading","name":"Example Domain"}]}'
print("HTML tokens:", tokens(raw_html))
print("Snapshot tokens:", tokens(snapshot_json))
Typical output demonstrates the dramatic reduction:
HTML tokens: 2,300
Snapshot tokens: 120
Summary
- Camofox-Browser's
lib/snapshot.jsgenerates accessibility snapshots via the/tabs/:tabId/snapshotendpoint instead of returning raw HTML. - The snapshot includes semantic roles, plain text, stable references, and hierarchy depth—nothing else.
- Payloads are 10-20× smaller than raw HTML because they exclude markup, scripts, styles, and DOM nesting.
- Stable references (
e1,e2) allow agents to interact with elements without transmitting full selectors or page context. - Token costs are driven only by visible text content, not structural noise, reducing LLM inference expenses significantly.
Frequently Asked Questions
What endpoint returns the accessibility snapshot?
The snapshot is available at /tabs/:tabId/snapshot. Pass the userId as a query parameter to authenticate the request. The endpoint returns a JSON object containing the elements array, as documented in the repository's AGENTS.md file.
How does the snapshot handle large pages?
Camofox-Browser implements payload truncation to prevent oversized responses. The logic in lib/snapshot.js safely limits the tree depth and element count, ensuring the JSON remains compact even for complex single-page applications. End-to-end tests in tests/e2e/snapshot-truncation.test.js verify this behavior.
Can I use the snapshot with any LLM?
Yes. The snapshot is LLM-agnostic JSON. Because it contains only plain text and semantic identifiers, it works with GPT-4, Claude, Llama, or any other model that accepts structured text input. The format is optimized for models that need to understand page structure without parsing HTML.
What is a "stable reference" in the snapshot?
A stable reference is a short string identifier (like e1 or e42) assigned to each accessible element in the snapshot. These references remain consistent for the duration of the page session, allowing your agent to issue commands like "click e7" or "type in e12" without constructing complex CSS selectors or XPaths, further reducing token usage in conversational loops.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →