# How Camofox-Browser's Accessibility Snapshot Reduces Token Usage Compared to Raw HTML

> Camofox-Browser's accessibility snapshot slashes token usage 10-20x by sending compact JSON instead of raw HTML. Get essential text and element data for LLMs, reducing payload size significantly.

- Repository: [jo/camofox-browser](https://github.com/jo-inc/camofox-browser)
- Tags: performance
- Published: 2026-04-15

---

**Camofox-Browser's accessibility snapshot reduces token usage by returning a compact JSON representation of the page's semantic structure instead of raw HTML, typically shrinking payloads by 10-20x by stripping markup, scripts, and styles while preserving only the text content and element references an LLM needs.**

Camofox-Browser (from the `jo-inc/camofox-browser` repository) eliminates the cost penalty of sending full HTML to large language models by implementing an **accessibility-tree snapshot** system. Instead of transmitting thousands of characters of nested tags, inline CSS, and JavaScript, the browser extracts only the semantic elements and visible text content that an AI agent requires for decision-making. This approach fundamentally changes how automation tools consume web content, shifting from document parsing to intent-based interaction.

## What the Accessibility Snapshot Contains

The **accessibility snapshot** is not the HTML source. It is a purified data structure generated from Firefox's accessibility API that captures only the interactive and readable surface of a web page.

The snapshot payload includes:

- **Semantic element types** – Identifies headings, paragraphs, links, buttons, and other interactive roles without DOM tag noise.
- **Plain-text content** – The actual human-readable strings displayed on the page, stripped of HTML entities and formatting tags.
- **Stable element references** – Short identifiers like `e1`, `e2`, etc., that allow the agent to target specific elements for clicks or text input without transmitting the full DOM hierarchy.
- **Hierarchical depth** – Parent-child relationships that provide navigation context without the cost of nested tag structures.

Because the snapshot is essentially a **compact, plain-text outline**, its JSON payload avoids tokenizing the thousands of irrelevant characters found in standard HTML.

## How the Accessibility Snapshot Reduces Token Usage (10-20x Compression)

Raw HTML contains massive amounts of data that never contribute to an LLM's reasoning but still consume tokens: attribute declarations, inline styles, script blocks, and deeply nested `<div>` structures. The accessibility snapshot eliminates this waste entirely.

| Component | Raw HTML Impact | Snapshot Replacement | Token Savings |
|-----------|----------------|----------------------|---------------|
| **Markup syntax** | Every `<`, `>`, and `/` is tokenized | Removed; only key-value pairs remain | Eliminates structural noise |
| **CSS & scripts** | Inline and external references add hundreds of tokens | Excluded entirely | Zero styling or script tokens |
| **DOM depth** | Nested tags multiply with each level | Flattened hierarchy with depth integers | Reduces structural overhead |
| **Element targeting** | Complex XPath or CSS selectors | Short refs like `e7` | Minimal reference cost |

The result is a **10-20× smaller payload** compared to the raw HTML string. When tokenized, the snapshot's count is driven almost entirely by visible text content rather than markup characters, directly lowering API costs and speeding up inference.

## Technical Implementation: Building Snapshots in lib/snapshot.js

The snapshot generation logic resides in **[`lib/snapshot.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/snapshot.js)**. This module interfaces with Firefox's native accessibility API to construct the tree:

1. **Tree Walking** – Parses the live page and traverses the accessibility tree (not the DOM tree).
2. **Field Extraction** – Pulls only the role, name, and hierarchical position of each accessible element.
3. **Serialization** – Condenses the data into a concise JSON structure.
4. **Endpoint Delivery** – The HTTP endpoint `/tabs/:tabId/snapshot` returns this JSON to callers.

The implementation deliberately avoids `document.documentElement.outerHTML` or similar full-source methods. Instead, it queries Firefox's accessibility engine for the semantic relationships that matter to automation agents, discarding presentation-layer data that would bloat the token count.

## Working with the Snapshot API: Code Examples

### Fetching a Snapshot

Use the REST API to retrieve the accessibility snapshot for a specific tab:

```javascript
import fetch from 'node-fetch';

const tabId = 'abc123';
const userId = 'agent1';
const base = 'http://localhost:9377';

const resp = await fetch(
  `${base}/tabs/${tabId}/snapshot?userId=${userId}`
);
const snapshot = await resp.json();

console.log('Snapshot size (bytes):', JSON.stringify(snapshot).length);
console.log('First few elements:', snapshot.elements.slice(0, 5));

```

### Interacting via Stable References

The snapshot returns stable references (e.g., `e7`) that you use for subsequent interactions without resending the full page context:

```javascript
// Find the element by its semantic properties
const linkRef = snapshot.elements.find(
  e => e.role === 'link' && /more info/i.test(e.name)
).ref;

// Click using the lightweight reference identifier
await fetch(`${base}/tabs/${tabId}/click`, {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({ userId, ref: linkRef })
});

```

### Measuring Token Savings

Compare token counts between raw HTML and the snapshot using a tokenizer:

```python
import tiktoken

def tokens(text): 
    return len(tiktoken.encoding_for_model("gpt-4").encode(text))

raw_html = "<html>…very long …</html>"  # Full page source

snapshot_json = '{"elements":[{"ref":"e1","role":"heading","name":"Example Domain"}]}'

print("HTML tokens:", tokens(raw_html))
print("Snapshot tokens:", tokens(snapshot_json))

```

Typical output demonstrates the dramatic reduction:

```

HTML tokens: 2,300
Snapshot tokens: 120

```

## Summary

- Camofox-Browser's **[`lib/snapshot.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/snapshot.js)** generates accessibility snapshots via the **`/tabs/:tabId/snapshot`** endpoint instead of returning raw HTML.
- The snapshot includes **semantic roles**, **plain text**, **stable references**, and **hierarchy depth**—nothing else.
- Payloads are **10-20× smaller** than raw HTML because they exclude markup, scripts, styles, and DOM nesting.
- **Stable references** (`e1`, `e2`) allow agents to interact with elements without transmitting full selectors or page context.
- Token costs are driven only by **visible text content**, not structural noise, reducing LLM inference expenses significantly.

## Frequently Asked Questions

### What endpoint returns the accessibility snapshot?

The snapshot is available at **`/tabs/:tabId/snapshot`**. Pass the `userId` as a query parameter to authenticate the request. The endpoint returns a JSON object containing the `elements` array, as documented in the repository's [`AGENTS.md`](https://github.com/jo-inc/camofox-browser/blob/main/AGENTS.md) file.

### How does the snapshot handle large pages?

Camofox-Browser implements **payload truncation** to prevent oversized responses. The logic in [`lib/snapshot.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/snapshot.js) safely limits the tree depth and element count, ensuring the JSON remains compact even for complex single-page applications. End-to-end tests in [`tests/e2e/snapshot-truncation.test.js`](https://github.com/jo-inc/camofox-browser/blob/main/tests/e2e/snapshot-truncation.test.js) verify this behavior.

### Can I use the snapshot with any LLM?

Yes. The snapshot is **LLM-agnostic** JSON. Because it contains only plain text and semantic identifiers, it works with GPT-4, Claude, Llama, or any other model that accepts structured text input. The format is optimized for models that need to understand page structure without parsing HTML.

### What is a "stable reference" in the snapshot?

A **stable reference** is a short string identifier (like `e1` or `e42`) assigned to each accessible element in the snapshot. These references remain consistent for the duration of the page session, allowing your agent to issue commands like "click e7" or "type in e12" without constructing complex CSS selectors or XPaths, further reducing token usage in conversational loops.