How camofox-browser Element Refs (e1, e2, e3) Work for Reliable Click and Type Operations

camofox-browser assigns stable identifier strings like e1, e2, and e3 to interactive elements using ARIA role-based extraction, enabling the /click and /type APIs to target elements reliably without brittle CSS selectors.

The camofox-browser project (available at jo-inc/camofox-browser) solves automation fragility by mapping interactive page elements to stable references during snapshot capture. Instead of relying on CSS selectors that break when classes or IDs change, the system generates element refs that persist for the lifetime of the current page and automatically refreshes them when necessary.

How Element Refs Are Generated

When the server captures a page snapshot, it builds a map of interactive elements using two distinct paths depending on the page type. Both methods assign sequential identifiers in the format e${counter} (e.g., e1, e2) to create stable references.

The ARIA Snapshot Path

For standard pages, the buildRefs(page) function in server.js calls Playwright’s ariaSnapshot() method and parses each line matching an interactive role. The implementation at lines 1470–1477 creates a ref for each match and stores it in a Map alongside disambiguation metadata:

// server.js – buildRefs → map of refs
// https://github.com/jo-inc/camofox-browser/blob/master/server.js#L1470-L1477
const seenCounts = new Map();
for (const el of extracted.elements) {
  const key = `${el.role}:${el.name}`;
  const nth = seenCounts.get(key) || 0;
  seenCounts.set(key, nth + 1);
  refs.set(el.id, { role: normalizedRole, name: normalizedName, nth });
}

Each ref entry contains:

  • role: The ARIA role (e.g., button, textbox)
  • name: The accessible name extracted from the element
  • nth: A 0-based index used when multiple elements share the same role and name combination

The Google SERP Fast-Path

For Google search results pages, the extractGoogleSerp() helper uses a local addRef() function to generate refs in the same format. The implementation at lines 1154–1159 in server.js handles the counter increment:

// server.js – addRef helper used by Google SERP extraction
// https://github.com/jo-inc/camofox-browser/blob/master/server.js#L1154-L1159
function addRef(role, name) {
  const id = 'e' + refCounter++;
  elements.push({ id, role, name });
  return id;
}

The resulting snapshot YAML includes these refs next to element descriptions, such as:

- button "Click Me" [e3]:

Resolving Refs to Playwright Locators

When the server receives a request containing a ref parameter, it converts the reference to a robust Playwright locator using role-based querying. The refToLocator function at lines 1401–1408 in server.js performs this lookup:

// server.js – ref → locator conversion
// https://github.com/jo-inc/camofox-browser/blob/master/server.js#L1401-L1408
function refToLocator(page, ref, refs) {
  const info = refs.get(ref);
  if (!info) return null;
  const { role, name, nth } = info;
  let locator = page.getByRole(role, name ? { name } : undefined);
  locator = locator.nth(nth);          // disambiguate duplicates
  return locator;
}

This approach uses Playwright’s getByRole() locator combined with the nth index to guarantee the correct element is selected even when multiple elements share identical roles and accessible names.

The Click and Type API Workflow

The /click and /type endpoints follow a three-step pattern that ensures reliability through automatic ref refreshing:

  1. Capture: The client requests a snapshot via /snapshot, which triggers buildRefs() and returns the current ref mappings.
  2. Select: The client identifies the desired element by parsing the snapshot text (e.g., extracting [e5] from a button description).
  3. Execute: The client sends a request to /click or /type with the ref parameter.

If a ref cannot be resolved (for example, due to navigation since the snapshot was taken), the server auto-refreshes the refs before failing. The click endpoint implementation at lines 1002–1014 in server.js demonstrates this logic:

// server.js – auto-refresh before click when ref missing
// https://github.com/jo-inc/camofox-browser/blob/master/server.js#L1002-L1014
if (!locator) {
  log('info', 'auto-refreshing refs before click', { ref, hadRefs: tabState.refs.size });
  const preClickBudget = Math.min(4000, remainingBudget());
  tabState.refs = await refreshTabRefs(tabState, { reason: 'pre_click', timeoutMs: preClickBudget });
  locator = refToLocator(tabState.page, ref, tabState.refs);
}

The same auto-refresh pattern appears in the type endpoint at lines 2206–2208. If the ref remains unresolvable after refresh, the server returns a StaleRefsError (HTTP 422), signaling the client to request a fresh snapshot.

Why Refs Outperform CSS Selectors

The ref-based approach solves several reliability issues inherent to traditional CSS selectors:

Problem Traditional CSS Selector camofox-browser Element Ref
DOM mutations (class/ID rewrites) Breaks selector Role and name remain stable
Duplicate elements Ambiguous targeting nth index disambiguates
Dynamic async rendering Requires explicit waits Snapshot captures accessible tree after waitForPageReady
Navigation Manual ref rebuilding required Server auto-clears and rebuilds refs

Practical Implementation Example

The following example demonstrates the complete workflow using the camofox-browser API, based on patterns from tests/e2e/formSubmission.test.js:

// 1️⃣ Get a snapshot
const { tabId } = await client.createTab('https://example.com');
let snap = await client.getSnapshot(tabId);
console.log(snap.snapshot); // contains lines like: - button "Submit" [e5]:

// 2️⃣ Extract the ref for the “Submit” button
const match = snap.snapshot.match(/\[(e\d+)\].*button.*Submit/i);
if (!match) throw new Error('Submit button not found');
const submitRef = match[1];

// 3️⃣ Click using the ref (no selector needed)
await client.click(tabId, { ref: submitRef });

// 4️⃣ Type into a textbox using its ref
const txtMatch = snap.snapshot.match(/\[(e\d+)\].*textbox.*Email/i);
const emailRef = txtMatch[1];
await client.type(tabId, { ref: emailRef, text: 'alice@example.com' });

All API calls—including createTab, getSnapshot, click, and type—are implemented in server.js and demonstrated in the test suite. Supporting snapshot logic (truncation and pagination) resides in lib/snapshot.js.

Summary

  • camofox-browser element refs (e1, e2, e3…) are generated from ARIA snapshots in server.js using buildRefs() for standard pages and addRef() for Google SERP extraction.
  • Each ref stores role, name, and nth (disambiguation index) in a Map for stable lookup.
  • The refToLocator() function converts refs to Playwright role-based locators, making them immune to DOM attribute changes.
  • The /click and /type endpoints auto-refresh stale refs (lines 1002–1014 and 2206–2208) before returning a StaleRefsError (HTTP 422) if resolution fails.
  • This architecture eliminates selector brittleness and handles dynamic content through accessibility-tree targeting.

Frequently Asked Questions

How long do camofox-browser element refs remain valid?

Refs are stable for the lifetime of the current page. They persist as long as the page remains loaded and the tab state is maintained. When navigation occurs, the server automatically clears and rebuilds the ref map, requiring the client to obtain a fresh snapshot.

What happens if I try to click an element using a stale ref?

If a ref cannot be resolved to a locator, the server attempts an auto-refresh of the refs (with a budget of up to 4000ms by default) before giving up. If the ref is still unresolved after refresh, the server returns a StaleRefsError with HTTP status 422, prompting the client to request a new snapshot and extract current refs.

Why are element refs more reliable than CSS selectors?

Refs use Playwright’s getByRole() locator based on ARIA roles and accessible names, which remain consistent even when frameworks like React or Vue regenerate DOM attributes dynamically. The nth disambiguation index also ensures that duplicate elements with identical roles and names can be targeted precisely without complex XPath or selector construction.

Where is the ref generation logic located in the codebase?

The core ref generation logic resides in server.js at jo-inc/camofox-browser. Specifically, buildRefs() handles ARIA snapshot parsing (lines 1470–1477), addRef() manages Google SERP extraction (lines 1154–1159), and refToLocator() performs the conversion to Playwright locators (lines 1401–1408). Integration tests demonstrating ref usage are located in tests/e2e/formSubmission.test.js.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →