How Ego-Lite Captures Snapshots and Page State: A Complete Technical Guide
Ego-Lite captures page state through a three-stage pipeline—snapshotRaw() fetches structured DOM data from the native runtime, browserSnapshotRefsToRefMap() builds a global ref-map for element resolution, and snapshot() returns a simplified text view for AI agents.
The citrolabs/ego-lite browser automation framework bridges the gap between raw browser internals and semantic, LLM-friendly page representations. Its snapshot system enables AI agents to reason about web pages using stable references like @23 instead of fragile CSS selectors. This article breaks down the complete implementation, from the native runtime call to the ref-based element interactions you use in agent code.
The Three-Stage Snapshot Pipeline
Ego-Lite's snapshot and page state capture works through three tightly-coupled stages defined in driver/observe.ts.
Stage 1: Request Raw Snapshot Data
The internal helper snapshotRaw() initiates capture by calling the native ego runtime through browserEgo().snapshot(options):
// driver/observe.ts, lines 49-63
export async function snapshotRaw(options?: SnapshotOptions): Promise<SnapshotResult> {
const result = await browserEgo().snapshot({
scope: options?.scope ?? "full_page",
includeActionMarks: options?.includeActionMarks ?? true,
includeStableLocator: options?.includeStableLocator ?? true,
});
// result contains {content: string, refs: SnapshotRef[]}
return result;
}
The runtime returns a structured object with two critical properties:
content— The rendered, accessible text representation of the pagerefs— An array of snapshot references like@23that map to backend node IDs
Each ref is a short, stable identifier that survives DOM mutations better than CSS selectors.
Stage 2: Refresh the Internal Ref-Map
Once the raw snapshot arrives, the system populates the global ref-map (browserRefMap) to enable future element resolution:
// driver/observe.ts, lines 61-66
browserSnapshotRefsToRefMap(browserRefMap, result.refs || []);
registerSnapshotForRefRefresh(() => snapshotRaw());
This two-step update mechanism ensures:
- Immediate mapping — Each
@Nref is tied to its backend node ID and selector metadata - Automatic refresh — Future ref lookups trigger fresh snapshots if the map becomes stale
The ref-map lives in ref-state.ts and persists across agent actions, making await click("@23") possible even after minor page changes.
Stage 3: Deliver the Convenient Text View
The public snapshot() helper wraps snapshotRaw() with opinionated defaults and strips the structured metadata:
// driver/observe.ts, lines 73-80
export async function snapshot(options?: SnapshotOptions): Promise<string> {
const result = await snapshotRaw({
scope: options?.scope ?? "full_page",
includeActionMarks: options?.includeActionMarks ?? true,
includeStableLocator: options?.includeStableLocator ?? true,
});
return result.content; // Only the text surface
}
snapshot() is what most agents call (await page.snapshot()). It returns clean, LLM-ready text without exposing implementation details.
snapshotRaw() is reserved for advanced use cases where you need the full {content, refs} object or custom options.
How the Ref-Map Enables Element Resolution
The ref-map system—spanning ref-state.ts, ref-map.ts, and element-resolver.ts—translates snapshot references into executable element handles.
Core Components
| Component | File | Responsibility |
|---|---|---|
browserRefMap |
src/ref-state.ts |
Singleton Map<number, RefInfo> holding all active refs |
browserSnapshotRefsToRefMap() |
src/ref-state.ts |
Populates the map from snapshot results |
| Lookup utilities | src/ref-map.ts |
Resolves refs, triggers refresh on stale lookups |
| Element resolution | src/element-resolver.ts |
Converts refs to executable handles for click, evaluate, etc. |
Resolution Flow
When you call await click("@23"), the resolver:
- Checks
browserRefMapfor entry23 - If missing or stale, triggers
snapshotRaw()via the registered refresh hook - Retrieves the backend node ID from the refreshed map
- Executes the action through
browser-runtime.ts
This indirection insulates agent code from DOM volatility.
Snapshot Options and Page State Granularity
The SnapshotOptions interface controls how the ego runtime serializes page state:
| Option | Values | Effect on Page State Capture |
|---|---|---|
scope |
"full_page" (default) |
Captures entire scrollable document |
scope |
"only_within_viewport" |
Limits to currently visible area |
includeActionMarks |
true (default) |
Annotates interactable elements with action hints |
includeStableLocator |
true (default) |
Computes durable CSS/XPath selectors for reuse |
The ego runtime (closed-source binary) performs the heavy lifting: DOM serialization, accessible name computation, and optional annotation injection.
Practical Code Examples
Basic Text Snapshot
// Standard agent pattern—clean text for LLM consumption
const pageText = await page.snapshot();
console.log(pageText);
// Output: semantic text with @23-style refs embedded for key elements
Uses defaults: full_page, action marks enabled, stable locators enabled.
Structured Raw Snapshot
// Access refs and stable locators directly
const raw = await page.snapshotRaw({
scope: "only_within_viewport",
includeActionMarks: false,
includeStableLocator: true,
});
console.log(raw.content); // Visible text only
console.log(raw.refs);
// [{ref: "@12", nodeId: 12345, stableLocator: "css=button.submit", ...}, ...]
Ref-Based Element Interaction
// Snapshot, find element by content, click by ref
const snap = await page.snapshotRaw();
const targetRef = snap.refs.find(r => r.text.includes("Submit"))?.ref; // "@7"
await page.click(targetRef); // Resolver handles the @7 → nodeId translation
Manual Ref-Map Refresh
// Force synchronization (rarely needed—automatic on stale access)
await page.snapshot(); // Fresh snapshot; ref-map updated automatically
Key Source Files
| File | Role |
|---|---|
src/driver/observe.ts |
snapshotRaw(), snapshot(), refresh hook registration |
src/ref-state.ts |
browserRefMap singleton, browserSnapshotRefsToRefMap() |
src/ref-map.ts |
Ref lookup utilities, freshness checking |
src/browser-runtime.ts |
Bridge to native browserEgo().snapshot() |
src/helpers.ts |
Exposes snapshot/snapshotRaw to agent globals |
src/format.ts |
Public API type definitions |
src/element-resolver.ts |
Ref-to-handle resolution for actions |
Summary
snapshotRaw()indriver/observe.tscalls the native ego runtime and returns structured{content, refs}databrowserSnapshotRefsToRefMap()inref-state.tsbuilds the global ref-map that powers stable element referencessnapshot()provides the simplified text view most agents consume, with sensible defaults for scope and annotations- The ref-map system enables durable
@Nreferences that survive DOM changes better than raw selectors - Snapshot options control granularity—viewport vs. full page, action marks, and stable locator computation
Frequently Asked Questions
What is the difference between snapshot() and snapshotRaw()?
snapshot() returns a plain string of page content optimized for LLM prompts. snapshotRaw() returns the full structured object including the refs array and stable locators. Use snapshotRaw() when you need to resolve element references programmatically before acting.
How does the @23 ref system work under the hood?
Each @N reference maps to a backend node ID stored in browserRefMap. When you pass "@23" to click() or elementCenter(), the resolver looks up node ID 23, validates freshness, refreshes the snapshot if needed, and executes against the corresponding DOM node.
Can I capture only the visible viewport instead of the full page?
Yes. Pass scope: "only_within_viewport" to snapshotRaw() or snapshot(). This reduces payload size and processing time for large documents when you only need immediately visible content.
What happens if a ref becomes stale between snapshot and action?
The ref-map automatically triggers snapshotRaw() via the registered refresh hook in registerSnapshotForRefRefresh(). This ensures the map stays consistent without manual intervention, though it adds one snapshot latency on first stale access.
Where is the actual DOM serialization implemented?
The heavy lifting occurs in the closed-source ego runtime binary, invoked through browserEgo().snapshot() in browser-runtime.ts. The open-source TypeScript layer handles option marshaling, ref-map maintenance, and convenient API surfacing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →