How to Click, Fill, and Hover on Located Elements in Ego-Browser
Ego-Browser exposes a Playwright-style API that resolves CSS selectors, XPath expressions, or snapshot references to concrete DOM elements, then dispatches CDP (Chrome DevTools Protocol) events to perform clicks, text input, and hover actions.
Ego-Browser (citrolabs/ego-lite) provides automated agents with a high-level interaction layer for manipulating web pages through intuitive locator syntax. The library abstracts Chrome DevTools Protocol complexity into simple commands like click, fill, and hover, automatically handling element resolution and event dispatching regardless of whether you target elements by CSS selector, @ref snapshot IDs, or Playwright-style role selectors.
Locator Resolution Architecture
Before any mouse or keyboard action executes, Ego-Browser translates the target identifier into a concrete DOM element. The element resolver (src/element-resolver.ts) serves as the foundation for all interactions, parsing strings that may contain CSS selectors, XPath expressions, loc= shortcuts, or @ref snapshot references to locate elements and compute their screen coordinates.
Supported Locator Formats
The resolver accepts multiple syntaxes for locating elements:
- CSS selectors: Standard DOM selectors like
button[type=submit] - XPath expressions: XML path queries for complex DOM traversal
- @ref snapshots: Numeric identifiers like
@21referencing cached element snapshots - loc= shortcuts: Abbreviated locator syntax for common patterns
- Playwright-style role selectors: Semantic queries like
getByRoleorgetByLabel
Performing Mouse Actions
Mouse interactions—including click, dblclick, hover, and drag—are implemented in src/driver/pointer.ts. After resolution, these helpers construct sequences of CDP mouse events and transmit them via Input.dispatchMouseEvent.
Click and Hover Implementation Details
The pointer driver orchestrates three-phase event sequences:
mouseMoved– positions the cursor over the element's centermousePressed– initiates the button pressmouseReleased– completes the interaction
To ensure reliability, the driver installs verification probes through installClickProbe and installHoverProbe, confirming the DOM mutation or state change occurred before returning control to the caller.
Text Input and Keyboard Actions
Text entry operations reside in src/driver/keyboard.ts. The fill method implements a complete input workflow: it focuses the target element, optionally clears existing content using Input.insertText, then fires input and change events to trigger native browser validation and reactive frameworks.
Fill vs. Type Sequences
While fill replaces content atomically, the driver also exposes typeText and pressSequentially for simulating individual keystrokes. All keyboard methods rely on the same element resolution logic to ensure consistent targeting across the API surface.
Practical Usage Examples
Interact with pages through the unified page object exported from src/helpers.ts. This façade wraps the low-level drivers and supports Playwright-style chained locators.
// Click a button using a CSS selector
await page.locator('button[type=submit]').click();
// Click using a Playwright-style role selector
await page.getByRole('button', { name: 'Submit' }).click();
// Fill an input identified by its label
await page.getByLabel('Email').fill('joe@example.com');
// Hover over an element identified by an @ref snapshot
await page.locator('@21').hover();
// Direct mouse actions on arbitrary coordinates
await page.mouse.click(420, 260);
await page.mouse.move(100, 200);
await page.mouse.wheel({ deltaY: 120 });
Advanced Low-Level Driver Access
For scenarios requiring precise control over event timing or coordinate offsets, import the driver modules directly to bypass the locator façade.
import * as pointer from 'ego-browser/src/driver/pointer.js';
import * as keyboard from 'ego-browser/src/driver/keyboard.js';
// Click with specific offset from element top-left
await pointer.click({ selector: '#login', x: 10, y: 5 });
// Direct keyboard fill without helper wrapping
await keyboard.fill('#search', 'ego-browser');
// Hover with explicit selector
await pointer.hover('#menu');
The documentation generator (src/format.ts) parses JSDoc comments from these modules to auto-generate the help() output, providing discoverable syntax references for all available actions.
Summary
- Element resolution occurs in
src/element-resolver.ts, translating CSS, XPath,@ref, and role-based selectors into screen coordinates and object handles. - Mouse actions (
src/driver/pointer.ts) dispatch CDP events viaInput.dispatchMouseEventand verify success usinginstallClickProbeandinstallHoverProbe. - Keyboard actions (
src/driver/keyboard.ts) handle text entry throughfill, leveragingInput.insertTextand firing nativeinput/changeevents. - Public API (
src/helpers.ts) exposes a unifiedpageobject supporting Playwright-style locators likepage.getByLabel().fill(). - Documentation (
src/format.ts) auto-generates help text from JSDoc comments embedded in the source.
Frequently Asked Questions
What locator syntaxes does Ego-Browser support?
Ego-Browser supports CSS selectors, XPath expressions, @ref snapshot references (e.g., @21), loc= shortcuts, and Playwright-style semantic locators including getByRole and getByLabel. The element resolver in src/element-resolver.ts normalizes all formats before interaction.
How does Ego-Browser verify that clicks and hovers succeeded?
The pointer driver installs verification probes—installClickProbe for clicks and installHoverProbe for hovers—that monitor the DOM for expected state changes or event propagation before completing the promise. This ensures the element was actually interactive and the action registered.
Can I interact with elements using direct coordinates instead of selectors?
Yes. While the page.locator() API requires a selector, the low-level page.mouse API and direct driver imports from src/driver/pointer.ts accept absolute screen coordinates or offset objects, allowing interaction with elements that lack stable selectors or for pixel-perfect precision.
Where is the documentation for available helpers generated?
Runtime documentation derives from JSDoc comments parsed by src/format.ts. This module extracts function signatures, parameters, and examples to populate the help() output, ensuring the API remains self-documenting as src/helpers.ts and driver modules evolve.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →