How to Use Browse CLI Commands for Page Interaction: A Complete Guide
The browse CLI is a daemon-based automation tool that lets you control browsers via command-line commands like browse open, browse snapshot, and browse click, operating in either local Chrome or remote Browserbase cloud environments.
The browse command-line interface, maintained in the browserbase/skills repository, provides a persistent daemon architecture for reliable web automation. Unlike one-shot browser tools, it runs a background process that stays alive across multiple commands, enabling deterministic scripting for complex page interactions. This guide covers the essential browse CLI commands for page interaction based on the implementation details found in skills/browser/SKILL.md and skills/browser/REFERENCE.md.
Understanding the Daemon Architecture
The browse CLI operates through a background daemon that manages the browser process independently of your terminal session. According to skills/browser/REFERENCE.md, this daemon starts automatically on the first command (such as browse open) and continues running until you explicitly execute browse stop.
Each CLI command communicates with this daemon over RPC channels, which eliminates the overhead of launching a new browser instance for every action. Because the browser stays resident in memory, you can chain multiple interactions—navigations, clicks, form fills—without re-initializing the environment, resulting in faster and more reliable automation scripts.
Selecting Your Environment: Local vs Remote
Before executing page interactions, you must select between local and remote modes using browse env:
browse env local: Runs a clean, isolated Chrome instance on your machine. Use--auto-connectto attach to an existing Chrome session with your login states preserved.browse env remote: Connects to Browserbase cloud browsers with anti-bot stealth, automatic CAPTCHA solving, and residential proxies. RequiresBROWSERBASE_API_KEYto be set.
You can switch environments mid-session or use the --session <name> flag to maintain multiple isolated browser contexts simultaneously.
Core Browse CLI Commands for Page Interaction
Navigation with browse open
Start every workflow by opening a target URL. The daemon initializes here if it isn't already running.
browse open https://example.com --wait networkidle
The --wait flag accepts load, domcontentloaded, or networkidle to control when the command returns.
Inspecting Page State
browse snapshot: Returns an accessibility tree containing element references (e.g., @0-5) that map to specific DOM nodes. This is the preferred inspection method for obtaining clickable references.
browse get: Retrieves specific page properties:
browse get title # Current page title
browse get text # Extracted text content
browse get html # Raw HTML source
browse get url # Current address
Always verify state with browse snapshot or browse get before and after interactions to ensure deterministic behavior.
Interacting with Elements
Once you have element references from browse snapshot, use these commands:
browse click @ref: Clicks the exact element corresponding to the reference (e.g., browse click @0-5). Using refs guarantees element identity won't shift due to DOM changes.
browse fill <selector> <value>: Enters text into input fields using CSS selectors:
browse fill "#email" "user@example.com"
browse fill "#password" "s3cr3t!"
browse type <text> and browse press <key>: Send keyboard input or press specific keys like Enter, Tab, or Escape.
browse scroll ...: Scrolls the page vertically or horizontally to reveal elements before interaction.
Managing Sessions and Contexts
browse stop: Shuts down the daemon, closes the browser, and clears any environment overrides. Always run this when your script completes to free resources and reset detection logic.
For remote Browserbase sessions, use --context-id <id> to load saved cookies and storage states. Add --persist to save changes back to that context when stopping.
Complete Workflow Example
This example demonstrates a full interaction flow using local mode:
# 1. Select clean local environment
browse env local
# 2. Navigate with network idle wait
browse open https://example.com --wait networkidle
# 3. Capture accessibility tree and element refs
browse snapshot
# 4. Click specific element using ref from snapshot
browse click @0-5
# 5. Verify navigation succeeded
browse get title
# 6. Interact with form
browse fill "#search-input" "browser automation"
browse press Enter
# 7. Capture visual confirmation (slower than snapshot)
browse screenshot ./results.png --full-page
# 8. Clean shutdown
browse stop
For protected sites requiring stealth capabilities:
export BROWSERBASE_API_KEY="YOUR_KEY"
browse env remote
# Load existing context to maintain login state
browse open https://protected-site.com \
--context-id ctx_abc123 \
--persist
# Execute interactions
browse fill "#email" "user@example.com"
browse click @0-12
# Verify success
browse get text ".welcome-banner"
browse stop
Working with Output Formats and Machine Reading
Add --json to any command for structured output suitable for piping to other tools:
browse --json snapshot | jq '.elements[0]'
This produces machine-readable JSON containing element references, text content, and accessibility properties.
Summary
- The browse CLI uses a persistent daemon architecture for efficient browser control via
skills/browser/REFERENCE.mdimplementation. - Select environments using
browse env local(isolated Chrome) orbrowse env remote(Browserbase cloud with stealth). - Always start with
browse openand validate state withbrowse snapshotbefore interactions. - Use element references (e.g.,
@0-5) from snapshots for reliable clicking rather than volatile CSS selectors. - Clean up sessions with
browse stopto release resources and reset environment detection.
Frequently Asked Questions
How do I maintain login state between browse CLI sessions?
Use browse env local --auto-connect to attach to an existing Chrome instance where you're already logged in. For remote Browserbase mode, specify --context-id <id> when running browse open, and include --persist to save cookies and storage changes back to that context when you run browse stop.
What is the difference between browse snapshot and browse screenshot?
browse snapshot returns an accessibility tree with element references (@0-5) and text content instantly, making it ideal for scripting and obtaining click targets. browse screenshot renders the actual visual page to a PNG file, which is slower but necessary for visual verification or documentation.
Can I run multiple browse CLI sessions simultaneously?
Yes. Use the --session <name> flag or set the BROWSE_SESSION environment variable to create isolated sessions. Each session maintains its own daemon and browser instance, allowing you to automate multiple sites or parallel workflows without interference.
Why does my browse command fail with "daemon not running"?
The daemon starts automatically on the first command like browse open, but if you receive this error, ensure you haven't manually killed the process or encountered a crash. Run browse stop to clear any stale locks, then retry your command sequence starting with browse open as documented in skills/browser/SKILL.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →