How to Use Browse CLI Commands for Page Interaction: A Complete Guide

The browse CLI is a daemon-based automation tool that lets you control browsers via command-line commands like browse open, browse snapshot, and browse click, operating in either local Chrome or remote Browserbase cloud environments.

The browse command-line interface, maintained in the browserbase/skills repository, provides a persistent daemon architecture for reliable web automation. Unlike one-shot browser tools, it runs a background process that stays alive across multiple commands, enabling deterministic scripting for complex page interactions. This guide covers the essential browse CLI commands for page interaction based on the implementation details found in skills/browser/SKILL.md and skills/browser/REFERENCE.md.

Understanding the Daemon Architecture

The browse CLI operates through a background daemon that manages the browser process independently of your terminal session. According to skills/browser/REFERENCE.md, this daemon starts automatically on the first command (such as browse open) and continues running until you explicitly execute browse stop.

Each CLI command communicates with this daemon over RPC channels, which eliminates the overhead of launching a new browser instance for every action. Because the browser stays resident in memory, you can chain multiple interactions—navigations, clicks, form fills—without re-initializing the environment, resulting in faster and more reliable automation scripts.

Selecting Your Environment: Local vs Remote

Before executing page interactions, you must select between local and remote modes using browse env:

  • browse env local: Runs a clean, isolated Chrome instance on your machine. Use --auto-connect to attach to an existing Chrome session with your login states preserved.
  • browse env remote: Connects to Browserbase cloud browsers with anti-bot stealth, automatic CAPTCHA solving, and residential proxies. Requires BROWSERBASE_API_KEY to be set.

You can switch environments mid-session or use the --session <name> flag to maintain multiple isolated browser contexts simultaneously.

Core Browse CLI Commands for Page Interaction

Start every workflow by opening a target URL. The daemon initializes here if it isn't already running.

browse open https://example.com --wait networkidle

The --wait flag accepts load, domcontentloaded, or networkidle to control when the command returns.

Inspecting Page State

browse snapshot: Returns an accessibility tree containing element references (e.g., @0-5) that map to specific DOM nodes. This is the preferred inspection method for obtaining clickable references.

browse get: Retrieves specific page properties:

browse get title          # Current page title

browse get text           # Extracted text content

browse get html           # Raw HTML source

browse get url            # Current address

Always verify state with browse snapshot or browse get before and after interactions to ensure deterministic behavior.

Interacting with Elements

Once you have element references from browse snapshot, use these commands:

browse click @ref: Clicks the exact element corresponding to the reference (e.g., browse click @0-5). Using refs guarantees element identity won't shift due to DOM changes.

browse fill <selector> <value>: Enters text into input fields using CSS selectors:

browse fill "#email" "user@example.com"
browse fill "#password" "s3cr3t!"

browse type <text> and browse press <key>: Send keyboard input or press specific keys like Enter, Tab, or Escape.

browse scroll ...: Scrolls the page vertically or horizontally to reveal elements before interaction.

Managing Sessions and Contexts

browse stop: Shuts down the daemon, closes the browser, and clears any environment overrides. Always run this when your script completes to free resources and reset detection logic.

For remote Browserbase sessions, use --context-id <id> to load saved cookies and storage states. Add --persist to save changes back to that context when stopping.

Complete Workflow Example

This example demonstrates a full interaction flow using local mode:


# 1. Select clean local environment

browse env local

# 2. Navigate with network idle wait

browse open https://example.com --wait networkidle

# 3. Capture accessibility tree and element refs

browse snapshot

# 4. Click specific element using ref from snapshot

browse click @0-5

# 5. Verify navigation succeeded

browse get title

# 6. Interact with form

browse fill "#search-input" "browser automation"
browse press Enter

# 7. Capture visual confirmation (slower than snapshot)

browse screenshot ./results.png --full-page

# 8. Clean shutdown

browse stop

For protected sites requiring stealth capabilities:

export BROWSERBASE_API_KEY="YOUR_KEY"
browse env remote

# Load existing context to maintain login state

browse open https://protected-site.com \
      --context-id ctx_abc123 \
      --persist

# Execute interactions

browse fill "#email" "user@example.com"
browse click @0-12

# Verify success

browse get text ".welcome-banner"

browse stop

Working with Output Formats and Machine Reading

Add --json to any command for structured output suitable for piping to other tools:

browse --json snapshot | jq '.elements[0]'

This produces machine-readable JSON containing element references, text content, and accessibility properties.

Summary

  • The browse CLI uses a persistent daemon architecture for efficient browser control via skills/browser/REFERENCE.md implementation.
  • Select environments using browse env local (isolated Chrome) or browse env remote (Browserbase cloud with stealth).
  • Always start with browse open and validate state with browse snapshot before interactions.
  • Use element references (e.g., @0-5) from snapshots for reliable clicking rather than volatile CSS selectors.
  • Clean up sessions with browse stop to release resources and reset environment detection.

Frequently Asked Questions

How do I maintain login state between browse CLI sessions?

Use browse env local --auto-connect to attach to an existing Chrome instance where you're already logged in. For remote Browserbase mode, specify --context-id <id> when running browse open, and include --persist to save cookies and storage changes back to that context when you run browse stop.

What is the difference between browse snapshot and browse screenshot?

browse snapshot returns an accessibility tree with element references (@0-5) and text content instantly, making it ideal for scripting and obtaining click targets. browse screenshot renders the actual visual page to a PNG file, which is slower but necessary for visual verification or documentation.

Can I run multiple browse CLI sessions simultaneously?

Yes. Use the --session <name> flag or set the BROWSE_SESSION environment variable to create isolated sessions. Each session maintains its own daemon and browser instance, allowing you to automate multiple sites or parallel workflows without interference.

Why does my browse command fail with "daemon not running"?

The daemon starts automatically on the first command like browse open, but if you receive this error, ensure you haven't manually killed the process or encountered a crash. Run browse stop to clear any stale locks, then retry your command sequence starting with browse open as documented in skills/browser/SKILL.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →