How to Perform Headless Browser Automation with Probe’s Browser Plugin

Probe’s browser plugin enables headless Chrome automation through YAML workflows using the chromedp library, supporting navigation, element interaction, screenshots, and data extraction without displaying a UI.

The linyows/probe repository provides a powerful browser automation plugin that integrates with the Chrome DevTools Protocol (CDP) via chromedp. This plugin allows you to drive Chrome instances programmatically through declarative YAML configurations, making it ideal for CI/CD pipelines, automated testing, and web scraping tasks that require JavaScript execution.

Understanding the Browser Plugin Architecture

The browser plugin follows a clean architecture that separates the core Chrome automation logic from the plugin wrapper and CLI integration.

Core Implementation in browser/client.go

The primary implementation resides in browser/client.go, which defines the Req struct to hold execution parameters including the headless flag, window dimensions, timeout duration, and a slice of ChromeDPActions. The file implements a BrowserRunner interface that abstracts the actual Chrome driver, with ChromeDPRunner handling real chromedp.Run calls and MockRunner provided for unit testing.

The createBrowserContext function builds a ChromeDP ExecAllocator with security-conscious defaults: disabling GPU acceleration, running with --no-sandbox for containerized environments, setting configurable window sizes, and applying the headless flag. Each action receives an individual timeout context to prevent hanging operations.

Plugin Wrapper and CLI Integration

The actions/browser/main.go file serves as the HashiCorp-compatible plugin wrapper. It creates logging callbacks (WithInBrowser, WithBefore, WithAfter) to provide visibility into Chrome operations and forwards requests to br.Request. The Probe CLI (cmd/probe/main.go) registers this plugin under the browser name, spawning the plugin server via gRPC when workflows contain uses: browser steps.

Configuring Headless Browser Automation Workflows

Probe workflows use declarative YAML to define browser automation sequences. The headless: true parameter (default) ensures Chrome runs without a visible window, making it suitable for server environments.

Basic Navigation and Text Extraction

The following workflow demonstrates navigating to a webpage and extracting heading text:

name: Browser Headless Example
jobs:
- name: Get Heading
  steps:
  - name: Grab heading from example.com
    uses: browser
    with:
      headless: true
      timeout: 30s
      actions:
      - name: navigate
        url: https://example.com/
      - name: wait_visible
        selector: h1
      - name: text
        id: heading
        selector: "h1"
    test: res.code == 0 && res.results.heading != ""
    echo: |
      Heading: {{res.results.heading}}

Execute this workflow with: probe examples/browser.yml

The actions array supports multiple ChromeDP operations executed sequentially. Results are stored in res.results using the action's id as the key (or the action name if no ID is specified).

Capturing Screenshots in Headless Mode

Screenshot capture works in both headless and visible modes. The capture_screenshot action writes binary data to temporary files while optionally preserving user-specified paths:

- name: Capture screenshot
  steps:
  - name: Google screenshot
    uses: browser
    with:
      headless: false
      window_w: 1280
      window_h: 800
      actions:
      - name: navigate
        url: https://www.google.com/
      - name: capture_screenshot
        path: ./google.png
        quality: 80
    test: res.code == 0

The res.filepaths map contains the temporary file location for the screenshot, while ./google.png receives a permanent copy when the path parameter is provided.

Advanced Usage: Direct Go Integration

For applications requiring programmatic control, Probe's browser package exposes the same functionality through a Go API:

package main

import (
    "fmt"
    br "github.com/linyows/probe/browser"
)

func main() {
    data := map[string]any{
        "headless": true,
        "timeout":  "45s",
        "actions": []any{
            map[string]any{
                "name": "navigate",
                "url":  "https://golang.org/",
            },
            map[string]any{
                "name": "text",
                "id":   "title",
                "selector": "h1",
            },
        },
    }

    within := br.WithInBrowser(func(s string, i ...any) {
        fmt.Printf("[chromedp] %s\n", fmt.Sprintf(s, i...))
    })
    before := br.WithBefore(func(req *br.Req) {
        fmt.Println("Request prepared:", req)
    })
    after := br.WithAfter(func(res *br.Res) {
        fmt.Println("Response:", res)
    })

    result, err := br.Request(data, within, before, after)
    if err != nil {
        panic(err)
    }
    fmt.Println("Results:", result["results"])
}

This approach utilizes br.Request with optional callbacks for debugging (WithInBrowser), request inspection (WithBefore), and response handling (WithAfter), providing the same capabilities as the YAML workflow interface.

Summary

  • Probe's browser plugin provides headless Chrome automation through the chromedp library, implemented primarily in browser/client.go.
  • Workflow configuration uses declarative YAML with uses: browser, supporting actions like navigate, text, click, send_keys, and capture_screenshot.
  • Headless mode is controlled via the headless: true parameter (default), with window_w and window_h configuring viewport dimensions.
  • Results are stored in res.results (text extraction) and res.filepaths (screenshot temporary files), with optional permanent file paths for screenshots.
  • Go API allows direct programmatic access through br.Request with debugging callbacks for advanced integration scenarios.

Frequently Asked Questions

What Chrome actions does Probe's browser plugin support?

Probe supports navigation (navigate), waiting for elements (wait_visible), text extraction (text), clicking (click), keyboard input (send_keys), and screenshot capture (capture_screenshot). Each action maps to a corresponding chromedp.Action in browser/client.go through the buildActionTasks function. Unsupported action names return an error during workflow parsing.

How does Probe handle screenshot file storage?

When using the capture_screenshot action, Probe writes binary image data to a temporary file using probe.SaveBinaryToTempFile and stores the path in res.filepaths keyed by the action ID. If the action includes a path parameter, Probe additionally writes the file to that user-specified location for backward compatibility and permanent storage.

Can I run Probe browser automation in visible mode for debugging?

Yes, set headless: false in the step's with map to launch a visible Chrome window. This is useful for debugging complex interactions or visual verification. The window_w and window_h parameters control the viewport size in both headless and visible modes. You can also use Jinja-style expressions to toggle headless mode dynamically based on environment variables.

How do I extract specific data from web pages using Probe?

Use the text action with a CSS selector to extract element content. Assign an id to the action for a custom result key, or omit it to use the action name as the key. The extracted value appears in res.results accessible via templating (e.g., {{res.results.heading}}). Combine with wait_visible to ensure elements are rendered before extraction, preventing race conditions in dynamic web applications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →