How to Perform Diff-Driven Testing for Pull Requests with the ui-test Skill

The ui-test skill performs diff-driven testing for pull requests by analyzing git diffs to identify changed UI components, launching an isolated Chromium instance to execute adversarial checks only on modified elements, and generating a structured HTML report for reviewers.

The ui-test skill in the browserbase/skills repository provides an agentic approach to UI validation that focuses exclusively on changes introduced in a pull request. Unlike traditional end-to-end testing that validates entire applications, this skill implements Workflow A (diff-driven testing) to reduce runtime and provide feedback tightly coupled to the PR's intent. By leveraging the browse CLI and real browser environments, it transforms how teams validate UI modifications before merging.

How Diff-Driven Testing Works

According to skills/ui-test/README.md in the browserbase/skills repository, the ui-test skill implements a focused validation pipeline that avoids testing unchanged application surfaces. The process follows Workflow A, which contrasts with full-exploratory (Workflow B) and parallel-session (Workflow C) modes.

Analyzing the Git Diff

The skill begins by ingesting the PR's git diff or a supplied diff URL to discover which UI components, routes, or assets were modified. This analysis determines the scope of testing, ensuring the agent only interacts with changed elements rather than performing comprehensive application exploration.

Launching Isolated Browser Sessions

For each identified change, the skill spins up a Chromium instance via the browse CLI. By default, it uses browse env local for local execution, but automatically switches to Browserbase cloud infrastructure when a BROWSERBASE_API_KEY environment variable is detected.

Executing Adversarial Checks

The skill runs deterministic validation checks against each modified UI element. As documented in skills/ui-test/references/adversarial-patterns.md, these checks include:

  • axe-core accessibility validation
  • Screenshot diff comparisons
  • Console error detection
  • XSS injection tests
  • Rapid-click stress tests

Structured Output Generation

Every test step outputs a machine-readable line following the protocol defined in skills/ui-test/SKILL.md:

  • STEP_PASS|id|evidence for successful validations
  • STEP_FAIL|id|expected → actual for failed assertions

These results feed into skills/ui-test/references/report-template.html to produce a single HTML file suitable for attachment to pull requests.

Running Diff-Driven Tests on Pull Requests

The ui-test skill exposes a CLI interface that integrates directly with GitHub PR workflows. Execution requires the browse CLI and optionally the Browserbase API key for cloud-based testing.

Installation

Install the skill globally using the skills CLI:

npx skills add browserbase/ui-test

This registers the ui-test command for local invocation.

Testing PR Changes

Execute diff-driven testing against a specific pull request by passing the diff URL:

skills run browserbase/ui-test "Test the UI changes in my PR" \
  --diff-url https://api.github.com/repos/<owner>/<repo>/pulls/<num>/files

After installation, use the shorthand command:

ui-test "Test the UI changes in my PR"

Local Execution Without API Keys

Run tests entirely locally without Browserbase cloud resources:


# Install the browse CLI

npm install -g @browserbasehq/browse-cli

# Initialize local environment

browse env local

# Execute diff-driven tests

ui-test "Test the UI changes in my PR"

This mode creates isolated Chromium instances on your local machine and outputs ui-test-report.html in the current directory.

Parallel Execution for Large PRs

For pull requests modifying numerous components, distribute testing across parallel workers using named sessions:

export BROWSE_SESSION=pr-1234-worker-1
ui-test "Test the UI changes in my PR" &

export BROWSE_SESSION=pr-1234-worker-2
ui-test "Test the UI changes in my PR" &

Each worker requires a unique BROWSE_SESSION value to prevent collision, as detailed in skills/ui-test/references/parallel-testing.md.

Key Source Files and References

The ui-test skill implementation spans several configuration and documentation files within the browserbase/skills repository:

Summary

  • The ui-test skill performs diff-driven testing for pull requests by analyzing git diff output to identify modified UI components, testing only what changed rather than the entire application.
  • It launches Chromium instances via the browse CLI (locally or on Browserbase cloud) to execute adversarial checks including accessibility, screenshot comparison, and XSS injection tests.
  • Results follow a structured protocol (STEP_PASS|id|evidence or STEP_FAIL|id|expected → actual) and compile into an HTML report using skills/ui-test/references/report-template.html.
  • Workflow A (diff-driven) minimizes runtime and provides PR-specific feedback, while Workflow C supports parallel execution for large changes via BROWSE_SESSION environment variables.
  • Installation requires npx skills add browserbase/ui-test, with execution via ui-test "<prompt>" --diff-url <url> or local invocation through browse env local.

Frequently Asked Questions

How does ui-test determine which UI elements to test?

The skill examines the PR's git diff—either from a GitHub API URL via --diff-url or supplied diff content—to identify modified components, routes, and assets. According to skills/ui-test/README.md, this Workflow A approach restricts testing to changed surfaces only, avoiding unnecessary validation of unchanged code.

Can I run ui-test without a Browserbase API key?

Yes. When no BROWSERBASE_API_KEY environment variable is present, the skill automatically uses browse env local to launch Chromium on your machine. Install the browse CLI with npm install -g @browserbasehq/browse-cli and ensure you run browse env local before executing ui-test for fully local testing.

What is the difference between Workflow A, B, and C?

Workflow A (diff-driven) tests only modified UI elements based on git diff analysis. Workflow B performs full exploratory testing across the entire application surface. Workflow C enables parallel execution using named BROWSE_SESSION environments to distribute tests across multiple workers, as documented in skills/ui-test/references/parallel-testing.md.

How do I interpret the STEP_PASS and STEP_FAIL output lines?

These structured log lines follow the protocol defined in skills/ui-test/SKILL.md. STEP_PASS|id|evidence indicates successful validation with supporting context, while STEP_FAIL|id|expected → actual documents assertion failures with specific mismatch details. The skill aggregates these into a human-readable HTML report using skills/ui-test/references/report-template.html.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →