How to Leverage Browser Automation in Strix for XSS and CSRF Testing

Strix exposes a Playwright-powered headless Chrome through the browser_action tool, which automatically routes all traffic through the built-in Caido proxy to enable fully automated XSS and CSRF testing without manual session or cookie management.

Strix is an open-source security testing framework that embeds browser automation directly into its agent architecture. The browser_action tool provides LLM-driven agents with a programmatic interface to a real Chrome instance, making it possible to perform complex cross-site scripting (XSS) and cross-site request forgery (CSRF) exploits while capturing every network request via the integrated Caido proxy. This architecture allows agents to verify vulnerabilities by inspecting both the DOM state and the underlying HTTP traffic in a single workflow.

Understanding Strix’s Browser Automation Architecture

Strix’s browser automation is implemented as a three-layer stack that abstracts Playwright operations into agent-friendly actions.

Core Components

The system relies on three primary components defined in the strix/tools/browser/ directory:

  • BrowserInstance (browser_instance.py) – The low-level driver that creates a shared Playwright browser, manages individual pages (tabs), and captures screenshots, console logs, and page source.
  • BrowserTabManager (tab_manager.py) – Maps each Strix agent (identified via get_current_agent_id) to its own BrowserInstance and implements high-level actions like navigate, click, type, and execute_js.
  • browser_action (browser_actions.py) – The registered tool that agents invoke. It validates arguments using helpers like _validate_url and _validate_coordinate, then routes calls to private handlers before normalizing the response.

The Request Flow

When an agent calls browser_action, the request flows through six distinct stages:

  1. Agent Request – The LLM emits a function call such as browser_action(action="goto", url="https://target/app").
  2. Tool Dispatch – Strix’s tool registry (strix/tools/registry.py) receives the call and invokes the browser_action function.
  3. Validation – Argument validators ensure mandatory parameters are present and well-formed.
  4. Action Routing – The call is dispatched to one of four private handlers: _handle_navigation_actions, _handle_interaction_actions, _handle_tab_actions, or _handle_utility_actions.
  5. Tab Management – The handler forwards the request to BrowserTabManager, which retrieves the caller’s BrowserInstance or creates a new one on launch.
  6. Execution – The concrete Playwright operation runs in an event-loop thread (_run_async in BrowserInstance), returning a payload containing the screenshot, current URL, page title, tab IDs, and—when relevant—console logs, JS results, or page source.

XSS Testing with Strix Browser Automation

The browser automation enables automated XSS detection by combining DOM manipulation with network request inspection.

Step-by-Step XSS Exploitation

  1. Navigate – Use goto or launch to load the vulnerable page.
  2. Inject Payload – Call type to enter data into input fields, or use execute_js to directly manipulate the DOM.
  3. Trigger Execution – Use click to submit forms or execute_js to fire JavaScript events.
  4. Verify Reflection – Inspect the returned console_logs and Caido proxy logs to confirm the payload was reflected in HTTP requests or responses.

Because the proxy logs are included in the tool’s return payload, agents can programmatically determine if an XSS vulnerability exists. According to the source code in browser_actions.py, the _handle_utility_actions method exposes console logs via get_console_logs, which aggregates data from BrowserInstance._get_console_logs and BrowserInstance._setup_console_logging.

Practical Example: Basic XSS Injection


# Launch browser and navigate to target

agent.browser_action(action="launch", url="https://demo.vulnapp.com/search")

# Inject XSS payload into search field

xss_payload = '<script>alert("XSS")</script>'
agent.browser_action(action="type", text=xss_payload)

# Trigger the search

agent.browser_action(action="click", coordinate="200,350")

# Retrieve console logs to check for payload reflection

result = agent.browser_action(action="get_console_logs", clear=True)
for log in result.get("console_logs", []):
    if "XSS" in log.get("text", ""):
        print("Potential XSS reflected:", log["text"])

CSRF Testing with Strix Browser Automation

Browser automation in Strix is particularly effective for CSRF testing because cookies and session state are automatically shared between the browser and the Caido proxy, eliminating manual cookie handling.

CSRF Token Bypass Workflow

  1. Authenticate – Navigate to the login page, use type to enter credentials, and click to submit. The proxy (BrowserInstance._setup_console_logging) records the authentication request and stores session cookies.
  2. Open Target Tab – Use new_tab to open a CSRF-protected endpoint. The BrowserTabManager.new_tab method ensures the new tab shares the same cookie jar.
  3. Craft Malicious Request – Use type and click to fill hidden forms, or execute JavaScript via execute_js to perform fetch calls without CSRF tokens.
  4. Capture Evidence – The proxy logs contain the raw HTTP request, allowing the agent to verify that the CSRF token was omitted or that the request succeeded despite missing protections.

Practical Example: CSRF Token Reuse


# Step 1: Log in to establish session

agent.browser_action(action="launch", url="https://secure.app/login")
agent.browser_action(action="type", text="alice@example.com")
agent.browser_action(action="press_key", key="Tab")
agent.browser_action(action="type", text="SuperSecret123")
agent.browser_action(action="click", coordinate="250,420")  # Submit button

# Step 2: Open CSRF-protected endpoint in new tab

new_tab = agent.browser_action(action="new_tab", url="https://secure.app/transfer")
tab_id = new_tab["tab_id"]

# Step 3: Execute JavaScript to perform unauthorized POST

js = """
fetch('/transfer', {
  method: 'POST',
  credentials: 'include',
  headers: {'Content-Type': 'application/json'},
  body: JSON.stringify({to:'bob', amount:1000})
});
"""
agent.browser_action(action="execute_js", js_code=js, tab_id=tab_id)

# Step 4: Verify CSRF bypass in proxy logs

logs = agent.browser_action(action="get_console_logs", clear=True, tab_id=tab_id)
for entry in logs.get("console_logs", []):
    if entry.get("type") == "response" and "/transfer" in entry.get("text", ""):
        print("CSRF request captured:", entry["text"])

The Role of the Built-In Caido Proxy

Strix’s browser automation does not operate in isolation; it is tightly coupled with the Caido proxy to provide network-level visibility.

Automatic Request Capture

Every HTTP request generated by the Playwright instance is intercepted by the proxy configuration established in BrowserInstance._setup_console_logging. This gives agents access to:

  • Request headers and bodies – Essential for detecting XSS payloads reflected in JSON API responses.
  • Response data – Allows verification of CSRF protections by examining returned status codes and error messages.
  • Cookie propagation – Session tokens are automatically shared across tabs and requests, enabling realistic multi-step CSRF exploitation.

Advanced Automation Patterns

Full-Cycle XSS Detection

For automated scanning, agents can implement a closed-loop verification function that injects, triggers, and confirms XSS without human intervention:

def test_xss(agent, target_url, payload):
    # Load target page

    agent.browser_action(action="launch", url=target_url)
    
    # Bypass complex DOM selectors by injecting via JavaScript

    injection_js = f'document.body.innerHTML += `{payload}`;'
    agent.browser_action(action="execute_js", js_code=injection_js)
    
    # Trigger page reload to check for persistent XSS

    agent.browser_action(action="refresh")
    
    # Check console logs for execution evidence

    logs = agent.browser_action(action="get_console_logs", clear=True)
    for entry in logs.get("console_logs", []):
        if entry.get("type") == "log" and "XSS" in entry.get("text", ""):
            return True
    return False

# Execute test

if test_xss(agent, "https://demo.vulnapp.com/comment", '<script>alert("XSS")</script>'):
    print("XSS vulnerability confirmed")

Summary

  • Three-layer architecture – Strix separates concerns into BrowserInstance (driver), BrowserTabManager (state), and browser_action (agent API), defined in strix/tools/browser/browser_instance.py, tab_manager.py, and browser_actions.py respectively.
  • Integrated proxy visibility – All browser traffic routes through the Caido proxy automatically, providing agents with HTTP request/response data essential for confirming XSS and CSRF vulnerabilities.
  • Session continuity – The BrowserTabManager maintains per-agent browser instances, ensuring cookies and authentication state persist across navigation, tab creation (new_tab), and JavaScript execution (execute_js).
  • Programmatic verification – Agents can validate vulnerabilities by analyzing screenshots, console logs, and proxy data returned by browser_action calls without manual intervention.

Frequently Asked Questions

How does Strix handle session state during CSRF testing?

Strix maintains session continuity through the BrowserTabManager, which maps each agent to a dedicated BrowserInstance. Because the browser runs in the same process as the Caido proxy, cookies are automatically shared across tabs and requests. When an agent calls new_tab to open a CSRF-protected endpoint, the new tab inherits the authenticated session established in the original tab, enabling realistic exploitation scenarios without manual cookie extraction or Set-Cookie parsing.

Can Strix detect XSS payloads that are only reflected in HTTP responses?

Yes. While the browser_action tool captures screenshots and console logs, it also returns data from the integrated Caido proxy that records every HTTP request and response. Agents can inspect these proxy logs (accessible via the tool’s return payload) to detect XSS payloads reflected in JSON responses, custom headers, or other non-DOM contexts that would not appear in console logs or screenshots alone.

What is the difference between using type and execute_js for payload injection?

The type action simulates user keyboard input into focused input fields, which is useful for testing stored XSS through form submissions or DOM-based XSS via input event handlers. The execute_js action, implemented in BrowserInstance._execute_js, allows direct JavaScript execution to manipulate the DOM, bypass complex input validation, or trigger fetch requests for CSRF testing. Use type for realistic user simulation and execute_js when you need to bypass frontend restrictions or perform actions not exposed through UI elements.

Where is the browser tool documentation located in the repository?

User-facing documentation for the browser tool is located in docs/tools/browser.mdx, while the UI rendering components (including emoji formatting and colorization for CLI output) are implemented in strix/interface/tool_components/browser_renderer.py. The core tool logic resides in strix/tools/browser/browser_actions.py, which defines the browser_action function and its four private handler methods (_handle_navigation_actions, _handle_interaction_actions, _handle_tab_actions, _handle_utility_actions).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →