How to Use MCP for Browser Automation: Server Setup and Tool Examples

The Model Context Protocol (MCP) enables LLMs to automate browsers by sending JSON tool calls to specialized MCP servers that wrap headless Chrome, Playwright, or Selenium, returning structured data like HTML, screenshots, or extracted content without writing browser control code.

The Model Context Protocol (MCP) standardizes how AI agents interact with external tools, making browser automation accessible through simple JSON payloads. According to the punkpeye/awesome-mcp-servers repository, dozens of ready-to-use MCP servers expose browser capabilities ranging from headless scraping to real-user Chrome extensions. This guide explains how to leverage MCP for browser automation using actual server implementations documented in the README.md Browser Automation section (lines 408–450).

MCP Architecture for Browser Automation

The MCP browser automation stack consists of four distinct layers that handle request routing, transport, and execution.

The MCP Client and Agent

The MCP client generates tool-call JSON payloads and manages communication with servers. Clients can be Claude Desktop, Cursor IDE, or custom scripts. When an agent decides to scrape a dynamic page, it emits a structured request such as {"tool":"scrape","args":{"url":"https://example.com"}}. The client packages this call and transmits it via the selected transport protocol.

The MCP Server Layer

MCP servers receive tool-call JSON, execute the corresponding browser operation, and return structured responses. In punkpeye/awesome-mcp-servers, the Browser Automation category lists implementations including browserless-mcp, playwright-mcp, selenium-mcp, and opentabs. Each server exposes a specific set of tools—such as scrape, click, type, or screenshot—that map to underlying browser APIs.

Transport Protocols

MCP supports two primary transports for carrying JSON payloads:

  • STDIO: Line-delimited JSON over standard input/output, ideal for local processes
  • Streamable HTTP: POST requests to endpoints like https://glama.ai/mcp/servers/<owner>/<repo>/tools/<tool>, used by cloud-hosted servers listed in the repository badges (e.g., line 428 for browserless-mcp)

Browser Backend Execution

The server launches and controls the actual browser instance. Backends include:

  • Headless browsers: Playwright (Node/Python), Selenium WebDriver, or Chrome DevTools Protocol (CDP)
  • Real-user browsers: Chrome extensions like opentabs (documented around line 208 in README.md) that drive authenticated user sessions
  • Lightweight binaries: Rust implementations such as browser-use-rs (line 422) for zero-dependency execution

Selecting an MCP Browser Server

The awesome-mcp-servers repository catalogs over 70 browser-automation servers. Choose based on your specific requirements:

  • Cloud-hosted headless: browserless-mcp offers pay-per-call access to headless Chrome with smart scraping APIs, accessible via the Glama gateway at https://glama.ai/mcp/servers/browserless/browserless-mcp (reference line 428)
  • Residential proxy support: aethyn-browser-mcp combines Playwright with proxy rotation for geo-targeted browsing (line 415)
  • Real authenticated sessions: opentabs uses a Chrome extension to drive the user's actual browser, preserving cookies and login state
  • Rust-native performance: browser-use-rs provides a small binary with fast startup and no Node.js dependency
  • Python ecosystem: selenium-mcp wraps Selenium WebDriver for integration with existing Python test suites (line 486)
  • Anti-bot protection: Ceki-me/mcp-server provides real-human Chrome instances with residential IPs and per-minute billing (line 433)

Installing and Configuring a Server

Most Node-based MCP servers follow a consistent installation pattern. For browserless-mcp:


# Install and start the server

npx -y browserless-mcp

The server prints its listening address (e.g., http://127.0.0.1:3000 for local STDIO/HTTP or the Glama gateway URL). Once running, discover available tools using the list meta-tool:

curl -X POST https://glama.ai/mcp/servers/browserless/browserless-mcp/tools/list \
     -H "Content-Type: application/json" \
     -d '{}'

The response includes all tool names and their JSON argument schemas, allowing the LLM to discover capabilities at runtime without hard-coding implementation details.

Executing Browser Automation Tasks

Scraping Dynamic Content

To fetch rendered HTML from JavaScript-heavy pages, use the scrape tool with a selector wait condition:

curl -X POST https://glama.ai/mcp/servers/browserless/browserless-mcp/tools/scrape \
     -H "Content-Type: application/json" \
     -d '{"url":"https://news.ycombinator.com/","wait_for_selector":".itemlist"}'

The server launches a headless browser, waits for the selector, and returns structured JSON:

{
  "status": "success",
  "content": "<html>...</html>",
  "metadata": {
    "url": "https://news.ycombinator.com/",
    "duration_ms": 842
  }
}

Interacting with Page Elements

For authenticated workflows or form submission, chain interaction tools. Using a real-browser server like opentabs:

curl -X POST https://glama.ai/mcp/servers/opentabs/opentabs/tools/click \
     -H "Content-Type: application/json" \
     -d '{"selector":"#login-button","page_id":"user-1234"}'

Follow up with a scrape call using the same page_id to capture the post-login state.

Capturing Screenshots

Playwright-based servers expose a screenshot tool that returns base64-encoded images:

curl -X POST https://glama.ai/mcp/servers/automatalabs/mcp-server-playwright/tools/screenshot \
     -H "Content-Type: application/json" \
     -d '{"url":"https://example.com","full_page":true}'

Decode the screenshot_base64 response field to render images in chat UIs or save them to disk.

Building a Complete Agent

The following Node.js script demonstrates discovery, scraping, and screenshot automation using the Glama HTTP transport:

const axios = require('axios');

async function run() {
  const base = 'https://glama.ai/mcp/servers/browserless/browserless-mcp/tools';

  // Discover available tools
  const list = await axios.post(`${base}/list`, {});
  console.log('Available tools:', list.data.tools.map(t => t.name));

  // Scrape dynamic content
  const scrape = await axios.post(`${base}/scrape`, {
    url: 'https://news.ycombinator.com/',
    wait_for_selector: '.itemlist'
  });
  console.log('Scraped HTML length:', scrape.data.content.length);

  // Capture full-page screenshot
  const shot = await axios.post(`${base}/screenshot`, {
    url: 'https://news.ycombinator.com/',
    full_page: true
  });
  require('fs').writeFileSync('hn.png',
    Buffer.from(shot.data.screenshot_base64, 'base64'));
  console.log('Saved screenshot as hn.png');
}

run().catch(console.error);

This script executes the full MCP flow: the agent discovers tools, invokes browser automation, and consumes structured JSON responses without managing browser binaries directly.

Summary

  • MCP standardizes browser automation by exposing headless and real-browser capabilities through JSON tool calls
  • Transport flexibility allows local STDIO connections or remote HTTP gateways like Glama for cloud-hosted servers
  • Server selection depends on your needs: browserless-mcp for pay-per-call scraping, opentabs for authenticated sessions, or browser-use-rs for lightweight Rust binaries
  • Tool discovery happens at runtime via the list endpoint, enabling LLMs to dynamically understand available actions
  • Structured responses (HTML, screenshots, metadata) integrate directly into agent reasoning loops

Frequently Asked Questions

What is the Model Context Protocol (MCP) in browser automation?

MCP is an open-source protocol that standardizes how LLMs interact with external tools through JSON request/response formats. For browser automation, MCP servers wrap headless browsers or real Chrome instances, exposing operations like navigation, clicking, and extraction as discrete tools that agents can invoke without writing browser-specific code.

How do I choose between STDIO and HTTP transports for MCP browser servers?

Use STDIO (line-delimited JSON over standard input/output) when running the MCP server as a local subprocess on the same machine as your agent, typical for development or desktop applications like Claude Desktop. Use streamable HTTP when connecting to cloud-hosted servers (such as those listed on Glama) or when deploying across network boundaries, as it supports stateless request/response patterns over standard HTTP POST requests.

Can MCP browser servers handle authenticated websites and cookies?

Yes, depending on the server implementation. Servers like opentabs (documented in README.md around line 208) drive real Chrome extensions that leverage the user's existing authenticated session, cookies, and local storage. For headless servers, you can often pass session tokens or credentials via tool arguments, though persistent authentication typically requires servers that support stateful page_id parameters to maintain context across multiple tool calls.

Do I need to install Chrome or Playwright separately to use MCP for browser automation?

It depends on the server. Cloud-hosted options like browserless-mcp require no local browser installation—they run headless Chrome on remote infrastructure. Local servers such as playwright-mcp or selenium-mcp require the corresponding browser binaries or WebDrivers installed on your system, while browser-use-rs bundles its own lightweight browser engine as a standalone Rust binary with zero external dependencies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →