Configuring and Controlling Browser Automation with Playwright and Selenium in MetaGPT

MetaGPT provides a unified browser automation interface that lets you switch between Playwright and Selenium by changing a single configuration value, with both engines exposing an identical async run() API for fetching and parsing web pages.

MetaGPT's browser automation layer abstracts away the differences between Playwright and Selenium, allowing developers to configure and control web scraping tools through a single, consistent interface. By leveraging the BrowserConfig and WebBrowserEngine classes found in the FoundationAgents/MetaGPT repository, you can seamlessly switch between automation backends without rewriting your scraping logic.

Architecture Overview

BrowserConfig

The configuration layer lives in metagpt/configs/browser_config.py. The BrowserConfig class extends YamlModel and defines which automation engine to use and which browser binary to launch:

class BrowserConfig(YamlModel):
    engine: WebBrowserEngineType = WebBrowserEngineType.PLAYWRIGHT
    browser_type: Literal[
        "chromium", "firefox", "webkit",   # Playwright

        "chrome", "firefox", "edge", "ie" # Selenium

    ] = "chromium"
  • engine accepts WebBrowserEngineType.PLAYWRIGHT, SELENIUM, or CUSTOM.
  • browser_type is passed directly to the underlying wrapper, letting you choose between Chromium, Firefox, WebKit (Playwright) or Chrome, Edge, IE (Selenium).

WebBrowserEngine

The unified façade is implemented in metagpt/tools/web_browser_engine.py. The WebBrowserEngine class dynamically imports the correct wrapper at runtime based on the engine field:

class WebBrowserEngine(BaseModel):
    engine: WebBrowserEngineType = WebBrowserEngineType.PLAYWRIGHT
    run_func: Optional[Callable[..., Coroutine[Any, Any, Union[WebPage, List[WebPage]]]]] = None
    proxy: Optional[str] = None

    @model_validator(mode="after")
    def validate_extra(self):
        # ... dynamic import logic ...

        return self

    async def run(self, url: str, *urls: str, per_page_timeout: float = None):
        return await self.run_func(url, *urls, per_page_timeout=per_page_timeout)

Key behaviors:

  • Dynamic import: Uses importlib.import_module to load PlaywrightWrapper or SeleniumWrapper only when needed.
  • Factory method: from_browser_config(cls, config: BrowserConfig, **kwargs) builds an instance directly from a BrowserConfig object.
  • Unified API: The run() method accepts one or many URLs and returns either a single WebPage or a List[WebPage].

Playwright Implementation

PlaywrightWrapper

The async Playwright driver is defined in metagpt/tools/web_browser_engine_playwright.py. The PlaywrightWrapper class handles browser launch, proxy configuration, and automatic installation of browser binaries:

class PlaywrightWrapper(BaseModel):
    browser_type: Literal["chromium", "firefox", "webkit"] = "chromium"
    launch_kwargs: dict = Field(default_factory=dict)
    proxy: Optional[str] = None
    context_kwargs: dict = Field(default_factory=dict)
    _has_run_precheck: bool = PrivateAttr(False)

    async def run(self, url: str, *urls: str, per_page_timeout: float = None) -> WebPage | List[WebPage]:
        async with async_playwright() as ap:
            browser_type = getattr(ap, self.browser_type)
            await self._run_precheck(browser_type)
            browser = await browser_type.launch(**self.launch_kwargs)
            if urls:
                return await asyncio.gather(
                    self._scrape(browser, url, per_page_timeout),
                    *(self._scrape(browser, u, per_page_timeout) for u in urls),
                )
            return await self._scrape(browser, url, per_page_timeout)

Notable features:

  • Auto‑installation: The _run_precheck method installs missing browsers using playwright install logic and resolves OS‑specific fallback paths.
  • Proxy support: Passed through launch_kwargs to the browser launcher.
  • Concurrency: Uses asyncio.gather to fetch multiple URLs in parallel within the same browser context.

Selenium Implementation

SeleniumWrapper

The synchronous Selenium driver is wrapped in metagpt/tools/web_browser_engine_selenium.py. The SeleniumWrapper bridges Selenium’s blocking API with MetaGPT’s async interface using a thread pool:

class SeleniumWrapper(BaseModel):
    browser_type: Literal["chrome", "firefox", "edge", "ie"] = "chrome"
    launch_kwargs: dict = Field(default_factory=dict)
    proxy: Optional[str] = None
    loop: Optional[asyncio.AbstractEventLoop] = None
    executor: Optional[futures.Executor] = None
    _has_run_precheck: bool = PrivateAttr(False)
    _get_driver: Optional[Callable] = PrivateAttr(None)

    async def run(self, url: str, *urls: str, per_page_timeout: float = None) -> WebPage | List[WebPage]:
        await self._run_precheck()
        _scrape = lambda u, t: self.loop.run_in_executor(
            self.executor, self._scrape_website, u, t
        )
        if urls:
            return await asyncio.gather(_scrape(url, per_page_timeout), *(_scrape(u, per_page_timeout) for u in urls))
        return await _scrape(url, per_page_timeout)

    def _scrape_website(self, url, timeout: float = None):
        with self._get_driver() as driver:
            try:
                driver.get(url)
                WebDriverWait(driver, timeout or 30).until(
                    EC.presence_of_element_located((By.TAG_NAME, "body"))
                )
                inner_text = driver.execute_script("return document.body.innerText;")
                html = driver.page_source
            except Exception as e:
                inner_text = f"Fail to load page content for {e}"
                html = ""
            return WebPage(inner_text=inner_text, html=html, url=url)

Key implementation details:

  • Thread‑pooled execution: run_in_executor keeps the public API async while Selenium runs in a background thread.
  • Automatic driver management: _run_precheck uses webdriver_manager to download the correct chromedriver (or geckodriver) on first use.
  • Proxy support: Configured via WDMHttpProxyClient and passed to the driver constructor.

WebPage Result Model

Regardless of which engine you choose, both wrappers return a WebPage object defined in metagpt/utils/parse_html.py:

class WebPage(BaseModel):
    inner_text: str
    html: str
    url: str

    _soup: Optional[BeautifulSoup] = PrivateAttr(default=None)
    _title: Optional[str] = PrivateAttr(default=None)

    @property
    def soup(self) -> BeautifulSoup:
        if self._soup is None:
            self._soup = BeautifulSoup(self.html, "html.parser")
        return self._soup

    @property
    def title(self):
        if self._title is None:
            title_tag = self.soup.find("title")
            self._title = title_tag.text.strip() if title_tag else ""
        return self._title

    def get_links(self) -> Generator[str, None, None]:
        for i in self.soup.find_all("a", href=True):
            url = i["href"]
            result = urlparse(url)
            if not result.scheme and result.path:
                yield urljoin(self.url, url)
            elif url.startswith(("http://", "https://")):
                yield urljoin(self.url, url)

The WebPage model provides:

  • Lazy parsing: soup property initializes BeautifulSoup only when accessed.
  • Link extraction: get_links() yields absolute URLs, resolving relative paths against the base URL.
  • Content access: inner_text for clean text, html for raw markup.

Practical Usage Examples

Basic Playwright Example

import asyncio
from metagpt.configs.browser_config import BrowserConfig, WebBrowserEngineType
from metagpt.tools.web_browser_engine import WebBrowserEngine

async def main():
    # Choose Playwright + Chromium, optional proxy

    cfg = BrowserConfig(
        engine=WebBrowserEngineType.PLAYWRIGHT,
        browser_type="chromium",
    )
    engine = WebBrowserEngine.from_browser_config(cfg, proxy="http://my-proxy:8080")
    page = await engine.run("https://example.com")
    print("Title:", page.title)
    print("First 200 chars of text:", page.inner_text[:200])
    # iterate links

    for link in page.get_links():
        print("→", link)

asyncio.run(main())

The wrapper automatically installs the required Playwright browsers on first run.

Selenium Example (Headless Chrome)

import asyncio
from metagpt.configs.browser_config import BrowserConfig, WebBrowserEngineType
from metagpt.tools.web_browser_engine import WebBrowserEngine

async def main():
    cfg = BrowserConfig(
        engine=WebBrowserEngineType.SELENIUM,
        browser_type="chrome",
    )
    # Selenium runs in a thread pool; you may provide a custom executor if desired.

    engine = WebBrowserEngine.from_browser_config(cfg, proxy="http://my-proxy:8080")
    page = await engine.run("https://news.ycombinator.com")
    print("Page title:", page.title)
    print("Links on the front page:")
    for link in page.get_links():
        print(link)

asyncio.run(main())

The first call downloads the matching chromedriver via webdriver_manager and then launches Chrome headlessly.

Fetch Multiple URLs in Parallel

import asyncio
from metagpt.configs.browser_config import BrowserConfig, WebBrowserEngineType
from metagpt.tools.web_browser_engine import WebBrowserEngine

async def scrape_many():
    cfg = BrowserConfig(engine=WebBrowserEngineType.PLAYWRIGHT, browser_type="firefox")
    engine = WebBrowserEngine.from_browser_config(cfg)

    urls = [
        "https://python.org",
        "https://pypi.org",
        "https://github.com",
    ]
    pages = await engine.run(urls[0], *urls[1:])   # returns List[WebPage]

    for p in pages:
        print(p.url, "→", p.title)

asyncio.run(scrape_many())

Both wrappers internally use asyncio.gather (Playwright) or asyncio.gather over executor calls (Selenium) to achieve concurrency.

Key Source Files

File Role
metagpt/configs/browser_config.py Defines BrowserConfig and WebBrowserEngineType enumeration
metagpt/tools/web_browser_engine.py Unified façade WebBrowserEngine with dynamic wrapper loading
metagpt/tools/web_browser_engine_playwright.py Async PlaywrightWrapper with auto-install and proxy support
metagpt/tools/web_browser_engine_selenium.py Thread-pooled SeleniumWrapper using webdriver_manager
metagpt/utils/parse_html.py WebPage result model with BeautifulSoup parsing utilities

Summary

  • Unified Configuration: BrowserConfig in metagpt/configs/browser_config.py lets you select PLAYWRIGHT or SELENIUM via a single enum value.
  • Runtime Abstraction: WebBrowserEngine dynamically imports the correct wrapper and exposes a single async run() method regardless of backend.
  • Async-First Design: Playwright runs natively async with async_playwright(), while Selenium is bridged via run_in_executor to maintain the same async API.
  • Automatic Setup: Both wrappers handle dependency management—Playwright via _run_precheck auto-install, Selenium via webdriver_manager for driver binaries.
  • Structured Results: All engines return WebPage objects from metagpt/utils/parse_html.py, providing lazy BeautifulSoup parsing, title extraction, and absolute link resolution via get_links().

Frequently Asked Questions

How do I switch from Playwright to Selenium in MetaGPT?

Change the engine field in your BrowserConfig to WebBrowserEngineType.SELENIUM and adjust the browser_type to a Selenium-compatible value like "chrome" or "firefox". The WebBrowserEngine façade automatically loads the SeleniumWrapper and maintains the same async run() interface, so no other code changes are required.

Does MetaGPT handle browser installation automatically?

Yes. When using Playwright, the PlaywrightWrapper runs _run_precheck on first use to install missing browsers via Playwright’s install API. For Selenium, the SeleniumWrapper uses webdriver_manager to download the correct ChromeDriver, GeckoDriver, or EdgeDriver binary on the fly, eliminating manual driver management.

Can I scrape multiple URLs in parallel?

Absolutely. Both PlaywrightWrapper and SeleniumWrapper implement parallel fetching. Pass multiple URLs to engine.run(url1, url2, url3) and the engine returns a List[WebPage]. Playwright uses asyncio.gather over native async contexts, while Selenium achieves concurrency via asyncio.gather over run_in_executor calls across a thread pool.

What data structure does MetaGPT return from web scraping?

All browser engines return a WebPage Pydantic model defined in metagpt/utils/parse_html.py. This object contains inner_text (clean page text), html (raw markup), and url (final URL). It also provides lazy-loaded soup (BeautifulSoup instance), title property, and get_links() generator that yields absolute URLs by resolving relative paths against the base URL.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →