Configuring and Controlling Browser Automation with Playwright and Selenium in MetaGPT
MetaGPT provides a unified browser automation interface that lets you switch between Playwright and Selenium by changing a single configuration value, with both engines exposing an identical async run() API for fetching and parsing web pages.
MetaGPT's browser automation layer abstracts away the differences between Playwright and Selenium, allowing developers to configure and control web scraping tools through a single, consistent interface. By leveraging the BrowserConfig and WebBrowserEngine classes found in the FoundationAgents/MetaGPT repository, you can seamlessly switch between automation backends without rewriting your scraping logic.
Architecture Overview
BrowserConfig
The configuration layer lives in metagpt/configs/browser_config.py. The BrowserConfig class extends YamlModel and defines which automation engine to use and which browser binary to launch:
class BrowserConfig(YamlModel):
engine: WebBrowserEngineType = WebBrowserEngineType.PLAYWRIGHT
browser_type: Literal[
"chromium", "firefox", "webkit", # Playwright
"chrome", "firefox", "edge", "ie" # Selenium
] = "chromium"
engineacceptsWebBrowserEngineType.PLAYWRIGHT,SELENIUM, orCUSTOM.browser_typeis passed directly to the underlying wrapper, letting you choose between Chromium, Firefox, WebKit (Playwright) or Chrome, Edge, IE (Selenium).
WebBrowserEngine
The unified façade is implemented in metagpt/tools/web_browser_engine.py. The WebBrowserEngine class dynamically imports the correct wrapper at runtime based on the engine field:
class WebBrowserEngine(BaseModel):
engine: WebBrowserEngineType = WebBrowserEngineType.PLAYWRIGHT
run_func: Optional[Callable[..., Coroutine[Any, Any, Union[WebPage, List[WebPage]]]]] = None
proxy: Optional[str] = None
@model_validator(mode="after")
def validate_extra(self):
# ... dynamic import logic ...
return self
async def run(self, url: str, *urls: str, per_page_timeout: float = None):
return await self.run_func(url, *urls, per_page_timeout=per_page_timeout)
Key behaviors:
- Dynamic import: Uses
importlib.import_moduleto loadPlaywrightWrapperorSeleniumWrapperonly when needed. - Factory method:
from_browser_config(cls, config: BrowserConfig, **kwargs)builds an instance directly from aBrowserConfigobject. - Unified API: The
run()method accepts one or many URLs and returns either a singleWebPageor aList[WebPage].
Playwright Implementation
PlaywrightWrapper
The async Playwright driver is defined in metagpt/tools/web_browser_engine_playwright.py. The PlaywrightWrapper class handles browser launch, proxy configuration, and automatic installation of browser binaries:
class PlaywrightWrapper(BaseModel):
browser_type: Literal["chromium", "firefox", "webkit"] = "chromium"
launch_kwargs: dict = Field(default_factory=dict)
proxy: Optional[str] = None
context_kwargs: dict = Field(default_factory=dict)
_has_run_precheck: bool = PrivateAttr(False)
async def run(self, url: str, *urls: str, per_page_timeout: float = None) -> WebPage | List[WebPage]:
async with async_playwright() as ap:
browser_type = getattr(ap, self.browser_type)
await self._run_precheck(browser_type)
browser = await browser_type.launch(**self.launch_kwargs)
if urls:
return await asyncio.gather(
self._scrape(browser, url, per_page_timeout),
*(self._scrape(browser, u, per_page_timeout) for u in urls),
)
return await self._scrape(browser, url, per_page_timeout)
Notable features:
- Auto‑installation: The
_run_precheckmethod installs missing browsers usingplaywright installlogic and resolves OS‑specific fallback paths. - Proxy support: Passed through
launch_kwargsto the browser launcher. - Concurrency: Uses
asyncio.gatherto fetch multiple URLs in parallel within the same browser context.
Selenium Implementation
SeleniumWrapper
The synchronous Selenium driver is wrapped in metagpt/tools/web_browser_engine_selenium.py. The SeleniumWrapper bridges Selenium’s blocking API with MetaGPT’s async interface using a thread pool:
class SeleniumWrapper(BaseModel):
browser_type: Literal["chrome", "firefox", "edge", "ie"] = "chrome"
launch_kwargs: dict = Field(default_factory=dict)
proxy: Optional[str] = None
loop: Optional[asyncio.AbstractEventLoop] = None
executor: Optional[futures.Executor] = None
_has_run_precheck: bool = PrivateAttr(False)
_get_driver: Optional[Callable] = PrivateAttr(None)
async def run(self, url: str, *urls: str, per_page_timeout: float = None) -> WebPage | List[WebPage]:
await self._run_precheck()
_scrape = lambda u, t: self.loop.run_in_executor(
self.executor, self._scrape_website, u, t
)
if urls:
return await asyncio.gather(_scrape(url, per_page_timeout), *(_scrape(u, per_page_timeout) for u in urls))
return await _scrape(url, per_page_timeout)
def _scrape_website(self, url, timeout: float = None):
with self._get_driver() as driver:
try:
driver.get(url)
WebDriverWait(driver, timeout or 30).until(
EC.presence_of_element_located((By.TAG_NAME, "body"))
)
inner_text = driver.execute_script("return document.body.innerText;")
html = driver.page_source
except Exception as e:
inner_text = f"Fail to load page content for {e}"
html = ""
return WebPage(inner_text=inner_text, html=html, url=url)
Key implementation details:
- Thread‑pooled execution:
run_in_executorkeeps the public API async while Selenium runs in a background thread. - Automatic driver management:
_run_precheckuseswebdriver_managerto download the correctchromedriver(or geckodriver) on first use. - Proxy support: Configured via
WDMHttpProxyClientand passed to the driver constructor.
WebPage Result Model
Regardless of which engine you choose, both wrappers return a WebPage object defined in metagpt/utils/parse_html.py:
class WebPage(BaseModel):
inner_text: str
html: str
url: str
_soup: Optional[BeautifulSoup] = PrivateAttr(default=None)
_title: Optional[str] = PrivateAttr(default=None)
@property
def soup(self) -> BeautifulSoup:
if self._soup is None:
self._soup = BeautifulSoup(self.html, "html.parser")
return self._soup
@property
def title(self):
if self._title is None:
title_tag = self.soup.find("title")
self._title = title_tag.text.strip() if title_tag else ""
return self._title
def get_links(self) -> Generator[str, None, None]:
for i in self.soup.find_all("a", href=True):
url = i["href"]
result = urlparse(url)
if not result.scheme and result.path:
yield urljoin(self.url, url)
elif url.startswith(("http://", "https://")):
yield urljoin(self.url, url)
The WebPage model provides:
- Lazy parsing:
soupproperty initializes BeautifulSoup only when accessed. - Link extraction:
get_links()yields absolute URLs, resolving relative paths against the base URL. - Content access:
inner_textfor clean text,htmlfor raw markup.
Practical Usage Examples
Basic Playwright Example
import asyncio
from metagpt.configs.browser_config import BrowserConfig, WebBrowserEngineType
from metagpt.tools.web_browser_engine import WebBrowserEngine
async def main():
# Choose Playwright + Chromium, optional proxy
cfg = BrowserConfig(
engine=WebBrowserEngineType.PLAYWRIGHT,
browser_type="chromium",
)
engine = WebBrowserEngine.from_browser_config(cfg, proxy="http://my-proxy:8080")
page = await engine.run("https://example.com")
print("Title:", page.title)
print("First 200 chars of text:", page.inner_text[:200])
# iterate links
for link in page.get_links():
print("→", link)
asyncio.run(main())
The wrapper automatically installs the required Playwright browsers on first run.
Selenium Example (Headless Chrome)
import asyncio
from metagpt.configs.browser_config import BrowserConfig, WebBrowserEngineType
from metagpt.tools.web_browser_engine import WebBrowserEngine
async def main():
cfg = BrowserConfig(
engine=WebBrowserEngineType.SELENIUM,
browser_type="chrome",
)
# Selenium runs in a thread pool; you may provide a custom executor if desired.
engine = WebBrowserEngine.from_browser_config(cfg, proxy="http://my-proxy:8080")
page = await engine.run("https://news.ycombinator.com")
print("Page title:", page.title)
print("Links on the front page:")
for link in page.get_links():
print(link)
asyncio.run(main())
The first call downloads the matching chromedriver via webdriver_manager and then launches Chrome headlessly.
Fetch Multiple URLs in Parallel
import asyncio
from metagpt.configs.browser_config import BrowserConfig, WebBrowserEngineType
from metagpt.tools.web_browser_engine import WebBrowserEngine
async def scrape_many():
cfg = BrowserConfig(engine=WebBrowserEngineType.PLAYWRIGHT, browser_type="firefox")
engine = WebBrowserEngine.from_browser_config(cfg)
urls = [
"https://python.org",
"https://pypi.org",
"https://github.com",
]
pages = await engine.run(urls[0], *urls[1:]) # returns List[WebPage]
for p in pages:
print(p.url, "→", p.title)
asyncio.run(scrape_many())
Both wrappers internally use asyncio.gather (Playwright) or asyncio.gather over executor calls (Selenium) to achieve concurrency.
Key Source Files
| File | Role |
|---|---|
metagpt/configs/browser_config.py |
Defines BrowserConfig and WebBrowserEngineType enumeration |
metagpt/tools/web_browser_engine.py |
Unified façade WebBrowserEngine with dynamic wrapper loading |
metagpt/tools/web_browser_engine_playwright.py |
Async PlaywrightWrapper with auto-install and proxy support |
metagpt/tools/web_browser_engine_selenium.py |
Thread-pooled SeleniumWrapper using webdriver_manager |
metagpt/utils/parse_html.py |
WebPage result model with BeautifulSoup parsing utilities |
Summary
- Unified Configuration:
BrowserConfiginmetagpt/configs/browser_config.pylets you selectPLAYWRIGHTorSELENIUMvia a single enum value. - Runtime Abstraction:
WebBrowserEnginedynamically imports the correct wrapper and exposes a singleasync run()method regardless of backend. - Async-First Design: Playwright runs natively async with
async_playwright(), while Selenium is bridged viarun_in_executorto maintain the same async API. - Automatic Setup: Both wrappers handle dependency management—Playwright via
_run_precheckauto-install, Selenium viawebdriver_managerfor driver binaries. - Structured Results: All engines return
WebPageobjects frommetagpt/utils/parse_html.py, providing lazy BeautifulSoup parsing, title extraction, and absolute link resolution viaget_links().
Frequently Asked Questions
How do I switch from Playwright to Selenium in MetaGPT?
Change the engine field in your BrowserConfig to WebBrowserEngineType.SELENIUM and adjust the browser_type to a Selenium-compatible value like "chrome" or "firefox". The WebBrowserEngine façade automatically loads the SeleniumWrapper and maintains the same async run() interface, so no other code changes are required.
Does MetaGPT handle browser installation automatically?
Yes. When using Playwright, the PlaywrightWrapper runs _run_precheck on first use to install missing browsers via Playwright’s install API. For Selenium, the SeleniumWrapper uses webdriver_manager to download the correct ChromeDriver, GeckoDriver, or EdgeDriver binary on the fly, eliminating manual driver management.
Can I scrape multiple URLs in parallel?
Absolutely. Both PlaywrightWrapper and SeleniumWrapper implement parallel fetching. Pass multiple URLs to engine.run(url1, url2, url3) and the engine returns a List[WebPage]. Playwright uses asyncio.gather over native async contexts, while Selenium achieves concurrency via asyncio.gather over run_in_executor calls across a thread pool.
What data structure does MetaGPT return from web scraping?
All browser engines return a WebPage Pydantic model defined in metagpt/utils/parse_html.py. This object contains inner_text (clean page text), html (raw markup), and url (final URL). It also provides lazy-loaded soup (BeautifulSoup instance), title property, and get_links() generator that yields absolute URLs by resolving relative paths against the base URL.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →