How to Use Playwright CDP Mode with MediaCrawler: A Complete Guide
MediaCrawler provides a thin wrapper around Playwright that enables Chrome DevTools Protocol (CDP) access through the CdpBrowser class in tools/cdp_browser.py, allowing you to send raw CDP commands like Network.enable while retaining high-level Playwright APIs.
MediaCrawler, an open-source scraping framework, ships with built-in support for Playwright's Chrome DevTools Protocol (CDP) mode to give developers low-level browser control. By leveraging Playwright CDP mode with MediaCrawler, you can intercept network traffic, manipulate performance metrics, and access CDP-only APIs that standard Playwright methods don't expose. This integration bridges the gap between high-level browser automation and raw protocol-level interactions.
What is CDP Mode and Why Use It?
The Chrome DevTools Protocol (CDP) provides granular access to Chromium's internals, enabling operations like network interception, JavaScript debugging, and performance monitoring. While standard Playwright offers robust automation APIs, CDP mode unlocks advanced capabilities such as blocking specific resource types, emulating network conditions, or extracting browser metrics directly from the protocol layer.
How MediaCrawler Implements CDP Support
Core Architecture in tools/cdp_browser.py
The primary implementation resides in tools/cdp_browser.py, which creates a Playwright BrowserContext with CDP capabilities enabled. According to the MediaCrawler source code, this module exposes helper methods including send_cdp_cmd and wait_for_cdp_event to simplify protocol interactions without manual JSON serialization.
The launcher abstraction in tools/browser_launcher.py handles the initialization logic, automatically injecting --remote-debugging-port=0 into the launch arguments when CDP mode is requested. This allows Playwright to establish a debug session on an available port while maintaining the standard browser context.
The CDP Session Lifecycle
When you instantiate a CdpBrowser through the public API (as demonstrated in tests/test_cdp_browser.py), the following sequence occurs:
-
Playwright Launch: A headless Chromium instance starts with
args=["--remote-debugging-port=0"]andheadless=True. -
Session Creation: The context calls
await context.newCDPSession(page)to return aCDPSessionobject attached to the specific page. -
Command Interface: The wrapper exposes
get_cdp_session(page)to retrieve the session, enabling direct calls to any CDP command such asNetwork.enable,Page.navigate, orRuntime.evaluate.
Because the CDP session attaches to a specific page, you can freely mix low-level protocol commands with high-level Playwright APIs like page.goto() or page.click().
Using Playwright CDP Mode with MediaCrawler: Step-by-Step
Below is a complete workflow that demonstrates initializing the CDP browser, intercepting network requests, and evaluating JavaScript through the protocol:
# example_use_cdp.py
import asyncio
from tools.cdp_browser import CdpBrowser
async def main():
# 1️⃣ Initialise the CDP‑enabled browser
browser = await CdpBrowser.launch(headless=True, use_cdp=True)
# 2️⃣ Open a new page (Playwright page object)
page = await browser.new_page()
# 3️⃣ Get the low‑level CDP session for this page
cdp = await browser.get_cdp_session(page)
# 4️⃣ Enable network tracking via CDP
await cdp.send("Network.enable")
# 5️⃣ Intercept and log every request URL
async def on_request(**kwargs):
print("Request:", kwargs.get("request", {}).get("url"))
cdp.on("Network.requestWillBeSent", on_request)
# 6️⃣ Navigate using the high‑level Playwright API (still works)
await page.goto("https://example.com")
# 7️⃣ Evaluate a script via CDP (alternative to page.evaluate)
result = await cdp.send(
"Runtime.evaluate",
{"expression": "document.title", "returnByValue": True},
)
print("Page title (CDP):", result["result"]["value"])
# 8️⃣ Close everything
await browser.close()
asyncio.run(main())
Execute the script with:
python example_use_cdp.py
This example prints each network request URL via CDP event listeners, then retrieves the page title using Runtime.evaluate, demonstrating the hybrid approach of combining protocol-level access with standard Playwright navigation.
Advanced CDP Commands for Web Scraping
MediaCrawler's CDP integration supports sophisticated scraping scenarios. Here are practical implementations for common tasks:
Block Image Resources
Prevent bandwidth waste by blocking specific file extensions:
await cdp.send("Network.setBlockedURLs", {"urls": ["*.png", "*.jpg", "*.gif"]})
Emulate Slow Network Conditions
Test behavior under constrained bandwidth or high latency:
await cdp.send("Network.emulateNetworkConditions", {
"offline": False,
"latency": 200,
"downloadThroughput": 500 * 1024,
"uploadThroughput": 500 * 1024
})
Capture Performance Metrics
Enable and retrieve detailed browser performance data:
await cdp.send("Performance.enable")
metrics = await cdp.send("Performance.getMetrics")
Testing and Validation
The tests/test_cdp_browser.py file contains the official test suite verifying CDP functionality, including session creation, command sending, and event handling. When extending MediaCrawler or debugging CDP interactions, reference these tests to understand expected API patterns and validation strategies.
Summary
- MediaCrawler wraps Playwright's CDP capabilities in
tools/cdp_browser.py, providing convenient methods likesend_cdp_cmdandget_cdp_session. - The
tools/browser_launcher.pymodule handles the--remote-debugging-port=0configuration automatically when CDP mode is enabled. - You can mix high-level Playwright APIs (
page.goto) with low-level CDP commands (Network.enable,Runtime.evaluate) on the same page instance. - Common use cases include network interception, resource blocking, performance monitoring, and network condition emulation.
- Reference
tests/test_cdp_browser.pyfor implementation examples and validation patterns.
Frequently Asked Questions
How do I enable CDP mode in MediaCrawler?
Instantiate the CdpBrowser class with use_cdp=True (or rely on the default configuration in tools/cdp_browser.py). The launcher automatically appends --remote-debugging-port=0 to the Chromium arguments and establishes a CDP session via await context.newCDPSession(page).
Can I use regular Playwright methods alongside CDP commands?
Yes. The CDP session attaches to a specific page while preserving the standard Playwright Page object. You can call page.goto() or page.click() for high-level actions, then use cdp.send() for protocol-level operations like Network.enable or Runtime.evaluate on the same page.
What CDP commands are most useful for web scraping?
The most valuable commands include Network.setBlockedURLs to prevent loading images and scripts, Network.emulateNetworkConditions to simulate slow connections, and Performance.getMetrics to capture loading statistics. The Network.requestWillBeSent event also enables detailed request interception and logging.
Where can I find working examples of CDP usage with MediaCrawler?
The tests/test_cdp_browser.py file in the repository contains verified test cases demonstrating session creation, command execution, and event listening. Additionally, the example_use_cdp.py snippet in the documentation above provides a runnable template for common scraping workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →