How to Use Playwright CDP Mode with MediaCrawler: A Complete Guide

MediaCrawler provides a thin wrapper around Playwright that enables Chrome DevTools Protocol (CDP) access through the CdpBrowser class in tools/cdp_browser.py, allowing you to send raw CDP commands like Network.enable while retaining high-level Playwright APIs.

MediaCrawler, an open-source scraping framework, ships with built-in support for Playwright's Chrome DevTools Protocol (CDP) mode to give developers low-level browser control. By leveraging Playwright CDP mode with MediaCrawler, you can intercept network traffic, manipulate performance metrics, and access CDP-only APIs that standard Playwright methods don't expose. This integration bridges the gap between high-level browser automation and raw protocol-level interactions.

What is CDP Mode and Why Use It?

The Chrome DevTools Protocol (CDP) provides granular access to Chromium's internals, enabling operations like network interception, JavaScript debugging, and performance monitoring. While standard Playwright offers robust automation APIs, CDP mode unlocks advanced capabilities such as blocking specific resource types, emulating network conditions, or extracting browser metrics directly from the protocol layer.

How MediaCrawler Implements CDP Support

Core Architecture in tools/cdp_browser.py

The primary implementation resides in tools/cdp_browser.py, which creates a Playwright BrowserContext with CDP capabilities enabled. According to the MediaCrawler source code, this module exposes helper methods including send_cdp_cmd and wait_for_cdp_event to simplify protocol interactions without manual JSON serialization.

The launcher abstraction in tools/browser_launcher.py handles the initialization logic, automatically injecting --remote-debugging-port=0 into the launch arguments when CDP mode is requested. This allows Playwright to establish a debug session on an available port while maintaining the standard browser context.

The CDP Session Lifecycle

When you instantiate a CdpBrowser through the public API (as demonstrated in tests/test_cdp_browser.py), the following sequence occurs:

  1. Playwright Launch: A headless Chromium instance starts with args=["--remote-debugging-port=0"] and headless=True.

  2. Session Creation: The context calls await context.newCDPSession(page) to return a CDPSession object attached to the specific page.

  3. Command Interface: The wrapper exposes get_cdp_session(page) to retrieve the session, enabling direct calls to any CDP command such as Network.enable, Page.navigate, or Runtime.evaluate.

Because the CDP session attaches to a specific page, you can freely mix low-level protocol commands with high-level Playwright APIs like page.goto() or page.click().

Using Playwright CDP Mode with MediaCrawler: Step-by-Step

Below is a complete workflow that demonstrates initializing the CDP browser, intercepting network requests, and evaluating JavaScript through the protocol:


# example_use_cdp.py

import asyncio
from tools.cdp_browser import CdpBrowser

async def main():
    # 1️⃣ Initialise the CDP‑enabled browser

    browser = await CdpBrowser.launch(headless=True, use_cdp=True)

    # 2️⃣ Open a new page (Playwright page object)

    page = await browser.new_page()

    # 3️⃣ Get the low‑level CDP session for this page

    cdp = await browser.get_cdp_session(page)

    # 4️⃣ Enable network tracking via CDP

    await cdp.send("Network.enable")

    # 5️⃣ Intercept and log every request URL

    async def on_request(**kwargs):
        print("Request:", kwargs.get("request", {}).get("url"))
    cdp.on("Network.requestWillBeSent", on_request)

    # 6️⃣ Navigate using the high‑level Playwright API (still works)

    await page.goto("https://example.com")

    # 7️⃣ Evaluate a script via CDP (alternative to page.evaluate)

    result = await cdp.send(
        "Runtime.evaluate", 
        {"expression": "document.title", "returnByValue": True},
    )
    print("Page title (CDP):", result["result"]["value"])

    # 8️⃣ Close everything

    await browser.close()

asyncio.run(main())

Execute the script with:

python example_use_cdp.py

This example prints each network request URL via CDP event listeners, then retrieves the page title using Runtime.evaluate, demonstrating the hybrid approach of combining protocol-level access with standard Playwright navigation.

Advanced CDP Commands for Web Scraping

MediaCrawler's CDP integration supports sophisticated scraping scenarios. Here are practical implementations for common tasks:

Block Image Resources

Prevent bandwidth waste by blocking specific file extensions:

await cdp.send("Network.setBlockedURLs", {"urls": ["*.png", "*.jpg", "*.gif"]})

Emulate Slow Network Conditions

Test behavior under constrained bandwidth or high latency:

await cdp.send("Network.emulateNetworkConditions", {
    "offline": False,
    "latency": 200,
    "downloadThroughput": 500 * 1024,
    "uploadThroughput": 500 * 1024
})

Capture Performance Metrics

Enable and retrieve detailed browser performance data:

await cdp.send("Performance.enable")
metrics = await cdp.send("Performance.getMetrics")

Testing and Validation

The tests/test_cdp_browser.py file contains the official test suite verifying CDP functionality, including session creation, command sending, and event handling. When extending MediaCrawler or debugging CDP interactions, reference these tests to understand expected API patterns and validation strategies.

Summary

  • MediaCrawler wraps Playwright's CDP capabilities in tools/cdp_browser.py, providing convenient methods like send_cdp_cmd and get_cdp_session.
  • The tools/browser_launcher.py module handles the --remote-debugging-port=0 configuration automatically when CDP mode is enabled.
  • You can mix high-level Playwright APIs (page.goto) with low-level CDP commands (Network.enable, Runtime.evaluate) on the same page instance.
  • Common use cases include network interception, resource blocking, performance monitoring, and network condition emulation.
  • Reference tests/test_cdp_browser.py for implementation examples and validation patterns.

Frequently Asked Questions

How do I enable CDP mode in MediaCrawler?

Instantiate the CdpBrowser class with use_cdp=True (or rely on the default configuration in tools/cdp_browser.py). The launcher automatically appends --remote-debugging-port=0 to the Chromium arguments and establishes a CDP session via await context.newCDPSession(page).

Can I use regular Playwright methods alongside CDP commands?

Yes. The CDP session attaches to a specific page while preserving the standard Playwright Page object. You can call page.goto() or page.click() for high-level actions, then use cdp.send() for protocol-level operations like Network.enable or Runtime.evaluate on the same page.

What CDP commands are most useful for web scraping?

The most valuable commands include Network.setBlockedURLs to prevent loading images and scripts, Network.emulateNetworkConditions to simulate slow connections, and Performance.getMetrics to capture loading statistics. The Network.requestWillBeSent event also enables detailed request interception and logging.

Where can I find working examples of CDP usage with MediaCrawler?

The tests/test_cdp_browser.py file in the repository contains verified test cases demonstrating session creation, command execution, and event listening. Additionally, the example_use_cdp.py snippet in the documentation above provides a runnable template for common scraping workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →