When to Use Playwright versus CDP Mode in MediaCrawler

Use standard Playwright mode for isolated, stateless crawls with dynamic proxy support, and choose CDP Mode when you need to reuse an existing Chrome profile with cookies and extensions to minimize anti-bot detection.

MediaCrawler supports two distinct browser automation strategies: standard Playwright (which launches a fresh browser instance) and CDP Mode (Chrome DevTools Protocol, which connects to an already-running Chrome/Edge process). Understanding the architectural differences between these modes is essential for configuring reliable crawls that balance isolation against detection avoidance.

How Browser Launch Modes Work in MediaCrawler

The crawler determines which mode to use by checking config.ENABLE_CDP_MODE at runtime. In media_platform/zhihu/core.py lines 84-95, the start method conditionally selects the launch strategy:

if config.ENABLE_CDP_MODE:
    browser_ctx = await self.launch_browser_with_cdp(...)
else:
    browser_ctx = await self.launch_browser(...)

Standard Playwright launches a new Chromium process via chromium.launch or chromium.launch_persistent_context. This occurs in launch_browser (lines 36-48 of core.py), where the browser starts with a clean slate, accepting proxy configurations and custom launch arguments directly.

CDP Mode connects to an existing browser instance via playwright.chromium.connect_over_cdp. The CDPBrowserManager class in tools/cdp_browser.py handles detection of running Chrome instances or optionally launches a new one with --remote-debugging-port, then attaches via WebSocket.

When to Use Standard Playwright Mode

Choose the standard Playwright implementation when you need fresh, isolated browser sessions and full control over the launch environment.

Ideal use cases include:

  • Stateless crawls – You do not need persistent login cookies, browser extensions, or browsing history from previous sessions.
  • CI/CD pipelines – Automated testing environments require guaranteed isolation between runs to prevent state contamination.
  • Dynamic proxy requirements – Each crawl may require a different proxy server. Playwright accepts proxy parameters directly during launch (proxy=playwright_proxy in launch_browser).
  • Simplified debugging – Native Playwright tracing, screenshots, and video recording work without manual browser management.

Implementation example (from media_platform/zhihu/core.py lines 94-99):

chromium = playwright.chromium
self.browser_context = await self.launch_browser(
    chromium, playwright_proxy, self.user_agent, headless=config.HEADLESS
)

When to Use CDP Mode

Select CDP Mode when realism and persistence matter more than isolation.

Ideal use cases include:

  • Profile reuse – You need the exact cookies, localStorage, and extension state from a user's real Chrome profile to maintain logged-in sessions.
  • Maximum anti-detection – CDP connects to the user's actual Chrome installation, making the browser fingerprint indistinguishable from normal human browsing.
  • Interactive debugging – You can manually solve CAPTCHAs, complete slide verifications, or inspect the DOM in the visible browser while the crawler executes.
  • Restricted environments – Corporate or sandboxed platforms that prevent launching new executable processes but allow connecting to existing browser instances.

Implementation example (from media_platform/zhihu/core.py lines 85-92):

self.cdp_manager = CDPBrowserManager()
browser_context = await self.cdp_manager.launch_and_connect(
    playwright, playwright_proxy, self.user_agent, headless=config.CDP_HEADLESS,
)

The CDPBrowserManager implementation in tools/cdp_browser.py lines 97-108 handles the connection logic:

if config.CDP_CONNECT_EXISTING:
    return await self._connect_existing_browser(playwright, playwright_proxy, user_agent)

# otherwise launch a new browser and connect

await self._launch_browser(browser_path, headless)
await self._connect_via_cdp(playwright)

Key Configuration Differences

Understanding the operational divergence between these modes prevents configuration errors.

Proxy Handling

  • Playwright: Fully supported via the proxy argument passed to chromium.launch.
  • CDP Mode: Proxy settings may be ignored because the browser is already running. The system logs a warning in tools/cdp_browser.py lines 88-94 when proxy configuration is detected in CDP mode.

Headless Operation

  • Playwright: headless=True functions reliably in all environments.
  • CDP Mode: Controlled by config.CDP_HEADLESS. Even when headless, some stealth techniques may be less effective (as noted in config/base_config.py line 73).

Resource Cleanup

  • Playwright: Automatic cleanup when the browser context closes.
  • CDP Mode: Requires explicit cleanup via cdp_manager.cleanup() (see close method in core.py lines 92-99) because the browser process may be user-owned. The auto_close_browser flag controls whether MediaCrawler terminates the Chrome process or leaves it running.

Performance Characteristics

  • Playwright: Slower startup due to new process creation, but consistent performance.
  • CDP Mode: Faster connection after initial setup since it attaches to an existing process.

Complete Configuration Examples

Switching Modes via Configuration

Set the mode in config/base_config.py (lines 55-60) before running the crawler:


# For standard Playwright

ENABLE_CDP_MODE = False
HEADLESS = True

# For CDP Mode

ENABLE_CDP_MODE = True
CDP_CONNECT_EXISTING = True  # Connect to already running Chrome

CDP_HEADLESS = False

Standard Playwright Connection

from playwright.async_api import async_playwright

async def run_standard():
    async with async_playwright() as pw:
        browser = await pw.chromium.launch(
            headless=False, 
            proxy={"server": "http://127.0.0.1:8080"}
        )
        context = await browser.new_context()
        page = await context.new_page()
        await page.goto("https://example.com")
        # Crawl operations...

        await browser.close()

CDP Connection to Existing Chrome

Ensure Chrome is started with --remote-debugging-port=9222 before executing:

from playwright.async_api import async_playwright
from tools.cdp_browser import CDPBrowserManager

async def run_cdp():
    async with async_playwright() as pw:
        cdp_mgr = CDPBrowserManager()
        ctx = await cdp_mgr.launch_and_connect(
            playwright=pw,
            playwright_proxy=None,
            user_agent=None,
            headless=False,
        )
        page = await ctx.new_page()
        await page.goto("https://example.com")
        # Crawl operations...

        await cdp_mgr.cleanup()

Summary

  • Standard Playwright (ENABLE_CDP_MODE = False) provides isolated, configurable browser sessions ideal for CI pipelines and stateless scraping where proxies must change between runs.
  • CDP Mode (ENABLE_CDP_MODE = True) connects to existing Chrome instances, preserving user profiles, cookies, and extensions to minimize anti-bot detection at the cost of reduced isolation.
  • Proxy configurations work reliably only in standard Playwright mode; CDP mode issues warnings when proxies are specified since the browser is already running.
  • Cleanup requirements differ significantly: Playwright handles cleanup automatically, while CDP Mode requires explicit cleanup() calls to manage the potentially shared browser process.

Frequently Asked Questions

Can I switch between Playwright and CDP mode without modifying the source code?

Yes. Set ENABLE_CDP_MODE in config/base_config.py to True or False before running the crawler. The media_platform/zhihu/core.py (and other platform cores) read this flag at lines 84-95 to determine which launch method to invoke, allowing mode switching without code changes.

Does CDP mode support headless operation?

Yes, but with limitations. Configure CDP_HEADLESS in config/base_config.py. However, as noted in line 73 of base_config.py, some stealth techniques may be less effective when running headless via CDP compared to standard Playwright, potentially increasing detection risk on heavily protected sites.

Why does my proxy configuration not work in CDP mode?

CDP mode connects to an already-running Chrome process, so proxy settings passed via Playwright parameters are ignored. The CDPBrowserManager in tools/cdp_browser.py lines 88-94 detects this configuration conflict and logs a warning. To use proxies with CDP, you must configure the proxy in the Chrome launch arguments before MediaCrawler connects, not via the crawler's proxy settings.

Which mode is better for avoiding anti-bot detection?

CDP Mode generally provides better anti-detection capabilities because it utilizes the user's actual Chrome installation with real browser extensions, saved cookies, and consistent fingerprinting. Standard Playwright launches a generic Chromium instance that may lack the entropy of a real user profile, making it more susceptible to sophisticated bot detection algorithms.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →