How to Use CDP Mode in MediaCrawler to Connect to an Existing Chrome Browser

MediaCrawler’s CDP Mode allows you to attach crawlers to an already-running Chrome instance via the Chrome DevTools Protocol, preserving logged-in sessions and avoiding launch overhead.

You can use CDP Mode in MediaCrawler to bypass the standard Playwright browser launch and instead connect to an existing Chrome window. This approach retains cookies, extensions, and authentication states, significantly reducing the risk of anti-bot detection. The implementation centers on the CDPBrowserManager class in tools/cdp_browser.py, which handles detection, WebSocket negotiation, and context management.

Prerequisites: Enable Remote Debugging in Chrome

Before MediaCrawler can attach to your browser, you must enable remote debugging.

  1. Open Chrome and navigate to chrome://inspect/#remote-debugging
  2. Check the option to allow remote debugging for the current browser instance
  3. Verify the server is running at 127.0.0.1:9222 (or your custom port)

You can verify the port is accessible using a simple socket probe:

python - <<'PY'
import socket
port = 9222
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
s.settimeout(5)
result = s.connect_ex(('localhost', port))
print('Port open' if result == 0 else 'Port closed')
PY

Configuration Settings

Connection behavior is controlled through config/base_config.py. Two flags determine whether MediaCrawler launches a new browser or attaches to an existing one:

  • CDP_CONNECT_EXISTING: Set to True to enable attachment mode
  • CDP_DEBUG_PORT: Specifies the port (default: 9222)

# config/base_config.py

CDP_CONNECT_EXISTING = True
CDP_DEBUG_PORT = 9222

How CDP Browser Management Works

The CDPBrowserManager class in tools/cdp_browser.py orchestrates the connection workflow when CDP_CONNECT_EXISTING is enabled.

Port Detection and Waiting

When launch_and_connect() is invoked, the manager calls _connect_existing_browser(). This method repeatedly tests the debug port using _test_cdp_connection(), which performs a socket probe to verify Chrome is listening. If the connection fails initially, the system logs progress and prompts you to enable remote debugging via the Chrome inspect interface.

WebSocket URL Resolution

If the direct CDP connection fails, the manager falls back to _get_browser_websocket_url(). This method queries the /json/version HTTP endpoint on the debug port to retrieve the precise WebSocket URL required for Playwright attachment.

Browser Context Creation

Once connected, _create_browser_context() either reuses the first existing browser context (if Chrome already has tabs open) or spawns a fresh context with sensible defaults. Fresh contexts are configured with a 1920×1080 viewport and enabled downloads, ensuring compatibility with MediaCrawler’s scraping requirements.

Step-by-Step Implementation

Enable Remote Debugging in Chrome

Launch Chrome with remote debugging enabled, or use the chrome://inspect method described above. Ensure no firewall rules block port 9222.

Configure MediaCrawler

Edit config/base_config.py to set:

CDP_CONNECT_EXISTING = True
CDP_DEBUG_PORT = 9222  # Change if you used a custom startup flag

Alternatively, set environment variables if your deployment prefers external configuration.

Run the Crawler

Execute your crawl command normally. The CDPBrowserManager automatically detects the CDP_CONNECT_EXISTING flag and attaches to your Chrome instance rather than launching a new browser:

uv run main.py --platform xhs --lt qrcode --type search

The CLI entry point in main.py parses arguments and instantiates CDPBrowserManager, which handles the underlying connection logic transparently.

Advanced Usage

Stealth Scripts and Session Management

Once connected, CDPBrowserManager provides methods to inject anti-detection scripts and manage cookies:

  • add_stealth_script(): Injects JavaScript to mask automation fingerprints
  • add_cookies(): Seeds the browser context with session cookies
  • get_cookies(): Extracts current cookies for persistence

These methods operate on the attached context, allowing you to maintain state across crawl sessions without re-authentication.

Programmatic Access

You can also instantiate the manager directly for custom workflows:

from tools.cdp_browser import CDPBrowserManager
from playwright.async_api import async_playwright

async def run():
    async with async_playwright() as pw:
        manager = CDPBrowserManager()
        context = await manager.launch_and_connect(
            playwright=pw,
            playwright_proxy=None,
            user_agent=None,
            headless=False,
        )
        page = await context.new_page()
        await page.goto('https://www.xiaohongshu.com')
        # Proceed with crawler logic

Cleanup Behavior

The cleanup() method ensures graceful shutdown. When CDP_CONNECT_EXISTING is True, it skips process termination to avoid killing your existing Chrome instance, closing only the contexts and connections managed by MediaCrawler.

Summary

  • CDP Mode in MediaCrawler connects to existing Chrome instances via the Chrome DevTools Protocol, preserving login states and reducing detection risks.
  • Configuration is controlled by CDP_CONNECT_EXISTING and CDP_DEBUG_PORT in config/base_config.py.
  • The CDPBrowserManager class in tools/cdp_browser.py handles port detection, WebSocket resolution, and context creation.
  • Use chrome://inspect/#remote-debugging to enable remote debugging on port 9222 before running crawlers.
  • Run commands normally via main.py; the manager automatically attaches to the existing browser when configured.
  • Cleanup routines preserve the external Chrome process while closing managed resources.

Frequently Asked Questions

What is CDP Mode in MediaCrawler?

CDP Mode allows MediaCrawler to connect to an existing Chrome browser instance using the Chrome DevTools Protocol rather than launching a fresh Playwright-managed browser. This preserves the browser’s cookies, extensions, and logged-in sessions, making crawls faster and less likely to trigger anti-bot measures.

How do I verify my Chrome browser is accepting CDP connections?

You can verify the debug port is accessible by running a socket connection test to localhost:9222 (or your configured port). In tools/cdp_browser.py, the _test_cdp_connection() method performs this check. If the port is closed, ensure you enabled remote debugging via chrome://inspect/#remote-debugging and that Chrome was started with the --remote-debugging-port flag if necessary.

Can I use CDP Mode with browsers other than Chrome?

Yes, the implementation supports any Chromium-based browser that implements the Chrome DevTools Protocol, including Microsoft Edge. The CDPBrowserManager uses standard CDP endpoints (/json/version for WebSocket discovery) that are compatible with any Chromium derivative running with remote debugging enabled.

Why should I use CDP Mode instead of launching a new browser?

Using CDP Mode eliminates the overhead of browser startup and allows you to reuse existing authentication states, cookies, and browser profiles. According to the source code in tools/cdp_browser.py, this approach greatly reduces the chance of being blocked by target platforms because the browser fingerprint matches a regular user’s long-running session rather than a fresh automation profile.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →