How to Configure CDP Mode in MediaCrawler with Existing Chrome and Login State

Enable CDP mode in config/base_config.py, launch Chrome with --remote-debugging-port=9222, and set CDP_CONNECT_EXISTING=True to attach MediaCrawler to your existing browser instance while preserving cookies and login state through SAVE_LOGIN_STATE=True.

The MediaCrawler repository supports Chrome DevTools Protocol (CDP) mode, allowing the crawler to control a real Chrome browser instead of a headless instance. This configuration significantly improves anti-detection capabilities and enables seamless reuse of existing login sessions by connecting to an already-running Chrome process with remote debugging enabled.

Understanding CDP Configuration Architecture

MediaCrawler’s CDP implementation relies on two primary components: the global configuration file that defines connection parameters, and the browser manager class that handles the actual CDP lifecycle.

Key Parameters in base_config.py

The central configuration resides in config/base_config.py (lines 55-84), where several boolean flags control CDP behavior:

  • ENABLE_CDP_MODE – Activates CDP support across all platforms. When True, crawlers instantiate CDPBrowserManager instead of standard Playwright launchers.
  • CDP_CONNECT_EXISTING – When set to True, the crawler attempts to attach to a Chrome instance already running with remote debugging. When False, it launches a fresh Chrome process.
  • SAVE_LOGIN_STATE – Enables persistent storage of browser data including cookies, localStorage, and session information in a dedicated user-data directory.
  • CDP_DEBUG_PORT – Defines the port for remote-debugging connections (default: 9222).
  • USER_DATA_DIR – Format string generating platform-specific paths like cdp_zhihu_user_data_dir under the repository root.

The CDPBrowserManager Implementation

Located in tools/cdp_browser.py, the CDPBrowserManager class orchestrates browser connections through three critical methods (lines 97-118, 150-176, 250-263):

  1. _connect_existing_browser – Polls the configured debug port until Chrome’s remote-debugging endpoint becomes available, then establishes a connection via playwright.chromium.connect_over_cdp.
  2. _launch_browser – Handles fresh browser launches with custom user-data directories when CDP_CONNECT_EXISTING is False.
  3. cleanup – Respects the AUTO_CLOSE_BROWSER setting, ensuring that externally launched Chrome instances remain running while properly closing only crawler-initiated processes.

Platform-specific crawlers in media_platform/*/core.py (such as media_platform/zhihu/core.py lines 458-470) integrate this manager by calling launch_browser_with_cdp when ENABLE_CDP_MODE is active.

Step-by-Step Configuration Guide

Follow these steps to configure MediaCrawler to use your existing Chrome installation with persistent login state.

Step 1: Modify base_config.py

Update the configuration file to enable CDP mode and connection preferences:


# config/base_config.py

ENABLE_CDP_MODE = True          # Activate CDP support

CDP_CONNECT_EXISTING = True     # Connect to running Chrome instance

CDP_DEBUG_PORT = 9222           # Must match Chrome's remote-debugging port

SAVE_LOGIN_STATE = True         # Preserve cookies and session data

When SAVE_LOGIN_STATE is enabled, MediaCrawler creates a directory at <repo_root>/browser_data/cdp_<platform>_user_data_dir to store browser profiles.

Step 2: Launch Chrome with Remote Debugging

Start your Chrome browser manually with the remote-debugging flag before running the crawler:


# Windows

"C:\Program Files\Google\Chrome\Application\chrome.exe" --remote-debugging-port=9222

# macOS

/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222

# Linux

google-chrome --remote-debugging-port=9222

This command exposes Chrome’s DevTools Protocol on port 9222, allowing MediaCrawler to attach to this specific process and inherit all active sessions, cookies, and extensions.

Step 3: Execute the Crawler

With Chrome running in the background, execute your target platform crawler normally:

python -m MediaCrawler zhihu --keyword "python"

The CDPBrowserManager automatically detects the debug port, establishes the connection via connect_over_cdp, and navigates using the existing browser context. Since SAVE_LOGIN_STATE persists data to browser_data/cdp_zhihu_user_data_dir, subsequent runs maintain your login state even if you restart Chrome.

Managing Login State Persistence

The interaction between CDP_CONNECT_EXISTING and SAVE_LOGIN_STATE determines how authentication persists across sessions:

  • Existing Chrome + Saved State: When connecting to your personal Chrome profile, cookies and logins from your daily browsing are immediately available to the crawler. The user_data_dir parameter ensures that any changes made during crawling (like new session tokens) persist to the filesystem.
  • Clearing Saved State: To force re-authentication, delete the platform-specific data directory:
rm -rf browser_data/cdp_zhihu_user_data_dir

After deletion, the next crawler run will present a fresh browser context requiring new login credentials.

Practical Implementation Example

Below is a standalone script demonstrating the CDP connection flow programmatically:

import asyncio
from tools.cdp_browser import CDPBrowserManager
from playwright.async_api import async_playwright

async def crawl_with_cdp():
    async with async_playwright() as playwright:
        # Initialize manager (reads config/base_config.py automatically)

        cdp_manager = CDPBrowserManager()
        
        # Connect to existing Chrome or launch new instance

        context = await cdp_manager.launch_and_connect(playwright, headless=False)
        
        # Create new page - inherits cookies from existing Chrome session

        page = await context.new_page()
        await page.goto("https://www.zhihu.com")
        await page.wait_for_load_state("networkidle")
        
        print(f"Page title: {await page.title()}")
        
        # Cleanup unregisters context but leaves external Chrome running

        await cdp_manager.cleanup()

if __name__ == "__main__":
    asyncio.run(crawl_with_cdp())

Run this script after starting Chrome with --remote-debugging-port=9222. The output will show the Zhihu homepage title without requiring manual login, confirming that the CDP connection successfully transferred your existing authentication state.

Summary

  • Enable CDP mode by setting ENABLE_CDP_MODE = True in config/base_config.py to switch from headless to CDP-based crawling.
  • Connect to existing Chrome by launching the browser with --remote-debugging-port=9222 and keeping CDP_CONNECT_EXISTING set to True.
  • Preserve login state using SAVE_LOGIN_STATE = True, which stores profile data in browser_data/cdp_<platform>_user_data_dir for persistence across crawler restarts.
  • Avoid duplicate launches by ensuring AUTO_CLOSE_BROWSER respects external Chrome instances, leaving your personal browser untouched when cleanup runs.

Frequently Asked Questions

What is CDP mode in MediaCrawler?

CDP mode (Chrome DevTools Protocol mode) is a configuration that allows MediaCrawler to control a real Chrome or Edge browser instance instead of using Playwright’s built-in Chromium downloads. According to the source code in tools/cdp_browser.py, this mode uses playwright.chromium.connect_over_cdp to attach to browsers exposing their debug protocol, enabling access to existing cookies, extensions, and real-user browser fingerprints that improve anti-detection capabilities.

How do I maintain my login state between crawler runs?

Set SAVE_LOGIN_STATE = True in config/base_config.py. This instructs the CDPBrowserManager to create a persistent user-data directory (e.g., browser_data/cdp_zhihu_user_data_dir) that Chrome uses for storing cookies, localStorage, and session tokens. When combined with CDP_CONNECT_EXISTING = True, your existing Chrome profile remains intact, and the crawler inherits all active logins without requiring manual authentication each time.

Can I use Microsoft Edge instead of Chrome for CDP mode?

Yes, the CDP implementation in tools/cdp_browser.py is browser-agnostic regarding the Chrome DevTools Protocol. Start Edge with the --remote-debugging-port flag (e.g., msedge --remote-debugging-port=9222), ensure CDP_DEBUG_PORT matches in your configuration, and MediaCrawler will connect successfully. You can also specify a custom browser path using the CUSTOM_BROWSER_PATH variable in base_config.py if Edge is not in your system PATH.

Why does MediaCrawler fail to connect to my existing Chrome instance?

Connection failures typically occur when the remote-debugging port is inaccessible or Chrome was not started with the correct flag. Verify that:

  1. Chrome was launched with --remote-debugging-port=9222 (or your configured CDP_DEBUG_PORT).
  2. No firewall or network policy blocks localhost connections to that port.
  3. The _test_cdp_connection method in cdp_browser.py confirms the port is reachable before attempting connect_over_cdp.

If Chrome starts but MediaCrawler times out, check that no other process occupies port 9222 using netstat -an | grep 9222 (Linux/macOS) or netstat -ano | findstr 9222 (Windows).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →