# How to Configure CDP Mode in MediaCrawler with Existing Chrome and Login State

> Configure CDP mode in MediaCrawler to attach to an existing Chrome instance preserving login state. Learn how to launch Chrome and set config options for seamless integration.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**Enable CDP mode in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), launch Chrome with `--remote-debugging-port=9222`, and set `CDP_CONNECT_EXISTING=True` to attach MediaCrawler to your existing browser instance while preserving cookies and login state through `SAVE_LOGIN_STATE=True`.**

The **MediaCrawler** repository supports **Chrome DevTools Protocol (CDP)** mode, allowing the crawler to control a real Chrome browser instead of a headless instance. This configuration significantly improves anti-detection capabilities and enables seamless reuse of existing login sessions by connecting to an already-running Chrome process with remote debugging enabled.

## Understanding CDP Configuration Architecture

MediaCrawler’s CDP implementation relies on two primary components: the global configuration file that defines connection parameters, and the browser manager class that handles the actual CDP lifecycle.

### Key Parameters in base_config.py

The central configuration resides in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) (lines 55-84), where several boolean flags control CDP behavior:

- **`ENABLE_CDP_MODE`** – Activates CDP support across all platforms. When `True`, crawlers instantiate `CDPBrowserManager` instead of standard Playwright launchers.
- **`CDP_CONNECT_EXISTING`** – When set to `True`, the crawler attempts to attach to a Chrome instance already running with remote debugging. When `False`, it launches a fresh Chrome process.
- **`SAVE_LOGIN_STATE`** – Enables persistent storage of browser data including cookies, localStorage, and session information in a dedicated user-data directory.
- **`CDP_DEBUG_PORT`** – Defines the port for remote-debugging connections (default: 9222).
- **`USER_DATA_DIR`** – Format string generating platform-specific paths like `cdp_zhihu_user_data_dir` under the repository root.

### The CDPBrowserManager Implementation

Located in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py), the `CDPBrowserManager` class orchestrates browser connections through three critical methods (lines 97-118, 150-176, 250-263):

1. **`_connect_existing_browser`** – Polls the configured debug port until Chrome’s remote-debugging endpoint becomes available, then establishes a connection via `playwright.chromium.connect_over_cdp`.
2. **`_launch_browser`** – Handles fresh browser launches with custom user-data directories when `CDP_CONNECT_EXISTING` is `False`.
3. **`cleanup`** – Respects the `AUTO_CLOSE_BROWSER` setting, ensuring that externally launched Chrome instances remain running while properly closing only crawler-initiated processes.

Platform-specific crawlers in `media_platform/*/core.py` (such as [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) lines 458-470) integrate this manager by calling `launch_browser_with_cdp` when `ENABLE_CDP_MODE` is active.

## Step-by-Step Configuration Guide

Follow these steps to configure MediaCrawler to use your existing Chrome installation with persistent login state.

### Step 1: Modify base_config.py

Update the configuration file to enable CDP mode and connection preferences:

```python

# config/base_config.py

ENABLE_CDP_MODE = True          # Activate CDP support

CDP_CONNECT_EXISTING = True     # Connect to running Chrome instance

CDP_DEBUG_PORT = 9222           # Must match Chrome's remote-debugging port

SAVE_LOGIN_STATE = True         # Preserve cookies and session data

```

When `SAVE_LOGIN_STATE` is enabled, MediaCrawler creates a directory at `<repo_root>/browser_data/cdp_<platform>_user_data_dir` to store browser profiles.

### Step 2: Launch Chrome with Remote Debugging

Start your Chrome browser manually with the remote-debugging flag **before** running the crawler:

```bash

# Windows

"C:\Program Files\Google\Chrome\Application\chrome.exe" --remote-debugging-port=9222

# macOS

/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222

# Linux

google-chrome --remote-debugging-port=9222

```

This command exposes Chrome’s DevTools Protocol on port 9222, allowing MediaCrawler to attach to this specific process and inherit all active sessions, cookies, and extensions.

### Step 3: Execute the Crawler

With Chrome running in the background, execute your target platform crawler normally:

```bash
python -m MediaCrawler zhihu --keyword "python"

```

The `CDPBrowserManager` automatically detects the debug port, establishes the connection via `connect_over_cdp`, and navigates using the existing browser context. Since `SAVE_LOGIN_STATE` persists data to `browser_data/cdp_zhihu_user_data_dir`, subsequent runs maintain your login state even if you restart Chrome.

## Managing Login State Persistence

The interaction between `CDP_CONNECT_EXISTING` and `SAVE_LOGIN_STATE` determines how authentication persists across sessions:

- **Existing Chrome + Saved State**: When connecting to your personal Chrome profile, cookies and logins from your daily browsing are immediately available to the crawler. The `user_data_dir` parameter ensures that any changes made during crawling (like new session tokens) persist to the filesystem.
- **Clearing Saved State**: To force re-authentication, delete the platform-specific data directory:

```bash
rm -rf browser_data/cdp_zhihu_user_data_dir

```

After deletion, the next crawler run will present a fresh browser context requiring new login credentials.

## Practical Implementation Example

Below is a standalone script demonstrating the CDP connection flow programmatically:

```python
import asyncio
from tools.cdp_browser import CDPBrowserManager
from playwright.async_api import async_playwright

async def crawl_with_cdp():
    async with async_playwright() as playwright:
        # Initialize manager (reads config/base_config.py automatically)

        cdp_manager = CDPBrowserManager()
        
        # Connect to existing Chrome or launch new instance

        context = await cdp_manager.launch_and_connect(playwright, headless=False)
        
        # Create new page - inherits cookies from existing Chrome session

        page = await context.new_page()
        await page.goto("https://www.zhihu.com")
        await page.wait_for_load_state("networkidle")
        
        print(f"Page title: {await page.title()}")
        
        # Cleanup unregisters context but leaves external Chrome running

        await cdp_manager.cleanup()

if __name__ == "__main__":
    asyncio.run(crawl_with_cdp())

```

Run this script after starting Chrome with `--remote-debugging-port=9222`. The output will show the Zhihu homepage title without requiring manual login, confirming that the CDP connection successfully transferred your existing authentication state.

## Summary

- **Enable CDP mode** by setting `ENABLE_CDP_MODE = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to switch from headless to CDP-based crawling.
- **Connect to existing Chrome** by launching the browser with `--remote-debugging-port=9222` and keeping `CDP_CONNECT_EXISTING` set to `True`.
- **Preserve login state** using `SAVE_LOGIN_STATE = True`, which stores profile data in `browser_data/cdp_<platform>_user_data_dir` for persistence across crawler restarts.
- **Avoid duplicate launches** by ensuring `AUTO_CLOSE_BROWSER` respects external Chrome instances, leaving your personal browser untouched when cleanup runs.

## Frequently Asked Questions

### What is CDP mode in MediaCrawler?

**CDP mode** (Chrome DevTools Protocol mode) is a configuration that allows MediaCrawler to control a real Chrome or Edge browser instance instead of using Playwright’s built-in Chromium downloads. According to the source code in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py), this mode uses `playwright.chromium.connect_over_cdp` to attach to browsers exposing their debug protocol, enabling access to existing cookies, extensions, and real-user browser fingerprints that improve anti-detection capabilities.

### How do I maintain my login state between crawler runs?

Set **`SAVE_LOGIN_STATE = True`** in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). This instructs the `CDPBrowserManager` to create a persistent user-data directory (e.g., `browser_data/cdp_zhihu_user_data_dir`) that Chrome uses for storing cookies, localStorage, and session tokens. When combined with `CDP_CONNECT_EXISTING = True`, your existing Chrome profile remains intact, and the crawler inherits all active logins without requiring manual authentication each time.

### Can I use Microsoft Edge instead of Chrome for CDP mode?

Yes, the CDP implementation in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) is browser-agnostic regarding the Chrome DevTools Protocol. Start Edge with the `--remote-debugging-port` flag (e.g., `msedge --remote-debugging-port=9222`), ensure `CDP_DEBUG_PORT` matches in your configuration, and MediaCrawler will connect successfully. You can also specify a custom browser path using the `CUSTOM_BROWSER_PATH` variable in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) if Edge is not in your system PATH.

### Why does MediaCrawler fail to connect to my existing Chrome instance?

Connection failures typically occur when the remote-debugging port is inaccessible or Chrome was not started with the correct flag. Verify that:
1. Chrome was launched with `--remote-debugging-port=9222` (or your configured `CDP_DEBUG_PORT`).
2. No firewall or network policy blocks localhost connections to that port.
3. The `_test_cdp_connection` method in [`cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cdp_browser.py) confirms the port is reachable before attempting `connect_over_cdp`.

If Chrome starts but MediaCrawler times out, check that no other process occupies port 9222 using `netstat -an | grep 9222` (Linux/macOS) or `netstat -ano | findstr 9222` (Windows).