How MediaCrawler Uses CDP Mode to Bypass Bot Detection
MediaCrawler bypasses bot detection by connecting to a real Chrome or Edge browser instance via the Chrome DevTools Protocol (CDP), allowing it to reuse user profiles, extensions, and cookies while injecting stealth scripts to mask automation fingerprints.
MediaCrawler is an open-source multi-platform content scraper that employs sophisticated anti-detection techniques to evade modern anti-bot systems. Instead of launching a standard headless Playwright instance—which many sites flag as automated traffic—the crawler leverages CDP mode to control an actual browser installation. This approach allows MediaCrawler to present the exact same fingerprints, session state, and behavioral characteristics as a genuine user browsing session.
What is CDP Mode and Why It Matters
CDP (Chrome DevTools Protocol) mode enables programmatic control over a running Chrome or Edge browser instance through a WebSocket connection. Unlike traditional browser automation that launches isolated, ephemeral browser processes, CDP mode allows MediaCrawler to attach to an existing browser profile that contains real user data, extensions, and persistent cookies. This significantly raises the barrier for bot detection systems, as the crawler inherits the target machine's actual browser fingerprint rather than a synthetic one.
How MediaCrawler Implements CDP Mode
The CDP implementation centers on the CDPBrowserManager class in tools/cdp_browser.py, which orchestrates browser detection, launch, connection, and context management.
Browser Detection and Launch
MediaCrawler first locates a valid Chrome or Edge binary to use as the automation target. The CDPBrowserManager checks config.CUSTOM_BROWSER_PATH for a user-specified browser location, or auto-detects installed Chrome/Edge executables on the system.
Once detected, the browser launches with remote debugging enabled on a configurable port (config.CDP_DEBUG_PORT). The launch parameters include --remote-debugging-port, an optional --headless flag controlled by config.CDP_HEADLESS, and a user-data directory to preserve cookies, extensions, and browser state between sessions.
# From tools/cdp_browser.py - browser detection logic
def _find_browser_executable(self):
custom_path = config.CUSTOM_BROWSER_PATH
if custom_path and Path(custom_path).exists():
return custom_path
# Auto-detection for Chrome/Edge on various platforms
...
Remote Debugging Connection
After launching the browser, MediaCrawler establishes a CDP connection using Playwright's chromium.connect_over_cdp method. The manager retrieves the WebSocket URL from the browser's /json/version endpoint and creates a persistent connection to the debugging protocol.
This connection grants the crawler full control over the already-running browser instance, including access to all existing tabs, cookies, and storage. As implemented in tools/cdp_browser.py (lines 313-327), this approach allows MediaCrawler to reuse the exact browser profile that the user employs for daily browsing.
# CDP connection establishment
browser = await playwright.chromium.connect_over_cdp(
f"http://localhost:{config.CDP_DEBUG_PORT}",
timeout=30000
)
Context Creation and Stealth Injection
Upon connecting, MediaCrawler either attaches to an existing browser context or creates a fresh one with normal viewport settings and download handling. Crucially, the crawler injects an anti-detection script (libs/stealth.min.js) into the context to remove common Playwright automation fingerprints.
The add_stealth_script method in tools/cdp_browser.py (lines 400-410) executes this JavaScript snippet, which patches navigator properties, masks webdriver flags, and modifies other indicators that bot detection systems typically probe.
# Stealth script injection
await self.cdp_manager.add_stealth_script("libs/stealth.min.js")
Cookie Sharing with HTTP Client
To maintain consistent authentication state across both browser and API interactions, MediaCrawler extracts cookies from the CDP browser context and feeds them to platform-specific HTTP clients. This happens in platform crawlers like media_platform/zhihu/core.py (lines 99-105), where the crawler converts browser cookies into a format compatible with the httpx or requests library.
The convert_browser_context_cookies utility extracts the current session cookies and passes them to the ZhiHuClient or other platform-specific clients, ensuring that HTTP requests carry the same authentication tokens as the browser sessions.
# Cookie extraction for HTTP client usage
cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
self.browser_context, urls=self.cookie_urls
)
zhihu_client = ZhiHuClient(
headers={... "cookie": cookie_str, ...},
cookie_dict=cookie_dict,
)
Configuration and Fallback Mechanisms
MediaCrawler provides granular control over CDP behavior through config/base_config.py. Key configuration options include:
- ENABLE_CDP_MODE: Toggle between CDP and standard Playwright launch
- CDP_DEBUG_PORT: Specify the remote debugging port (default: 9222)
- CDP_HEADLESS: Control whether the browser window is visible
- CDP_CONNECT_EXISTING: Connect to an already-running browser instance
- CUSTOM_BROWSER_PATH: Specify a non-standard browser location
The system includes graceful degradation logic in platform crawlers like media_platform/zhihu/core.py (lines 84-90). If CDP launch fails or the browser is unavailable, the crawler automatically reverts to Playwright's standard launch mode, ensuring reliability across different environments.
# Fallback logic from media_platform/zhihu/core.py
if config.ENABLE_CDP_MODE:
try:
self.browser_context = await self.launch_browser_with_cdp(...)
except Exception as e:
utils.logger.error(f"CDP mode failed: {e}, falling back to standard launch")
# Standard Playwright launch
Code Implementation Examples
Platform-specific crawlers integrate CDP mode through a consistent pattern. Here is how the Zhihu crawler initializes the CDP manager:
# ZhihuCrawler.start() implementation
if config.ENABLE_CDP_MODE:
utils.logger.info("[ZhihuCrawler] Launching browser in CDP mode")
self.cdp_manager = CDPBrowserManager()
self.browser_context = await self.cdp_manager.launch_and_connect(
playwright=playwright,
playwright_proxy=playwright_proxy,
user_agent=self.user_agent,
headless=config.CDP_HEADLESS,
)
await self.cdp_manager.add_stealth_script("libs/stealth.min.js")
else:
# Standard Playwright launch path
self.browser_context = await self.launch_browser(...)
The same CDP integration pattern appears across other platforms including XiaoHongShu (media_platform/xhs/core.py) and Weibo (media_platform/weibo/core.py), demonstrating the architecture's reusability.
Summary
- Real Browser Control: MediaCrawler uses
CDPBrowserManagerintools/cdp_browser.pyto launch and control actual Chrome/Edge instances via the Chrome DevTools Protocol. - Profile Reuse: By connecting to existing browsers with user-data directories, the crawler inherits genuine browser fingerprints, extensions, and persistent cookies.
- Stealth Capabilities: The
libs/stealth.min.jsscript removes automation fingerprints from the browser context, masking Playwright's presence. - Session Consistency: Cookie extraction utilities ensure HTTP API clients share the same authentication state as the browser, maintaining valid sessions across request types.
- Robust Fallback: If CDP mode fails, the system automatically falls back to standard Playwright launch, ensuring operational continuity.
Frequently Asked Questions
What is the advantage of CDP mode over standard Playwright?
CDP mode connects to a real browser installation that contains the user's actual browsing history, extensions, and saved cookies. Standard Playwright launches isolated, ephemeral browser instances with default fingerprints that bot detection systems easily identify. By using CDP, MediaCrawler presents the same browser characteristics as a human user, significantly reducing detection rates.
How does MediaCrawler handle cases where Chrome is not installed?
If CDPBrowserManager cannot locate a browser binary via config.CUSTOM_BROWSER_PATH or auto-detection, or if the CDP connection fails, the crawler falls back to Playwright's standard browser launch mode. This fallback logic ensures the scraper remains functional even in environments without Chrome or Edge installed, though with potentially higher detection risk.
Can CDP mode work with headless browsers?
Yes, but with caveats. The CDP_HEADLESS configuration option controls whether the browser window is visible. However, headless Chrome often exhibits different fingerprinting characteristics than headed mode, and some advanced bot detection systems can identify headless operation. For maximum stealth, MediaCrawler recommends running CDP mode with headless=False to mimic a real user desktop environment.
Where does MediaCrawler store the stealth script?
The stealth anti-detection script is located at libs/stealth.min.js in the repository root. The CDPBrowserManager.add_stealth_script() method reads this file and injects it into each browser context upon creation, patching automation-detectable properties before any navigation occurs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →