# How to Debug Common MediaCrawler Issues: Browser Crashes, Login Failures, and Data Gaps

> Troubleshoot MediaCrawler problems like browser crashes, login errors, and missing data. Learn to debug CDP connections, QR code selectors, and async store writes effectively.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-06-29

---

**Debug MediaCrawler by validating CDP browser connections in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py), verifying QR-code canvas selectors during login workflows, and ensuring async store writes complete in [`store/douyin/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/_store_impl.py).**

MediaCrawler is a comprehensive data scraping framework that supports multiple Chinese social media platforms. When crawling complex sites like Zhihu, Douyin, or Xiaohongshu, developers frequently encounter three critical failure modes: browser instability, authentication timeouts, and incomplete data persistence. Understanding how to debug common MediaCrawler issues requires tracing the interaction between the Chrome DevTools Protocol (CDP) browser manager, platform-specific login handlers, and asynchronous storage implementations.

## Understanding MediaCrawler's Core Architecture

MediaCrawler relies on three tightly integrated subsystems that interact during every crawl session:

1. **Browser Management (CDP mode)** – Handled by `CDPBrowserManager` in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py). This class launches or attaches to Chrome/Edge instances via CDP, with cleanup handlers that guarantee browser process termination on exit.

2. **Login Workflow** – Exemplified by the Zhihu implementation in [`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py). The `ZhiHuLogin` class receives a `BrowserContext` from the manager and executes one of three strategies (`qrcode`, `phone`, or `cookie`), polling cookies via `utils.convert_cookies` until a valid session token appears.

3. **Data Persistence** – Provided by store implementations like `DouyinCsvStoreImplement`, `DouyinDbStoreImplement`, and `DouyinMongoStoreImplement` in [`store/douyin/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/_store_impl.py). Each store receives a `Dict` from the crawler and writes to CSV, relational databases, or MongoDB using the async `AsyncFileWriter` for non-blocking I/O.

## Debugging Browser Crashes and CDP Connection Failures

Browser crashes typically manifest as "Cannot connect to existing browser" errors. These failures originate in `CDPBrowserManager._connect_existing_browser` or `_launch_browser` when the CDP handshake fails.

### Verify Remote Debugging Configuration

Before running MediaCrawler, confirm your Chrome or Edge instance is running with remote debugging enabled:

1. Navigate to `chrome://inspect/#remote-debugging` in your browser.
2. Verify the debug port matches `config.CDP_DEBUG_PORT` (default **9222**).
3. Ensure Chrome version is **≥ 144** (required for CDP attach functionality).

### Check Connection Logic

The manager splits connection logic between `_connect_existing_browser` (attach to running browser) and `_launch_browser` + `_connect_via_cdp` (spawn fresh instance). Review the log output from `utils.logger` for messages prefixed with `[CDPBrowserManager]`—successful connections print "CDP port … is accessible" while failures trigger warnings.

### Force Fresh Browser Launch

If attaching to an existing browser fails, set `CDP_CONNECT_EXISTING=False` in your configuration to force a fresh launch. This bypasses port conflicts and zombie processes that may block the default port.

## Fixing Login Failures Across Platforms

Login failures occur in `ZhiHuLogin.login_by_qrcode` or `check_login_state` when the crawler cannot detect successful authentication.

### Validate QR Code Rendering

For QR-code-based logins, verify the canvas element exists in the browser's DOM:

1. Open DevTools console in the attached browser.
2. Run `document.querySelector("canvas.Qrcode-qrcode")`.
3. If the selector returns `null`, the crawler exits at lines 92–94 of [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) because `utils.find_qrcode_img_from_canvas` cannot locate the image.

### Inspect Cookie Extraction

Insert temporary debug output inside `check_login_state` after `utils.convert_cookies` to dump the cookie dictionary. Successful Zhihu logins require the presence of the `z_c0` session token. If this field is missing after the tenacity retry settings (line 51) expire, the login has timed out.

### Configuration Checks

- Verify `config.LOGIN_TYPE` matches your intended method (`qrcode`, `phone`, or `cookie`).
- Ensure `SAVE_LOGIN_STATE` is not unintentionally clearing valid cookies.
- Check that browser UI dialogs are not blocking confirmation prompts (see the FAQ section regarding browser popups).

## Resolving Data Gaps and Missing Rows

Data gaps typically stem from store implementations in [`store/douyin/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/_store_impl.py) when async writes fail or required dictionary keys are missing.

### Verify Async Writer Completion

For CSV stores, ensure `await self.file_writer.write_to_csv` completes before the crawler exits. The `AsyncFileWriter` guarantees non-blocking I/O, but premature termination truncates the write buffer.

### Validate Data Dictionary Keys

Missing keys in the `item` dict cause database `INSERT` failures. Before calling `store_content`, verify all required fields (`aweme_id`, `title`, `desc`, `author`) are present. Check for early exits in the crawler logic (e.g., `if not aweme_id: return`) that skip the storage phase.

### Database Commit Verification

For database stores, confirm `await session.commit()` executes successfully. Review logs for `logger.warning` messages that might swallow exceptions during the commit phase.

## Step-by-Step Debug Workflow

Follow this systematic approach to isolate failures:

1. **Enable Verbose Logging**
   
   ```python
   import config, utils
   config.LOG_LEVEL = "DEBUG"
   utils.logger.setLevel("DEBUG")
   ```

   
   The manager logs every stage with `[CDPBrowserManager]` prefix; login logs use `[ZhiHu]`.

2. **Validate CDP Connection**
   
   Run a minimal script that creates a `CDPBrowserManager` and prints `await mgr.get_browser_info()` after `launch_and_connect`. If `is_connected` is `False`, re-run with `CDP_CONNECT_EXISTING=False` to force a fresh launch.

3. **Inspect QR-Code Flow**
   
   After `utils.show_qrcode` displays the code, confirm the canvas element exists in DevTools. If the selector fails, the crawler exits at lines 92–94 of [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py).

4. **Confirm Cookie Extraction**
   
   Dump cookies inside `check_login_state` to verify the `z_c0` field presence. Absence indicates timeout or authentication failure.

5. **Check Data Store Writes**
   
   For CSV stores, verify the file header matches the keys in `content_item`. For DB stores, query directly (`SELECT * FROM douyin_aweme WHERE aweme_id = ?`) to confirm insertion versus updates.

6. **Resolve Environment Issues**
   
   - **Node.js missing**: Required for [`libs/stealth.min.js`](https://github.com/NanmiCoder/MediaCrawler/blob/main/libs/stealth.min.js). Install Node ≥ v16.
   - **Playwright timeout**: Increase `BROWSER_LAUNCH_TIMEOUT` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) or check proxy/VPN stability.
   - **Slide-captcha on Xiaohongshu**: Switch to CDP mode (`CDP_CONNECT_EXISTING=True`) or delete the `browser_data` folder to reset the session.

## Practical Code Examples

### Minimal CDP Connection Sanity Check

```python
import asyncio
from playwright.async_api import async_playwright
from tools.cdp_browser import CDPBrowserManager
from tools import utils
import config

async def main():
    async with async_playwright() as pw:
        mgr = CDPBrowserManager()
        ctx = await mgr.launch_and_connect(
            playwright=pw,
            playwright_proxy=None,
            user_agent=None,
            headless=False,               # set False to see the UI during debugging

        )
        info = await mgr.get_browser_info()
        utils.logger.info(f"Browser info: {info}")

        # clean up

        await mgr.cleanup(force=True)

asyncio.run(main())

```

*Key lines:* `launch_and_connect` (`cdp_browser.py#L97-L112`), `_get_browser_websocket_url` (`cdp_browser.py#L88-L110`).

### Debug-Enhanced Zhihu Login

```python
import asyncio
from tools.cdp_browser import CDPBrowserManager
from media_platform.zhihu.login import ZhiHuLogin
from tools import utils
import config

async def zhihu_login_demo():
    async with async_playwright() as pw:
        mgr = CDPBrowserManager()
        ctx = await mgr.launch_and_connect(pw, headless=False)
        page = await ctx.new_page()
        await page.goto("https://www.zhihu.com/signin")
        login = ZhiHuLogin(
            login_type="qrcode",
            browser_context=ctx,
            context_page=page,
            login_phone="",
            cookie_str=""
        )
        await login.begin()
        # after login you can dump cookies:

        cookies = await ctx.cookies()
        utils.logger.debug(f"Final cookies: {cookies}")

        await mgr.cleanup(force=True)

asyncio.run(zhihu_login_demo())

```

*Key lines:* `ZhiHuLogin.begin` (`login.py#L65-L75`), `login_by_qrcode` (`login.py#L81-L110`), `check_login_state` (`login.py#L52-L64`).

### Verifying Douyin Store Writes

```python
import asyncio
from store.douyin._store_impl import DouyinCsvStoreImplement
from tools import utils

async def test_store():
    store = DouyinCsvStoreImplement()
    sample = {
        "aweme_id": "1234567890",
        "title": "Demo video",
        "desc": "Test description",
        "author": "tester"
    }
    await store.store_content(sample)
    utils.logger.info("CSV write completed")

asyncio.run(test_store())

```

*Key lines:* `DouyinCsvStoreImplement.store_content` (`_store_impl.py#L50-L63`), async file writer usage (`AsyncFileWriter.write_to_csv`).

## Summary

- **Browser crashes** require verifying CDP port configuration (default 9222) and Chrome version compatibility (≥144) in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py).
- **Login failures** depend on canvas element detection and `z_c0` cookie extraction in [`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py).
- **Data gaps** result from missing dictionary keys or incomplete async writes in [`store/douyin/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/_store_impl.py).
- Enable `DEBUG` logging and use isolated test scripts to isolate subsystems before running full crawls.
- Reference `docs/常见问题.md` for platform-specific edge cases like slide-captcha handling.

## Frequently Asked Questions

### Why does MediaCrawler fail to connect to my existing Chrome browser?

The connection fails when `CDPBrowserManager._connect_existing_browser` cannot reach the WebSocket URL at `localhost:9222`. Verify Chrome is running with the `--remote-debugging-port=9222` flag, no other process occupies that port, and your Chrome version supports CDP attach (version 144 or higher). Set `CDP_CONNECT_EXISTING=False` to force a fresh browser instance if the existing process is unresponsive.

### Why does the QR code login timeout even when I scan it successfully?

The `check_login_state` method polls cookies every few seconds until it detects the `z_c0` token. If the timeout persists despite successful scanning, the canvas element may not be rendering correctly, or a confirmation dialog is blocking the UI. Check that `document.querySelector("canvas.Qrcode-qrcode")` returns a valid element in DevTools, and ensure no browser popups are waiting for interaction.

### Why are my CSV files empty or missing rows?

Empty CSVs occur when `AsyncFileWriter` does not flush before the program exits, or when the crawler encounters an early return (e.g., `if not aweme_id: return`) before reaching `store_content`. Verify that `content_item` dictionaries contain all required keys, and await the `write_to_csv` call explicitly. For database stores, confirm `await session.commit()` executes without exceptions.

### How do I fix the "Node.js missing" error when running in stealth mode?

The stealth script at [`libs/stealth.min.js`](https://github.com/NanmiCoder/MediaCrawler/blob/main/libs/stealth.min.js) requires Node.js to execute JavaScript evasion techniques. Install Node.js version 16 or higher and ensure the `node` executable is in your system PATH. This dependency is documented in `docs/常见问题.md` along with other environment-specific troubleshooting steps.