# MediaCrawler Login State Caching and Cookie Persistence Management: A Technical Guide

> Master MediaCrawler login state caching and cookie persistence. Learn how NanmiCoder/MediaCrawler avoids repeated logins by managing user data and Playwright cookies for seamless sessions.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**MediaCrawler maintains authenticated sessions across executions by persisting browser user data directories and converting Playwright cookies into HTTP header strings, eliminating repeated manual logins.**

MediaCrawler, an open-source multi-platform content crawler, implements a sophisticated approach to MediaCrawler login state caching and cookie persistence management that bridges browser automation with API-level requests. The system stores Chrome profiles on disk and utilizes utility functions to translate between Playwright's native cookie format and standard HTTP headers, ensuring seamless session continuity whether using standard Playwright or CDP (Chrome DevTools Protocol) modes.

## Configuration: Enabling Persistent Browser Profiles

The foundation of session persistence begins in [[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py#L52-L55), where the `SAVE_LOGIN_STATE` flag (default `True`) controls whether the crawler retains authentication data between runs. When enabled, MediaCrawler creates a dedicated user data directory at `browser_data/<platform>_user_data_dir` that stores cookies, local storage, and browser extensions.

According to the source code, this configuration applies globally across all supported platforms, allowing each crawler instance to maintain isolated profiles for sites like Zhihu, Xiaohongshu, or Weibo.

## Browser Launch: Persistent Contexts vs. CDP Mode

MediaCrawler supports two distinct browser automation strategies, both respecting the login state caching configuration.

### Standard Playwright Mode

In platform-specific core modules such as [[`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py#L35-L45), the `launch_browser` method checks `SAVE_LOGIN_STATE` to determine the launch strategy:

```python
if config.SAVE_LOGIN_STATE:
    user_data_dir = os.path.join(
        os.getcwd(),
        "browser_data",
        config.USER_DATA_DIR % config.PLATFORM,
    )
    context = await chromium.launch_persistent_context(
        user_data_dir=user_data_dir,
        headless=config.HEADLESS,
        proxy=playwright_proxy,
        viewport={"width": 1920, "height": 1080},
        user_agent=self.user_agent,
    )
else:
    browser = await chromium.launch(headless=config.HEADLESS, proxy=playwright_proxy)
    context = await browser.new_context(viewport={"width": 1920, "height": 1080})

```

### CDP (Chrome DevTools Protocol) Mode

For CDP-based crawling, [[`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py#L54-L63) implements similar persistence logic within `CDPBrowserManager._launch_browser`. When `SAVE_LOGIN_STATE` is active, the method passes the identical `user_data_dir` path to the Chrome launch arguments, ensuring CDP-controlled browsers reuse stored profiles rather than spawning fresh anonymous sessions.

## Cookie Extraction and Format Conversion

After successful authentication via QR code, phone number, or manual injection, MediaCrawler extracts cookies from the browser context using utility functions defined in [[`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py#L48-L57).

The `convert_browser_context_cookies` function retrieves cookies via `browser_context.cookies()` and transforms them into two formats simultaneously:

```python
cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
    browser_context,
    urls=self.cookie_urls,
)

```

This dual-format output provides a semicolon-separated string suitable for HTTP headers and a dictionary for programmatic access. The implementation handles domain-specific cookie retrieval by accepting a `urls` parameter, ensuring only relevant authentication tokens are extracted.

## Header Injection and Session Reuse

Once extracted, the cookie string integrates directly into the HTTP client's default headers. The platform-specific client stores the value in `self.default_headers["cookie"] = cookie_str`, ensuring all subsequent API requests automatically carry the persisted session credentials.

This mechanism bridges the gap between browser-based authentication and direct HTTP requests, allowing MediaCrawler to use authenticated sessions for high-performance API calls while maintaining the option to fall back to browser automation when necessary.

## Login State Verification

Each platform implements a `check_login_state` method to validate existing cookies before initiating new authentication flows. In [[`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py#L58-L63), the Zhihu crawler specifically checks for the presence of the `z_c0` cookie key:

```python
if not await self.check_login_state():
    # Trigger login flow

```

The verification process utilizes `utils.convert_cookies` to parse the current browser state, comparing extracted values against platform-specific session indicators. If valid credentials exist, the crawler skips the login process entirely, reducing execution time and avoiding unnecessary authentication challenges.

## Manual Cookie Injection

MediaCrawler supports pre-authenticated sessions through manual cookie input via the `COOKIES` configuration field. The `login_by_cookies` method in [[`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py#L15-L25) parses raw cookie strings using `convert_str_cookie_to_dict` and injects them into the browser context:

```python
for key, value in utils.convert_str_cookie_to_dict(self.cookie_str).items():
    await self.browser_context.add_cookies([{
        "name": key,
        "value": value,
        "domain": ".zhihu.com",
        "path": "/",
    }])

```

This approach enables users to supply cookies obtained from external sources or previous sessions, bypassing the QR-code or phone-based authentication workflows entirely.

## Automatic Re-Login and Session Recovery

When API requests fail authentication checks, MediaCrawler triggers an automatic recovery sequence. The client class in [[`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py#L64-L78) implements an `update_cookies` method synchronized with the `pong` verification check:

```python
if not await zhihu_client.pong():          # Internally calls check_login_state

    await zhihu_login.begin()              # Trigger fresh login flow

    await zhihu_client.update_cookies(self.browser_context)

```

This resilience mechanism ensures that expired sessions automatically refresh by re-running the authentication flow and updating both the in-memory cookie cache and the persistent browser profile on disk.

## Summary

- **Persistent Profiles**: MediaCrawler stores browser data in `browser_data/<platform>_user_data_dir` when `SAVE_LOGIN_STATE` is enabled, maintaining cookies and local storage across executions.
- **Dual Browser Support**: Both standard Playwright (`launch_persistent_context`) and CDP modes (`_launch_browser`) respect the user data directory configuration for consistent session handling.
- **Cookie Conversion**: The `convert_browser_context_cookies` utility translates Playwright cookie objects into HTTP header strings and lookup dictionaries.
- **Automatic Validation**: Platform-specific `check_login_state` methods verify session validity by checking for unique cookie keys (e.g., `z_c0` on Zhihu) before attempting new logins.
- **Flexible Authentication**: Support for manual cookie injection via `login_by_cookies` and automatic session recovery through `update_cookies` ensures robust long-running crawl operations.

## Frequently Asked Questions

### Where does MediaCrawler store persistent browser data?

MediaCrawler creates platform-specific directories under `browser_data/` (formatted as `<platform>_user_data_dir`) to store Chrome/Edge profiles. This location is defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and passed to `chromium.launch_persistent_context()` or CDP launch arguments, preserving cookies, local storage, and extension data between runs.

### How does MediaCrawler convert Playwright cookies to HTTP headers?

The `convert_browser_context_cookies` function in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py) calls `browser_context.cookies()` and formats the returned list into a semicolon-separated string (e.g., `key1=value1; key2=value2`). This string is stored in `default_headers["cookie"]` for API requests, while a dictionary version enables programmatic cookie access.

### What triggers automatic re-login in MediaCrawler?

The `pong()` method in platform-specific clients (such as [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py)) periodically validates the session via `check_login_state`. When validation fails—indicating expired or invalid cookies—the crawler automatically triggers the login flow again and calls `update_cookies` to refresh both the HTTP client headers and the persistent browser profile.

### Can I use existing cookies instead of scanning QR codes?

Yes. Set the `COOKIES` field in your configuration with a raw cookie string, and MediaCrawler will parse it using `convert_str_cookie_to_dict` before injecting values into the browser context via `add_cookies()`. This bypasses interactive login methods and immediately establishes an authenticated session.