MediaCrawler Login State Caching and Cookie Persistence Management: A Technical Guide

MediaCrawler maintains authenticated sessions across executions by persisting browser user data directories and converting Playwright cookies into HTTP header strings, eliminating repeated manual logins.

MediaCrawler, an open-source multi-platform content crawler, implements a sophisticated approach to MediaCrawler login state caching and cookie persistence management that bridges browser automation with API-level requests. The system stores Chrome profiles on disk and utilizes utility functions to translate between Playwright's native cookie format and standard HTTP headers, ensuring seamless session continuity whether using standard Playwright or CDP (Chrome DevTools Protocol) modes.

Configuration: Enabling Persistent Browser Profiles

The foundation of session persistence begins in [config/base_config.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py#L52-L55), where the SAVE_LOGIN_STATE flag (default True) controls whether the crawler retains authentication data between runs. When enabled, MediaCrawler creates a dedicated user data directory at browser_data/<platform>_user_data_dir that stores cookies, local storage, and browser extensions.

According to the source code, this configuration applies globally across all supported platforms, allowing each crawler instance to maintain isolated profiles for sites like Zhihu, Xiaohongshu, or Weibo.

Browser Launch: Persistent Contexts vs. CDP Mode

MediaCrawler supports two distinct browser automation strategies, both respecting the login state caching configuration.

Standard Playwright Mode

In platform-specific core modules such as [media_platform/zhihu/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py#L35-L45), the launch_browser method checks SAVE_LOGIN_STATE to determine the launch strategy:

if config.SAVE_LOGIN_STATE:
    user_data_dir = os.path.join(
        os.getcwd(),
        "browser_data",
        config.USER_DATA_DIR % config.PLATFORM,
    )
    context = await chromium.launch_persistent_context(
        user_data_dir=user_data_dir,
        headless=config.HEADLESS,
        proxy=playwright_proxy,
        viewport={"width": 1920, "height": 1080},
        user_agent=self.user_agent,
    )
else:
    browser = await chromium.launch(headless=config.HEADLESS, proxy=playwright_proxy)
    context = await browser.new_context(viewport={"width": 1920, "height": 1080})

CDP (Chrome DevTools Protocol) Mode

For CDP-based crawling, [tools/cdp_browser.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py#L54-L63) implements similar persistence logic within CDPBrowserManager._launch_browser. When SAVE_LOGIN_STATE is active, the method passes the identical user_data_dir path to the Chrome launch arguments, ensuring CDP-controlled browsers reuse stored profiles rather than spawning fresh anonymous sessions.

After successful authentication via QR code, phone number, or manual injection, MediaCrawler extracts cookies from the browser context using utility functions defined in [tools/crawler_util.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py#L48-L57).

The convert_browser_context_cookies function retrieves cookies via browser_context.cookies() and transforms them into two formats simultaneously:

cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
    browser_context,
    urls=self.cookie_urls,
)

This dual-format output provides a semicolon-separated string suitable for HTTP headers and a dictionary for programmatic access. The implementation handles domain-specific cookie retrieval by accepting a urls parameter, ensuring only relevant authentication tokens are extracted.

Header Injection and Session Reuse

Once extracted, the cookie string integrates directly into the HTTP client's default headers. The platform-specific client stores the value in self.default_headers["cookie"] = cookie_str, ensuring all subsequent API requests automatically carry the persisted session credentials.

This mechanism bridges the gap between browser-based authentication and direct HTTP requests, allowing MediaCrawler to use authenticated sessions for high-performance API calls while maintaining the option to fall back to browser automation when necessary.

Login State Verification

Each platform implements a check_login_state method to validate existing cookies before initiating new authentication flows. In [media_platform/zhihu/login.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py#L58-L63), the Zhihu crawler specifically checks for the presence of the z_c0 cookie key:

if not await self.check_login_state():
    # Trigger login flow

The verification process utilizes utils.convert_cookies to parse the current browser state, comparing extracted values against platform-specific session indicators. If valid credentials exist, the crawler skips the login process entirely, reducing execution time and avoiding unnecessary authentication challenges.

MediaCrawler supports pre-authenticated sessions through manual cookie input via the COOKIES configuration field. The login_by_cookies method in [media_platform/zhihu/login.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py#L15-L25) parses raw cookie strings using convert_str_cookie_to_dict and injects them into the browser context:

for key, value in utils.convert_str_cookie_to_dict(self.cookie_str).items():
    await self.browser_context.add_cookies([{
        "name": key,
        "value": value,
        "domain": ".zhihu.com",
        "path": "/",
    }])

This approach enables users to supply cookies obtained from external sources or previous sessions, bypassing the QR-code or phone-based authentication workflows entirely.

Automatic Re-Login and Session Recovery

When API requests fail authentication checks, MediaCrawler triggers an automatic recovery sequence. The client class in [media_platform/zhihu/client.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py#L64-L78) implements an update_cookies method synchronized with the pong verification check:

if not await zhihu_client.pong():          # Internally calls check_login_state

    await zhihu_login.begin()              # Trigger fresh login flow

    await zhihu_client.update_cookies(self.browser_context)

This resilience mechanism ensures that expired sessions automatically refresh by re-running the authentication flow and updating both the in-memory cookie cache and the persistent browser profile on disk.

Summary

  • Persistent Profiles: MediaCrawler stores browser data in browser_data/<platform>_user_data_dir when SAVE_LOGIN_STATE is enabled, maintaining cookies and local storage across executions.
  • Dual Browser Support: Both standard Playwright (launch_persistent_context) and CDP modes (_launch_browser) respect the user data directory configuration for consistent session handling.
  • Cookie Conversion: The convert_browser_context_cookies utility translates Playwright cookie objects into HTTP header strings and lookup dictionaries.
  • Automatic Validation: Platform-specific check_login_state methods verify session validity by checking for unique cookie keys (e.g., z_c0 on Zhihu) before attempting new logins.
  • Flexible Authentication: Support for manual cookie injection via login_by_cookies and automatic session recovery through update_cookies ensures robust long-running crawl operations.

Frequently Asked Questions

Where does MediaCrawler store persistent browser data?

MediaCrawler creates platform-specific directories under browser_data/ (formatted as <platform>_user_data_dir) to store Chrome/Edge profiles. This location is defined in config/base_config.py and passed to chromium.launch_persistent_context() or CDP launch arguments, preserving cookies, local storage, and extension data between runs.

How does MediaCrawler convert Playwright cookies to HTTP headers?

The convert_browser_context_cookies function in tools/crawler_util.py calls browser_context.cookies() and formats the returned list into a semicolon-separated string (e.g., key1=value1; key2=value2). This string is stored in default_headers["cookie"] for API requests, while a dictionary version enables programmatic cookie access.

What triggers automatic re-login in MediaCrawler?

The pong() method in platform-specific clients (such as media_platform/zhihu/client.py) periodically validates the session via check_login_state. When validation fails—indicating expired or invalid cookies—the crawler automatically triggers the login flow again and calls update_cookies to refresh both the HTTP client headers and the persistent browser profile.

Can I use existing cookies instead of scanning QR codes?

Yes. Set the COOKIES field in your configuration with a raw cookie string, and MediaCrawler will parse it using convert_str_cookie_to_dict before injecting values into the browser context via add_cookies(). This bypasses interactive login methods and immediately establishes an authenticated session.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →