How to Handle Login State Persistence and Cookie-Based Authentication in MediaCrawler

MediaCrawler handles login state persistence through a combination of persistent browser user data directories, bidirectional cookie conversion utilities, and platform-specific login verification classes.

To avoid repetitive logins across crawler executions, the repository implements a layered authentication system. The core mechanism relies on Playwright's launch_persistent_context() and custom conversion utilities in tools/crawler_util.py. Below is a complete breakdown of how the system works and how to configure it for your own scraping workflows.


Enabling Persistent Login State in Configuration

Login state caching is controlled by a single boolean flag in the global configuration. When enabled, MediaCrawler creates a dedicated folder on disk that Chrome or Edge treats as its profile directory—preserving cookies, localStorage, and extension states between runs.

  • Configuration flag: SAVE_LOGIN_STATE (default: True)
  • Storage path: browser_data/<platform>_user_data_dir

# config/base_config.py (lines 52-55)

SAVE_LOGIN_STATE = True  # Enable persistent browser profile caching

USER_DATA_DIR = "%s_user_data_dir"  # Platform-specific subdirectory naming

The USER_DATA_DIR template string is formatted with the platform name (e.g., zhihu, xiaohongshu) to isolate profiles per target site.


Launching Browsers with Persistent Context

MediaCrawler supports two browser control modes, both respecting the SAVE_LOGIN_STATE flag.

Playwright Mode

In standard Playwright operation, launch_persistent_context() is invoked when state persistence is enabled. This returns a BrowserContext that automatically reads from and writes to the specified directory.


# media_platform/zhihu/core.py (simplified from lines 35-45)

import os
from playwright.async_api import async_playwright
from config import base_config as config

user_data_dir = os.path.join(
    os.getcwd(),
    "browser_data",
    config.USER_DATA_DIR % config.PLATFORM,  # e.g., "zhihu_user_data_dir"

)

if config.SAVE_LOGIN_STATE:
    context = await chromium.launch_persistent_context(
        user_data_dir=user_data_dir,
        headless=config.HEADLESS,
        proxy=playwright_proxy,
        viewport={"width": 1920, "height": 1080},
        user_agent=self.user_agent,
    )
else:
    browser = await chromium.launch(headless=config.HEADLESS, proxy=playwright_proxy)
    context = await browser.new_context(viewport={"width": 1920, "height": 1080})

CDP Mode (Chrome DevTools Protocol)

For attaching to existing browser instances, CDPBrowserManager passes the same user_data_dir parameter when launching Chrome with remote debugging enabled.


# tools/cdp_browser.py (lines 54-63)

def _launch_browser(self):
    args = [
        f"--remote-debugging-port={self.debug_port}",
        f"--user-data-dir={self.user_data_dir}" if config.SAVE_LOGIN_STATE else "",
        # additional chromium flags...

    ]
    subprocess.Popen([self.chrome_path] + [a for a in args if a])

This ensures that even CDP-controlled sessions retain authentication state across process restarts.


Extracting and Converting Cookies

Once a user authenticates—whether via QR code, SMS, or manual flow—the crawler extracts cookies from the browser context for reuse in HTTP API calls.

The convert_browser_context_cookies() function in tools/crawler_util.py bridges Playwright's native cookie format and standard HTTP header strings:


# tools/crawler_util.py (lines 48-57)

async def convert_browser_context_cookies(browser_context, urls=None):
    cookies = await browser_context.cookies(urls)  # Returns list of dicts

    cookie_str = "; ".join(f"{c['name']}={c['value']}" for c in cookies)
    cookie_dict = {c["name"]: c["value"] for c in cookies}
    return cookie_str, cookie_dict

Return values:

  • cookie_str: Semicolon-separated string ready for the Cookie: HTTP header
  • cookie_dict: Key-value mapping for programmatic access

Injecting Cookies into API Requests

After extraction, cookies are attached to the platform client's default headers:


# Typical usage in platform core classes

cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
    self.browser_context,
    urls=self.cookie_urls,  # e.g., ["https://www.zhihu.com"]

)
self.default_headers["cookie"] = cookie_str
self.cookie_dict = cookie_dict

All subsequent API calls through the client automatically carry this session identifier.


Verifying Existing Login State

Each platform implements a check_login_state() method to test whether cached credentials remain valid without triggering a full login flow.

Platform-Specific Verification

For Zhihu, the login class checks for the presence of the z_c0 authentication cookie:


# media_platform/zhihu/login.py (lines 58-63)

async def check_login_state(self):
    cookie_str, cookie_dict = await utils.convert_cookies(self.browser_context)
    if cookie_dict.get("z_c0"):
        logger.info("[ZhiHuLogin] Valid login state detected via z_c0 cookie")
        return True
    return False

The convert_cookies helper (aliased in crawler_util.py) handles the underlying context extraction.

Client-Level "Pong" Check

The platform client exposes a pong() method that wraps this verification, enabling automatic re-authentication:


# media_platform/zhihu/client.py (lines 64-78)

async def pong(self):
    """Verify session validity; return False if re-login required."""
    return await self.login_obj.check_login_state()

async def update_cookies(self, browser_context):
    """Refresh stored cookies after re-authentication."""
    cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
        browser_context,
        urls=self.cookie_urls,
    )
    self.headers["cookie"] = cookie_str
    self.cookie_dict = cookie_dict

Users can supply raw cookie strings via configuration to bypass interactive authentication entirely.

Parsing and Loading External Cookies

The COOKIES configuration value is parsed and injected into the browser context before any navigation:


# media_platform/zhihu/login.py (lines 15-25)

async def login_by_cookies(self):
    """Inject user-supplied cookies into browser context."""
    cookie_dict = utils.convert_str_cookie_to_dict(self.cookie_str)
    # Format: "key1=value1; key2=value2" → {"key1": "value1", "key2": "value2"}

    
    for key, value in cookie_dict.items():
        await self.browser_context.add_cookies([{
            "name": key,
            "value": value,
            "domain": ".zhihu.com",  # Platform-specific domain

            "path": "/",
        }])

The convert_str_cookie_to_dict() utility handles the string-to-mapping conversion, normalizing whitespace and duplicate delimiters.


When a session expires during crawling, MediaCrawler triggers the full authentication workflow and synchronizes the persistent storage:


# Typical orchestration pattern in crawler execution

if not await zhihu_client.pong():           # Session check fails

    await zhihu_login.begin()               # QR-code or cookie re-login

    await zhihu_client.update_cookies(      # Sync new cookies to client

        self.browser_context
    )
    # Persistent context automatically saves to user_data_dir

This loop ensures that the browser profile on disk always reflects the freshest valid session.


Summary

  • Enable persistence with SAVE_LOGIN_STATE = True in config/base_config.py
  • Store profiles per platform using the USER_DATA_DIR template
  • Launch browsers via chromium.launch_persistent_context() or CDP with --user-data-dir
  • Convert cookies with convert_browser_context_cookies() for HTTP reuse
  • Verify sessions through platform-specific check_login_state() methods
  • Inject external cookies using convert_str_cookie_to_dict() and browser_context.add_cookies()
  • Refresh automatically via update_cookies() when pong() detects expiry

Frequently Asked Questions

Where does MediaCrawler store the persistent browser profile?

MediaCrawler stores profiles in browser_data/<platform>_user_data_dir relative to the working directory, as configured by USER_DATA_DIR in config/base_config.py. This folder contains the full Chrome/Edge user data including cookies, localStorage, and cache.

Yes. Set the COOKIES configuration value to a semicolon-separated string, then call login_by_cookies() which parses and injects them via browser_context.add_cookies(). The crawler still launches a browser instance, but skips interactive authentication flows.

How does MediaCrawler detect if saved cookies have expired?

Each platform's check_login_state() method checks for signature cookies (e.g., z_c0 for Zhihu). The client calls pong() to verify validity; if absent or invalid, the crawler triggers begin() to re-authenticate and refreshes stored cookies with update_cookies().

What is the difference between Playwright and CDP modes for login persistence?

Both modes respect SAVE_LOGIN_STATE. Playwright uses launch_persistent_context() internally, while CDP mode passes --user-data-dir as a Chrome command-line argument. CDP attaches to an existing or newly launched process; Playwright manages the full lifecycle.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →