# How to Handle Login State Persistence and Cookie-Based Authentication in MediaCrawler

> Learn how MediaCrawler manages login state persistence and cookie authentication using persistent user data directories, cookie conversion, and platform-specific verification.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-14

---

**MediaCrawler handles login state persistence through a combination of persistent browser user data directories, bidirectional cookie conversion utilities, and platform-specific login verification classes.**

To avoid repetitive logins across crawler executions, the repository implements a layered authentication system. The core mechanism relies on Playwright's `launch_persistent_context()` and custom conversion utilities in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py). Below is a complete breakdown of how the system works and how to configure it for your own scraping workflows.

---

## Enabling Persistent Login State in Configuration

Login state caching is controlled by a single boolean flag in the global configuration. When enabled, MediaCrawler creates a dedicated folder on disk that Chrome or Edge treats as its profile directory—preserving cookies, localStorage, and extension states between runs.

- **Configuration flag:** `SAVE_LOGIN_STATE` (default: `True`)
- **Storage path:** `browser_data/<platform>_user_data_dir`

```python

# config/base_config.py (lines 52-55)

SAVE_LOGIN_STATE = True  # Enable persistent browser profile caching

USER_DATA_DIR = "%s_user_data_dir"  # Platform-specific subdirectory naming

```

The `USER_DATA_DIR` template string is formatted with the platform name (e.g., `zhihu`, `xiaohongshu`) to isolate profiles per target site.

---

## Launching Browsers with Persistent Context

MediaCrawler supports two browser control modes, both respecting the `SAVE_LOGIN_STATE` flag.

### Playwright Mode

In standard Playwright operation, `launch_persistent_context()` is invoked when state persistence is enabled. This returns a `BrowserContext` that automatically reads from and writes to the specified directory.

```python

# media_platform/zhihu/core.py (simplified from lines 35-45)

import os
from playwright.async_api import async_playwright
from config import base_config as config

user_data_dir = os.path.join(
    os.getcwd(),
    "browser_data",
    config.USER_DATA_DIR % config.PLATFORM,  # e.g., "zhihu_user_data_dir"

)

if config.SAVE_LOGIN_STATE:
    context = await chromium.launch_persistent_context(
        user_data_dir=user_data_dir,
        headless=config.HEADLESS,
        proxy=playwright_proxy,
        viewport={"width": 1920, "height": 1080},
        user_agent=self.user_agent,
    )
else:
    browser = await chromium.launch(headless=config.HEADLESS, proxy=playwright_proxy)
    context = await browser.new_context(viewport={"width": 1920, "height": 1080})

```

### CDP Mode (Chrome DevTools Protocol)

For attaching to existing browser instances, `CDPBrowserManager` passes the same `user_data_dir` parameter when launching Chrome with remote debugging enabled.

```python

# tools/cdp_browser.py (lines 54-63)

def _launch_browser(self):
    args = [
        f"--remote-debugging-port={self.debug_port}",
        f"--user-data-dir={self.user_data_dir}" if config.SAVE_LOGIN_STATE else "",
        # additional chromium flags...

    ]
    subprocess.Popen([self.chrome_path] + [a for a in args if a])

```

This ensures that even CDP-controlled sessions retain authentication state across process restarts.

---

## Extracting and Converting Cookies

Once a user authenticates—whether via QR code, SMS, or manual flow—the crawler extracts cookies from the browser context for reuse in HTTP API calls.

### The Cookie Conversion Pipeline

The `convert_browser_context_cookies()` function in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py) bridges Playwright's native cookie format and standard HTTP header strings:

```python

# tools/crawler_util.py (lines 48-57)

async def convert_browser_context_cookies(browser_context, urls=None):
    cookies = await browser_context.cookies(urls)  # Returns list of dicts

    cookie_str = "; ".join(f"{c['name']}={c['value']}" for c in cookies)
    cookie_dict = {c["name"]: c["value"] for c in cookies}
    return cookie_str, cookie_dict

```

**Return values:**
- `cookie_str`: Semicolon-separated string ready for the `Cookie:` HTTP header
- `cookie_dict`: Key-value mapping for programmatic access

### Injecting Cookies into API Requests

After extraction, cookies are attached to the platform client's default headers:

```python

# Typical usage in platform core classes

cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
    self.browser_context,
    urls=self.cookie_urls,  # e.g., ["https://www.zhihu.com"]

)
self.default_headers["cookie"] = cookie_str
self.cookie_dict = cookie_dict

```

All subsequent API calls through the client automatically carry this session identifier.

---

## Verifying Existing Login State

Each platform implements a `check_login_state()` method to test whether cached credentials remain valid without triggering a full login flow.

### Platform-Specific Verification

For Zhihu, the login class checks for the presence of the `z_c0` authentication cookie:

```python

# media_platform/zhihu/login.py (lines 58-63)

async def check_login_state(self):
    cookie_str, cookie_dict = await utils.convert_cookies(self.browser_context)
    if cookie_dict.get("z_c0"):
        logger.info("[ZhiHuLogin] Valid login state detected via z_c0 cookie")
        return True
    return False

```

The `convert_cookies` helper (aliased in [`crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/crawler_util.py)) handles the underlying context extraction.

### Client-Level "Pong" Check

The platform client exposes a `pong()` method that wraps this verification, enabling automatic re-authentication:

```python

# media_platform/zhihu/client.py (lines 64-78)

async def pong(self):
    """Verify session validity; return False if re-login required."""
    return await self.login_obj.check_login_state()

async def update_cookies(self, browser_context):
    """Refresh stored cookies after re-authentication."""
    cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
        browser_context,
        urls=self.cookie_urls,
    )
    self.headers["cookie"] = cookie_str
    self.cookie_dict = cookie_dict

```

---

## Injecting External Cookies for Cookie-Based Login

Users can supply raw cookie strings via configuration to bypass interactive authentication entirely.

### Parsing and Loading External Cookies

The `COOKIES` configuration value is parsed and injected into the browser context before any navigation:

```python

# media_platform/zhihu/login.py (lines 15-25)

async def login_by_cookies(self):
    """Inject user-supplied cookies into browser context."""
    cookie_dict = utils.convert_str_cookie_to_dict(self.cookie_str)
    # Format: "key1=value1; key2=value2" → {"key1": "value1", "key2": "value2"}

    
    for key, value in cookie_dict.items():
        await self.browser_context.add_cookies([{
            "name": key,
            "value": value,
            "domain": ".zhihu.com",  # Platform-specific domain

            "path": "/",
        }])

```

The `convert_str_cookie_to_dict()` utility handles the string-to-mapping conversion, normalizing whitespace and duplicate delimiters.

---

## Automatic Re-Login and Cookie Refresh

When a session expires during crawling, MediaCrawler triggers the full authentication workflow and synchronizes the persistent storage:

```python

# Typical orchestration pattern in crawler execution

if not await zhihu_client.pong():           # Session check fails

    await zhihu_login.begin()               # QR-code or cookie re-login

    await zhihu_client.update_cookies(      # Sync new cookies to client

        self.browser_context
    )
    # Persistent context automatically saves to user_data_dir

```

This loop ensures that the browser profile on disk always reflects the freshest valid session.

---

## Summary

- **Enable persistence** with `SAVE_LOGIN_STATE = True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)
- **Store profiles** per platform using the `USER_DATA_DIR` template
- **Launch browsers** via `chromium.launch_persistent_context()` or CDP with `--user-data-dir`
- **Convert cookies** with `convert_browser_context_cookies()` for HTTP reuse
- **Verify sessions** through platform-specific `check_login_state()` methods
- **Inject external cookies** using `convert_str_cookie_to_dict()` and `browser_context.add_cookies()`
- **Refresh automatically** via `update_cookies()` when `pong()` detects expiry

---

## Frequently Asked Questions

### Where does MediaCrawler store the persistent browser profile?

MediaCrawler stores profiles in `browser_data/<platform>_user_data_dir` relative to the working directory, as configured by `USER_DATA_DIR` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). This folder contains the full Chrome/Edge user data including cookies, localStorage, and cache.

### Can I use cookie-based authentication without browser automation?

Yes. Set the `COOKIES` configuration value to a semicolon-separated string, then call `login_by_cookies()` which parses and injects them via `browser_context.add_cookies()`. The crawler still launches a browser instance, but skips interactive authentication flows.

### How does MediaCrawler detect if saved cookies have expired?

Each platform's `check_login_state()` method checks for signature cookies (e.g., `z_c0` for Zhihu). The client calls `pong()` to verify validity; if absent or invalid, the crawler triggers `begin()` to re-authenticate and refreshes stored cookies with `update_cookies()`.

### What is the difference between Playwright and CDP modes for login persistence?

Both modes respect `SAVE_LOGIN_STATE`. Playwright uses `launch_persistent_context()` internally, while CDP mode passes `--user-data-dir` as a Chrome command-line argument. CDP attaches to an existing or newly launched process; Playwright manages the full lifecycle.