How to Use Cookie-Based Login in MediaCrawler: A Complete Guide

To use cookie-based login in MediaCrawler, set LOGIN_TYPE to "cookie" and provide your session cookies via the --cookies CLI argument or the COOKIES configuration variable, allowing the crawler to inject them directly into the Playwright browser context and bypass interactive authentication.

MediaCrawler by NanmiCoder supports multiple authentication methods, but cookie-based login offers the fastest way to start crawling without scanning QR codes or entering phone verification codes. When you have valid session cookies from platforms like Zhihu, XiaoHongShu, or Douyin, you can configure the crawler to use cookie-based login and immediately access authenticated content. This guide explains the complete implementation, from CLI arguments to the underlying convert_str_cookie_to_dict utility in tools/crawler_util.py.

The cookie authentication flow relies on converting raw cookie strings into Playwright-compatible dictionaries. According to the MediaCrawler source code, the process involves three core components: the configuration layer in config/base_config.py, the conversion utility in tools/crawler_util.py, and the platform-specific login handlers.

In tools/crawler_util.py, the convert_str_cookie_to_dict function (lines 59-74) parses semicolon-separated cookie strings into dictionaries suitable for Playwright's browser_context.add_cookies method. The function handles the transformation from HTTP header format (key=value; key2=value2) to the structured format required by the browser automation framework.

Platform-specific login classes implement the login_by_cookies() method to inject these cookies:

Each method adds cookies to the browser context with the appropriate domain (e.g., .zhihu.com, .xiaohongshu.com, or .douyin.com).

You can enable cookie-based login using either the global configuration file or command-line arguments defined in cmd_arg/arg.py.

Method 1: Configuration File

Edit config/base_config.py to set the login type and cookie string:


# config/base_config.py

LOGIN_TYPE = "cookie"
COOKIES = "z_c0=your_token_here; d_c0=another_cookie"

Method 2: CLI Arguments

The CLI parser in cmd_arg/arg.py exposes the --lt (login type) and --cookies flags:

python -m main --platform zhihu --lt cookie --cookies "z_c0=abc123; d_c0=def456"

Each platform requires specific cookie names and domains. Here are the implementation details for the three primary platforms.

For Zhihu, the ZhiHuLogin.login_by_cookies() method in media_platform/zhihu/login.py expects cookies like z_c0 and d_c0:

python -m main \
  --platform zhihu \
  --lt cookie \
  --cookies "z_c0=abc123; d_c0=def456"

Internally, the method iterates over the converted cookie dictionary and adds each cookie to the Playwright context with the domain .zhihu.com.

The XiaoHongShuLogin class in media_platform/xhs/login.py handles the web_session cookie. The domain is selected dynamically based on config.XHS_INTERNATIONAL:

python -m main \
  --platform xhs \
  --lt cookie \
  --cookies "web_session=YOUR_XHS_SESSION"

The cookies are injected into either .xiaohongshu.com or .rednote.com depending on your configuration.

For Douyin, implemented in media_platform/douyin/login.py, you typically need sessionid and LOGIN_STATUS:

python -m main \
  --platform dy \
  --lt cookie \
  --cookies "sessionid=YOUR_SESSION; LOGIN_STATUS=1"

The login_by_cookies() method adds these to the .douyin.com domain. The LOGIN_STATUS=1 cookie satisfies the check_login_state() validation, allowing immediate crawling.

Using the Python API Directly

To bypass the CLI entirely, instantiate platform login classes directly with the cookie_str parameter:

from media_platform.zhihu.login import ZhiHuLogin
from playwright.async_api import async_playwright

async def run_cookie_login():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=False)
        context = await browser.new_context()
        page = await context.new_page()
        
        login = ZhiHuLogin(
            login_type="cookie",
            browser_context=context,
            context_page=page,
            cookie_str="z_c0=abc123; d_c0=def456"
        )
        await login.begin()  # Injects cookies and starts crawling

This approach uses the same convert_str_cookie_to_dict utility internally but gives you full control over the browser context initialization.

Summary

  • Configure via CLI: Use --lt cookie --cookies "key=value" arguments parsed by cmd_arg/arg.py
  • Configure via file: Set LOGIN_TYPE = "cookie" and COOKIES string in config/base_config.py
  • Conversion utility: tools/crawler_util.py contains convert_str_cookie_to_dict to transform cookie strings for Playwright
  • Platform implementations: Each platform (Zhihu, XHS, Douyin) implements login_by_cookies() in their respective login.py files
  • Domain specificity: Cookies are scoped to platform-specific domains (.zhihu.com, .xiaohongshu.com, .douyin.com)
  • Cookie refresh: Use convert_browser_context_cookies in tools/crawler_util.py (lines 48-56) to retrieve updated cookies from the browser context after login

Frequently Asked Questions

MediaCrawler expects a semicolon-separated string following HTTP cookie header conventions, such as key1=value1; key2=value2. The convert_str_cookie_to_dict function in tools/crawler_util.py parses this format into dictionaries required by Playwright's browser_context.add_cookies method.

Where does MediaCrawler store cookies after injection?

After parsing, cookies are injected directly into the Playwright BrowserContext using browser_context.add_cookies. For session persistence, the convert_browser_context_cookies function (lines 48-56 in tools/crawler_util.py) can retrieve current cookies from the browser context to update HTTP headers.

Yes. When LOGIN_TYPE is set to cookie and valid cookies are provided via the --cookies argument or config/base_config.py, MediaCrawler skips the QR code and phone verification flows entirely. The platform-specific login_by_cookies() method proceeds directly to the crawling workflow after injecting cookies.

Which cookies are required for each platform?

Requirements vary by platform implementation:

  • Zhihu: Typically requires z_c0 for authentication and d_c0 for device tracking
  • XiaoHongShu: Primarily requires the web_session cookie
  • Douyin: Requires sessionid and LOGIN_STATUS=1 to satisfy check_login_state() validation

Check the specific platform's login.py file (e.g., media_platform/zhihu/login.py) to verify required cookie names for your target platform.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →