# How to Use Cookie-Based Login in MediaCrawler: A Complete Guide

> Master cookie-based login in MediaCrawler with this guide. Learn to set LOGIN_TYPE and provide cookies via CLI or config for seamless authentication bypass. Get started today.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-01

---

**To use cookie-based login in MediaCrawler, set `LOGIN_TYPE` to `"cookie"` and provide your session cookies via the `--cookies` CLI argument or the `COOKIES` configuration variable, allowing the crawler to inject them directly into the Playwright browser context and bypass interactive authentication.**

MediaCrawler by NanmiCoder supports multiple authentication methods, but cookie-based login offers the fastest way to start crawling without scanning QR codes or entering phone verification codes. When you have valid session cookies from platforms like Zhihu, XiaoHongShu, or Douyin, you can configure the crawler to use cookie-based login and immediately access authenticated content. This guide explains the complete implementation, from CLI arguments to the underlying `convert_str_cookie_to_dict` utility in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py).

## How Cookie-Based Login Works in MediaCrawler

The cookie authentication flow relies on converting raw cookie strings into Playwright-compatible dictionaries. According to the MediaCrawler source code, the process involves three core components: the configuration layer in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), the conversion utility in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py), and the platform-specific login handlers.

### The Cookie Conversion Pipeline

In [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py), the `convert_str_cookie_to_dict` function (lines 59-74) parses semicolon-separated cookie strings into dictionaries suitable for Playwright's `browser_context.add_cookies` method. The function handles the transformation from HTTP header format (`key=value; key2=value2`) to the structured format required by the browser automation framework.

Platform-specific login classes implement the `login_by_cookies()` method to inject these cookies:
- **ZhiHuLogin** in [`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py) (lines 73-82)
- **XiaoHongShuLogin** in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) (lines 71-78)
- **DouYinLogin** in [`media_platform/douyin/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/login.py) (lines 66-74)

Each method adds cookies to the browser context with the appropriate domain (e.g., `.zhihu.com`, `.xiaohongshu.com`, or `.douyin.com`).

## Configuring Cookie-Based Login via Command Line

You can enable cookie-based login using either the global configuration file or command-line arguments defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py).

### Method 1: Configuration File

Edit [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to set the login type and cookie string:

```python

# config/base_config.py

LOGIN_TYPE = "cookie"
COOKIES = "z_c0=your_token_here; d_c0=another_cookie"

```

### Method 2: CLI Arguments

The CLI parser in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) exposes the `--lt` (login type) and `--cookies` flags:

```bash
python -m main --platform zhihu --lt cookie --cookies "z_c0=abc123; d_c0=def456"

```

## Platform-Specific Cookie Login Examples

Each platform requires specific cookie names and domains. Here are the implementation details for the three primary platforms.

### Zhihu Cookie Login

For Zhihu, the `ZhiHuLogin.login_by_cookies()` method in [`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py) expects cookies like `z_c0` and `d_c0`:

```bash
python -m main \
  --platform zhihu \
  --lt cookie \
  --cookies "z_c0=abc123; d_c0=def456"

```

Internally, the method iterates over the converted cookie dictionary and adds each cookie to the Playwright context with the domain `.zhihu.com`.

### XiaoHongShu (XHS) Cookie Login

The `XiaoHongShuLogin` class in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) handles the `web_session` cookie. The domain is selected dynamically based on `config.XHS_INTERNATIONAL`:

```bash
python -m main \
  --platform xhs \
  --lt cookie \
  --cookies "web_session=YOUR_XHS_SESSION"

```

The cookies are injected into either `.xiaohongshu.com` or `.rednote.com` depending on your configuration.

### Douyin Cookie Login

For Douyin, implemented in [`media_platform/douyin/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/login.py), you typically need `sessionid` and `LOGIN_STATUS`:

```bash
python -m main \
  --platform dy \
  --lt cookie \
  --cookies "sessionid=YOUR_SESSION; LOGIN_STATUS=1"

```

The `login_by_cookies()` method adds these to the `.douyin.com` domain. The `LOGIN_STATUS=1` cookie satisfies the `check_login_state()` validation, allowing immediate crawling.

## Using the Python API Directly

To bypass the CLI entirely, instantiate platform login classes directly with the `cookie_str` parameter:

```python
from media_platform.zhihu.login import ZhiHuLogin
from playwright.async_api import async_playwright

async def run_cookie_login():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=False)
        context = await browser.new_context()
        page = await context.new_page()
        
        login = ZhiHuLogin(
            login_type="cookie",
            browser_context=context,
            context_page=page,
            cookie_str="z_c0=abc123; d_c0=def456"
        )
        await login.begin()  # Injects cookies and starts crawling

```

This approach uses the same `convert_str_cookie_to_dict` utility internally but gives you full control over the browser context initialization.

## Summary

- **Configure via CLI**: Use `--lt cookie --cookies "key=value"` arguments parsed by [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py)
- **Configure via file**: Set `LOGIN_TYPE = "cookie"` and `COOKIES` string in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)
- **Conversion utility**: [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py) contains `convert_str_cookie_to_dict` to transform cookie strings for Playwright
- **Platform implementations**: Each platform (Zhihu, XHS, Douyin) implements `login_by_cookies()` in their respective [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) files
- **Domain specificity**: Cookies are scoped to platform-specific domains (`.zhihu.com`, `.xiaohongshu.com`, `.douyin.com`)
- **Cookie refresh**: Use `convert_browser_context_cookies` in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py) (lines 48-56) to retrieve updated cookies from the browser context after login

## Frequently Asked Questions

### What format should the cookie string be in?

MediaCrawler expects a semicolon-separated string following HTTP cookie header conventions, such as `key1=value1; key2=value2`. The `convert_str_cookie_to_dict` function in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py) parses this format into dictionaries required by Playwright's `browser_context.add_cookies` method.

### Where does MediaCrawler store cookies after injection?

After parsing, cookies are injected directly into the Playwright `BrowserContext` using `browser_context.add_cookies`. For session persistence, the `convert_browser_context_cookies` function (lines 48-56 in [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py)) can retrieve current cookies from the browser context to update HTTP headers.

### Can I use cookie-based login to bypass QR code scanning?

Yes. When `LOGIN_TYPE` is set to `cookie` and valid cookies are provided via the `--cookies` argument or [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), MediaCrawler skips the QR code and phone verification flows entirely. The platform-specific `login_by_cookies()` method proceeds directly to the crawling workflow after injecting cookies.

### Which cookies are required for each platform?

Requirements vary by platform implementation:
- **Zhihu**: Typically requires `z_c0` for authentication and `d_c0` for device tracking
- **XiaoHongShu**: Primarily requires the `web_session` cookie
- **Douyin**: Requires `sessionid` and `LOGIN_STATUS=1` to satisfy `check_login_state()` validation

Check the specific platform's [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) file (e.g., [`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py)) to verify required cookie names for your target platform.