How to Use Cookie-Based Login in MediaCrawler: A Complete Guide
To use cookie-based login in MediaCrawler, set LOGIN_TYPE to "cookie" and provide your session cookies via the --cookies CLI argument or the COOKIES configuration variable, allowing the crawler to inject them directly into the Playwright browser context and bypass interactive authentication.
MediaCrawler by NanmiCoder supports multiple authentication methods, but cookie-based login offers the fastest way to start crawling without scanning QR codes or entering phone verification codes. When you have valid session cookies from platforms like Zhihu, XiaoHongShu, or Douyin, you can configure the crawler to use cookie-based login and immediately access authenticated content. This guide explains the complete implementation, from CLI arguments to the underlying convert_str_cookie_to_dict utility in tools/crawler_util.py.
How Cookie-Based Login Works in MediaCrawler
The cookie authentication flow relies on converting raw cookie strings into Playwright-compatible dictionaries. According to the MediaCrawler source code, the process involves three core components: the configuration layer in config/base_config.py, the conversion utility in tools/crawler_util.py, and the platform-specific login handlers.
The Cookie Conversion Pipeline
In tools/crawler_util.py, the convert_str_cookie_to_dict function (lines 59-74) parses semicolon-separated cookie strings into dictionaries suitable for Playwright's browser_context.add_cookies method. The function handles the transformation from HTTP header format (key=value; key2=value2) to the structured format required by the browser automation framework.
Platform-specific login classes implement the login_by_cookies() method to inject these cookies:
- ZhiHuLogin in
media_platform/zhihu/login.py(lines 73-82) - XiaoHongShuLogin in
media_platform/xhs/login.py(lines 71-78) - DouYinLogin in
media_platform/douyin/login.py(lines 66-74)
Each method adds cookies to the browser context with the appropriate domain (e.g., .zhihu.com, .xiaohongshu.com, or .douyin.com).
Configuring Cookie-Based Login via Command Line
You can enable cookie-based login using either the global configuration file or command-line arguments defined in cmd_arg/arg.py.
Method 1: Configuration File
Edit config/base_config.py to set the login type and cookie string:
# config/base_config.py
LOGIN_TYPE = "cookie"
COOKIES = "z_c0=your_token_here; d_c0=another_cookie"
Method 2: CLI Arguments
The CLI parser in cmd_arg/arg.py exposes the --lt (login type) and --cookies flags:
python -m main --platform zhihu --lt cookie --cookies "z_c0=abc123; d_c0=def456"
Platform-Specific Cookie Login Examples
Each platform requires specific cookie names and domains. Here are the implementation details for the three primary platforms.
Zhihu Cookie Login
For Zhihu, the ZhiHuLogin.login_by_cookies() method in media_platform/zhihu/login.py expects cookies like z_c0 and d_c0:
python -m main \
--platform zhihu \
--lt cookie \
--cookies "z_c0=abc123; d_c0=def456"
Internally, the method iterates over the converted cookie dictionary and adds each cookie to the Playwright context with the domain .zhihu.com.
XiaoHongShu (XHS) Cookie Login
The XiaoHongShuLogin class in media_platform/xhs/login.py handles the web_session cookie. The domain is selected dynamically based on config.XHS_INTERNATIONAL:
python -m main \
--platform xhs \
--lt cookie \
--cookies "web_session=YOUR_XHS_SESSION"
The cookies are injected into either .xiaohongshu.com or .rednote.com depending on your configuration.
Douyin Cookie Login
For Douyin, implemented in media_platform/douyin/login.py, you typically need sessionid and LOGIN_STATUS:
python -m main \
--platform dy \
--lt cookie \
--cookies "sessionid=YOUR_SESSION; LOGIN_STATUS=1"
The login_by_cookies() method adds these to the .douyin.com domain. The LOGIN_STATUS=1 cookie satisfies the check_login_state() validation, allowing immediate crawling.
Using the Python API Directly
To bypass the CLI entirely, instantiate platform login classes directly with the cookie_str parameter:
from media_platform.zhihu.login import ZhiHuLogin
from playwright.async_api import async_playwright
async def run_cookie_login():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=False)
context = await browser.new_context()
page = await context.new_page()
login = ZhiHuLogin(
login_type="cookie",
browser_context=context,
context_page=page,
cookie_str="z_c0=abc123; d_c0=def456"
)
await login.begin() # Injects cookies and starts crawling
This approach uses the same convert_str_cookie_to_dict utility internally but gives you full control over the browser context initialization.
Summary
- Configure via CLI: Use
--lt cookie --cookies "key=value"arguments parsed bycmd_arg/arg.py - Configure via file: Set
LOGIN_TYPE = "cookie"andCOOKIESstring inconfig/base_config.py - Conversion utility:
tools/crawler_util.pycontainsconvert_str_cookie_to_dictto transform cookie strings for Playwright - Platform implementations: Each platform (Zhihu, XHS, Douyin) implements
login_by_cookies()in their respectivelogin.pyfiles - Domain specificity: Cookies are scoped to platform-specific domains (
.zhihu.com,.xiaohongshu.com,.douyin.com) - Cookie refresh: Use
convert_browser_context_cookiesintools/crawler_util.py(lines 48-56) to retrieve updated cookies from the browser context after login
Frequently Asked Questions
What format should the cookie string be in?
MediaCrawler expects a semicolon-separated string following HTTP cookie header conventions, such as key1=value1; key2=value2. The convert_str_cookie_to_dict function in tools/crawler_util.py parses this format into dictionaries required by Playwright's browser_context.add_cookies method.
Where does MediaCrawler store cookies after injection?
After parsing, cookies are injected directly into the Playwright BrowserContext using browser_context.add_cookies. For session persistence, the convert_browser_context_cookies function (lines 48-56 in tools/crawler_util.py) can retrieve current cookies from the browser context to update HTTP headers.
Can I use cookie-based login to bypass QR code scanning?
Yes. When LOGIN_TYPE is set to cookie and valid cookies are provided via the --cookies argument or config/base_config.py, MediaCrawler skips the QR code and phone verification flows entirely. The platform-specific login_by_cookies() method proceeds directly to the crawling workflow after injecting cookies.
Which cookies are required for each platform?
Requirements vary by platform implementation:
- Zhihu: Typically requires
z_c0for authentication andd_c0for device tracking - XiaoHongShu: Primarily requires the
web_sessioncookie - Douyin: Requires
sessionidandLOGIN_STATUS=1to satisfycheck_login_state()validation
Check the specific platform's login.py file (e.g., media_platform/zhihu/login.py) to verify required cookie names for your target platform.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →