How to Handle Login State Persistence and Cookie-Based Authentication in MediaCrawler
MediaCrawler handles login state persistence through a combination of persistent browser user data directories, bidirectional cookie conversion utilities, and platform-specific login verification classes.
To avoid repetitive logins across crawler executions, the repository implements a layered authentication system. The core mechanism relies on Playwright's launch_persistent_context() and custom conversion utilities in tools/crawler_util.py. Below is a complete breakdown of how the system works and how to configure it for your own scraping workflows.
Enabling Persistent Login State in Configuration
Login state caching is controlled by a single boolean flag in the global configuration. When enabled, MediaCrawler creates a dedicated folder on disk that Chrome or Edge treats as its profile directory—preserving cookies, localStorage, and extension states between runs.
- Configuration flag:
SAVE_LOGIN_STATE(default:True) - Storage path:
browser_data/<platform>_user_data_dir
# config/base_config.py (lines 52-55)
SAVE_LOGIN_STATE = True # Enable persistent browser profile caching
USER_DATA_DIR = "%s_user_data_dir" # Platform-specific subdirectory naming
The USER_DATA_DIR template string is formatted with the platform name (e.g., zhihu, xiaohongshu) to isolate profiles per target site.
Launching Browsers with Persistent Context
MediaCrawler supports two browser control modes, both respecting the SAVE_LOGIN_STATE flag.
Playwright Mode
In standard Playwright operation, launch_persistent_context() is invoked when state persistence is enabled. This returns a BrowserContext that automatically reads from and writes to the specified directory.
# media_platform/zhihu/core.py (simplified from lines 35-45)
import os
from playwright.async_api import async_playwright
from config import base_config as config
user_data_dir = os.path.join(
os.getcwd(),
"browser_data",
config.USER_DATA_DIR % config.PLATFORM, # e.g., "zhihu_user_data_dir"
)
if config.SAVE_LOGIN_STATE:
context = await chromium.launch_persistent_context(
user_data_dir=user_data_dir,
headless=config.HEADLESS,
proxy=playwright_proxy,
viewport={"width": 1920, "height": 1080},
user_agent=self.user_agent,
)
else:
browser = await chromium.launch(headless=config.HEADLESS, proxy=playwright_proxy)
context = await browser.new_context(viewport={"width": 1920, "height": 1080})
CDP Mode (Chrome DevTools Protocol)
For attaching to existing browser instances, CDPBrowserManager passes the same user_data_dir parameter when launching Chrome with remote debugging enabled.
# tools/cdp_browser.py (lines 54-63)
def _launch_browser(self):
args = [
f"--remote-debugging-port={self.debug_port}",
f"--user-data-dir={self.user_data_dir}" if config.SAVE_LOGIN_STATE else "",
# additional chromium flags...
]
subprocess.Popen([self.chrome_path] + [a for a in args if a])
This ensures that even CDP-controlled sessions retain authentication state across process restarts.
Extracting and Converting Cookies
Once a user authenticates—whether via QR code, SMS, or manual flow—the crawler extracts cookies from the browser context for reuse in HTTP API calls.
The Cookie Conversion Pipeline
The convert_browser_context_cookies() function in tools/crawler_util.py bridges Playwright's native cookie format and standard HTTP header strings:
# tools/crawler_util.py (lines 48-57)
async def convert_browser_context_cookies(browser_context, urls=None):
cookies = await browser_context.cookies(urls) # Returns list of dicts
cookie_str = "; ".join(f"{c['name']}={c['value']}" for c in cookies)
cookie_dict = {c["name"]: c["value"] for c in cookies}
return cookie_str, cookie_dict
Return values:
cookie_str: Semicolon-separated string ready for theCookie:HTTP headercookie_dict: Key-value mapping for programmatic access
Injecting Cookies into API Requests
After extraction, cookies are attached to the platform client's default headers:
# Typical usage in platform core classes
cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
self.browser_context,
urls=self.cookie_urls, # e.g., ["https://www.zhihu.com"]
)
self.default_headers["cookie"] = cookie_str
self.cookie_dict = cookie_dict
All subsequent API calls through the client automatically carry this session identifier.
Verifying Existing Login State
Each platform implements a check_login_state() method to test whether cached credentials remain valid without triggering a full login flow.
Platform-Specific Verification
For Zhihu, the login class checks for the presence of the z_c0 authentication cookie:
# media_platform/zhihu/login.py (lines 58-63)
async def check_login_state(self):
cookie_str, cookie_dict = await utils.convert_cookies(self.browser_context)
if cookie_dict.get("z_c0"):
logger.info("[ZhiHuLogin] Valid login state detected via z_c0 cookie")
return True
return False
The convert_cookies helper (aliased in crawler_util.py) handles the underlying context extraction.
Client-Level "Pong" Check
The platform client exposes a pong() method that wraps this verification, enabling automatic re-authentication:
# media_platform/zhihu/client.py (lines 64-78)
async def pong(self):
"""Verify session validity; return False if re-login required."""
return await self.login_obj.check_login_state()
async def update_cookies(self, browser_context):
"""Refresh stored cookies after re-authentication."""
cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
browser_context,
urls=self.cookie_urls,
)
self.headers["cookie"] = cookie_str
self.cookie_dict = cookie_dict
Injecting External Cookies for Cookie-Based Login
Users can supply raw cookie strings via configuration to bypass interactive authentication entirely.
Parsing and Loading External Cookies
The COOKIES configuration value is parsed and injected into the browser context before any navigation:
# media_platform/zhihu/login.py (lines 15-25)
async def login_by_cookies(self):
"""Inject user-supplied cookies into browser context."""
cookie_dict = utils.convert_str_cookie_to_dict(self.cookie_str)
# Format: "key1=value1; key2=value2" → {"key1": "value1", "key2": "value2"}
for key, value in cookie_dict.items():
await self.browser_context.add_cookies([{
"name": key,
"value": value,
"domain": ".zhihu.com", # Platform-specific domain
"path": "/",
}])
The convert_str_cookie_to_dict() utility handles the string-to-mapping conversion, normalizing whitespace and duplicate delimiters.
Automatic Re-Login and Cookie Refresh
When a session expires during crawling, MediaCrawler triggers the full authentication workflow and synchronizes the persistent storage:
# Typical orchestration pattern in crawler execution
if not await zhihu_client.pong(): # Session check fails
await zhihu_login.begin() # QR-code or cookie re-login
await zhihu_client.update_cookies( # Sync new cookies to client
self.browser_context
)
# Persistent context automatically saves to user_data_dir
This loop ensures that the browser profile on disk always reflects the freshest valid session.
Summary
- Enable persistence with
SAVE_LOGIN_STATE = Trueinconfig/base_config.py - Store profiles per platform using the
USER_DATA_DIRtemplate - Launch browsers via
chromium.launch_persistent_context()or CDP with--user-data-dir - Convert cookies with
convert_browser_context_cookies()for HTTP reuse - Verify sessions through platform-specific
check_login_state()methods - Inject external cookies using
convert_str_cookie_to_dict()andbrowser_context.add_cookies() - Refresh automatically via
update_cookies()whenpong()detects expiry
Frequently Asked Questions
Where does MediaCrawler store the persistent browser profile?
MediaCrawler stores profiles in browser_data/<platform>_user_data_dir relative to the working directory, as configured by USER_DATA_DIR in config/base_config.py. This folder contains the full Chrome/Edge user data including cookies, localStorage, and cache.
Can I use cookie-based authentication without browser automation?
Yes. Set the COOKIES configuration value to a semicolon-separated string, then call login_by_cookies() which parses and injects them via browser_context.add_cookies(). The crawler still launches a browser instance, but skips interactive authentication flows.
How does MediaCrawler detect if saved cookies have expired?
Each platform's check_login_state() method checks for signature cookies (e.g., z_c0 for Zhihu). The client calls pong() to verify validity; if absent or invalid, the crawler triggers begin() to re-authenticate and refreshes stored cookies with update_cookies().
What is the difference between Playwright and CDP modes for login persistence?
Both modes respect SAVE_LOGIN_STATE. Playwright uses launch_persistent_context() internally, while CDP mode passes --user-data-dir as a Chrome command-line argument. CDP attaches to an existing or newly launched process; Playwright manages the full lifecycle.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →