When to Use login_by_cookies Instead of QR Code Login in MediaCrawler
Use login_by_cookies when running headless automation on servers without displays, when reusing existing authenticated sessions to avoid repeated QR scans, or in CI/CD pipelines where manual interaction is impossible.
MediaCrawler is an open-source social media scraping framework that supports multiple authentication strategies across platforms like Zhihu, XiaoHongShu, and Douyin. While the repository provides QR code login for interactive sessions, the login_by_cookies method offers a critical non-interactive alternative for production deployments. Understanding when to deploy each method ensures reliable automation and prevents authentication failures in environments lacking graphical interfaces.
Headless Servers and CI/CD Environments
QR code login requires a graphical display and manual user interaction to scan the generated image with a mobile device. In media_platform/zhihu/login.py, the login_by_qrcode method extracts a base64-encoded image from a canvas element and displays it locally, blocking execution until the scan completes.
Choose login_by_cookies for:
- Docker containers running without a display server
- CI/CD pipelines where manual intervention is prohibited
- Cloud servers accessed only via SSH
- Scheduled cron jobs that must run unattended
The login_by_cookies method bypasses visual elements entirely by injecting pre-authenticated session data directly into the Playwright browser context.
Reusing Sessions and Performance Optimization
Cookie-based authentication eliminates the 30-60 second overhead of QR code generation, display, and manual scanning. For daily automated crawls or frequent testing cycles, storing and reusing a valid session cookie significantly improves startup time and ensures deterministic execution.
The implementation in media_platform/xhs/login.py, media_platform/weibo/login.py, and other platform login modules follows the same pattern: the begin method checks config.LOGIN_TYPE and routes to either login_by_cookies or login_by_qrcode based on the string value "cookie" or "qrcode".
Implementation Comparison
How login_by_cookies Works
In media_platform/zhihu/login.py, the cookie method parses a semicolon-delimited string and injects each key-value pair into the browser context:
async def login_by_cookies(self):
utils.logger.info("[ZhiHu.login_by_cookies] Begin login zhihu by cookie ...")
for key, value in utils.convert_str_cookie_to_dict(self.cookie_str).items():
await self.browser_context.add_cookies([{
'name': key,
'value': value,
'domain': ".zhihu.com",
'path': "/"
}])
This function relies on utils.convert_str_cookie_to_dict from tools/utils.py to transform the input string into a dictionary before setting domain-specific cookies.
How QR Code Login Works
The QR alternative extracts visual authentication data and polls for completion:
async def login_by_qrcode(self):
utils.logger.info("[ZhiHu.login_by_qrcode] Begin login zhihu by qrcode ...")
base64_qrcode_img = await utils.find_qrcode_img_from_canvas(
self.context_page,
canvas_selector="canvas.Qrcode-qrcode"
)
# show QR code to the user…
partial_show_qrcode = functools.partial(utils.show_qrcode, base64_qrcode_img)
asyncio.get_running_loop().run_in_executor(None, partial_show_qrcode)
await self.check_login_state() # retries until logged‑in or timeout
This method requires the utils.find_qrcode_img_from_canvas helper and maintains an active polling loop until the user completes the scan or the timeout expires.
Configuration Guide
To activate cookie-based authentication:
-
Export cookies from an authenticated browser session using extensions like "EditThisCookie" or browser dev tools, formatting them as
name=value; name2=value2. -
Set the login type via command line (
--lt cookie) or by modifyingconfig/zhihu_config.py(or the respective platform config file) to setLOGIN_TYPE = "cookie". -
Pass the cookie string to the Login class constructor:
from media_platform.zhihu.login import ZhiHuLogin
from playwright.async_api import async_playwright
async def run():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context()
page = await context.new_page()
zhihu = ZhiHuLogin(
login_type="cookie",
browser_context=context,
context_page=page,
cookie_str="z_c0=ABC123; d_c0=XYZ789"
)
await zhihu.begin() # will call login_by_cookies internally
# ... continue with crawling logic ...
# asyncio.run(run())
For comparison, the QR code flow requires changing only the login_type parameter and omitting the cookie string:
zhihu = ZhiHuLogin(
login_type="qrcode",
browser_context=context,
context_page=page,
cookie_str="" # not used
)
await zhihu.begin() # triggers QR‑code flow
Summary
- Use
login_by_cookiesfor headless environments, automated schedules, and CI/CD pipelines where manual QR scanning is impossible. - Use QR code login for one-off interactive sessions when you cannot obtain valid session cookies.
- Configure via
LOGIN_TYPEin platform-specific config files (e.g.,config/zhihu_config.py) or via the--ltcommand line argument. - Cookie format requires semicolon-delimited key-value pairs that
utils.convert_str_cookie_to_dictparses before injection via Playwright'sadd_cookies.
Frequently Asked Questions
Can I switch between login methods without modifying the codebase?
Yes. Change the LOGIN_TYPE variable in the platform's config file (such as config/xhs_config.py or config/weibo_config.py) to "cookie" or "qrcode", or pass --lt cookie as a command line argument when launching the crawler. The begin method in each platform's login.py handles the routing automatically.
How do I export cookies from my browser for MediaCrawler?
Install a browser extension like "EditThisCookie" or "Cookie-Editor," navigate to the target platform while logged in, and export the cookies as a semicolon-separated string. Ensure the string includes essential session identifiers like z_c0 for Zhihu or equivalent tokens for other platforms.
Does login_by_cookies work for all platforms in MediaCrawler?
Yes. Every supported platform—including XiaoHongShu, Weibo, Tieba, Kuaishou, Douyin, and Bilibili—implements login_by_cookies in its respective media_platform/{platform}/login.py file following the same pattern as the Zhihu implementation.
What happens if my cookie expires during a crawl?
The crawler will lose authentication and subsequent requests will fail or return unauthenticated content. You must refresh the cookie by logging in manually again through a browser, export the new session data, and update the cookie_str parameter before restarting the automation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →