How MediaCrawler Extracts Cookies from Browser Context Using Playwright
MediaCrawler extracts cookies from browser context by calling Playwright's browser_context.cookies() method and converting the results into both a raw HTTP header string and a Python dictionary via the convert_browser_context_cookies utility in tools/crawler_util.py.
MediaCrawler, an open-source multi-platform content crawler maintained by NanmiCoder, relies on Playwright to automate browser sessions and manage authentication state. Understanding how MediaCrawler extracts cookies from browser context reveals the mechanism that bridges automated browser automation with subsequent HTTP API requests across platforms like Zhihu, Xiaohongshu, and Weibo.
How MediaCrawler Extracts Cookies from Browser Context
The Playwright Foundation
At the core of the extraction process lies Playwright's BrowserContext object. According to the source code in tools/crawler_util.py, MediaCrawler retrieves the raw cookie list by awaiting browser_context.cookies(urls=urls), which returns a list of cookie objects containing name, value, domain, path, and other standard attributes.
If the optional urls parameter is provided, Playwright filters the returned cookies to only those applicable to the specified URLs. This domain-scoped extraction is crucial for platforms that set cookies across multiple domains.
The Dual-Format Conversion Pipeline
Once retrieved, the raw cookie list undergoes transformation through two distinct stages:
convert_browser_context_cookies– An async wrapper that fetches cookies from the browser context and delegates to the conversion helper.convert_cookies– A synchronous utility that simultaneously builds both a semicolon-delimited string ("name=value; name2=value2") suitable for HTTP headers and a dictionary mapping for O(1) lookups.
Core Implementation in tools/crawler_util.py
The primary extraction interface resides in the convert_browser_context_cookies function:
# tools/crawler_util.py – async wrapper for cookie extraction
# https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py
async def convert_browser_context_cookies(
browser_context: BrowserContext, urls: Optional[List[str]] = None
) -> Tuple[str, Dict]:
cookies = (
await browser_context.cookies(urls=urls)
if urls
else await browser_context.cookies()
)
return convert_cookies(cookies)
This function returns a tuple (cookies_str, cookie_dict), where cookies_str is ready for direct injection into HTTP Cookie headers and cookie_dict facilitates fast token lookups for request signing.
Production Usage in Platform Crawlers
Zhihu Implementation
In media_platform/zhihu/core.py, the crawler invokes the utility after login or page navigation to populate the API client:
# media_platform/zhihu/core.py
cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
self.browser_context,
urls=self.cookie_urls,
)
The extracted cookie_str is subsequently passed to the HTTP client in media_platform/zhihu/client.py, while cookie_dict is stored for request signature generation.
Cross-Platform Pattern
Xiaohongshu, Weibo, and other platform implementations follow identical patterns, importing the same utility from tools/crawler_util.py to maintain consistent cookie handling. Additionally, tools/cdp_browser.py provides low-level add_cookies and get_cookies helpers for components requiring direct Chrome DevTools Protocol access.
Handling Edge Cases and URL Filtering
Empty Cookie Scenarios
When no cookies exist in the browser context, the utility gracefully returns ("", {}). This allows downstream HTTP clients to safely skip header injection without raising exceptions or requiring null checks.
Domain-Specific Extraction
The optional urls parameter enables fine-grained control over cookie extraction. By passing specific URLs, crawlers can isolate cookies relevant to particular API endpoints, preventing domain leakage and ensuring only necessary authentication tokens are transmitted.
Practical Example
from tools import crawler_util as utils
from playwright.async_api import async_playwright
async def fetch_cookies():
async with async_playwright() as p:
browser = await p.chromium.launch()
context = await browser.new_context()
await context.add_cookies([
{"name": "sessionid", "value": "abc123", "domain": "example.com", "path": "/"}
])
# Extract only cookies that belong to https://example.com
cookie_str, cookie_dict = await utils.convert_browser_context_cookies(
context, urls=["https://example.com"]
)
print(cookie_str) # → "sessionid=abc123"
print(cookie_dict) # → {"sessionid": "abc123"}
Summary
- MediaCrawler extracts cookies using Playwright's
browser_context.cookies()method through theconvert_browser_context_cookiesasync function. - The utility in
tools/crawler_util.pyconverts raw Playwright cookies into both a HTTP header string and a Python dictionary. - URL filtering is supported via the optional
urlsparameter to isolate domain-specific cookies. - Empty cookie scenarios return
("", {})to prevent downstream errors. - Platform implementations in
media_platform/zhihu/core.pyand similar files demonstrate production usage patterns.
Frequently Asked Questions
What Playwright method does MediaCrawler use to extract cookies?
MediaCrawler calls await browser_context.cookies(urls=urls) through the convert_browser_context_cookies function in tools/crawler_util.py. This retrieves the complete cookie list from the Playwright browser context, optionally filtered by the URLs provided.
Why does MediaCrawler return both a string and dictionary of cookies?
The dual format serves distinct purposes: the semicolon-delimited string ("name=value; name2=value2") is immediately usable for HTTP Cookie headers, while the dictionary enables O(1) lookups for request signing, token validation, and CSRF token extraction in API clients.
How does MediaCrawler handle domain-specific cookies?
By passing the urls parameter to convert_browser_context_cookies, the crawler filters cookies to only those matching the specified domains. This ensures that API requests contain only relevant authentication tokens, preventing cross-domain cookie leakage.
Where is the cookie extraction logic implemented in the MediaCrawler repository?
The core extraction and conversion logic resides in tools/crawler_util.py, with platform-specific usage examples found in media_platform/zhihu/core.py and analogous files for other platforms. Low-level CDP-based cookie operations are handled in tools/cdp_browser.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →