How to Parse Scraped Data with MediaCrawler: A Complete Guide
MediaCrawler parses scraped data through pure-function helpers in media_platform/<platform>/help.py that convert raw URLs, HTML, and API responses into strongly-typed data objects.
The NanmiCoder/MediaCrawler repository isolates all parsing logic in lightweight helper modules. Each platform—XiaoHongShu, Kuaishou, Douyin, Bilibili, Zhihu—ships its own help.py file containing data classes and transformation functions. This architecture keeps extraction logic predictable, testable, and reusable across different crawling workflows.
Understanding MediaCrawler's Parsing Architecture
MediaCrawler follows a pure-function design for all data extraction. The parsing layer never performs I/O or maintains state; it receives a raw string and returns a structured object.
The key design benefits include:
- Unit-testing isolation — each parser validates independently in
tests/ - Cross-platform consistency — identical function signatures across all
help.pymodules - Crawler integration — platform-specific crawlers call helpers to normalize data before persistence
Core Parsing Operations by Platform
XiaoHongShu (XHS) URL Parsing
The XiaoHongShu helper in media_platform/xhs/help.py handles note and creator URLs through two primary functions.
Parse a note URL with parse_note_info_from_note_url:
from media_platform.xhs.help import parse_note_info_from_note_url
url = "https://www.xiaohongshu.com/explore/66fad51c000000001b0224b8?xsec_token=AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=&xsec_source=pc_search"
note_info = parse_note_info_from_note_url(url)
print(note_info.note_id) # 66fad51c000000001b0224b8
print(note_info.xsec_token) # AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=
print(note_info.xsec_source) # pc_search
The function returns a NoteUrlInfo dataclass containing the note ID and optional security tokens required for downstream API calls.
Parse a creator URL with parse_creator_info_from_url:
from media_platform.xhs.help import parse_creator_info_from_url
creator = parse_creator_info_from_url(
"https://www.xiaohongshu.com/user/profile/5eb8e1d400000000010075ae?xsec_token=abc&xsec_source=pc_feed"
)
print(creator.user_id) # 5eb8e1d400000000010075ae
print(creator.xsec_token) # abc
print(creator.xsec_source) # pc_feed
This helper accepts either full profile URLs or raw user-ID strings, returning a CreatorUrlInfo instance.
Both functions rely on extract_url_params_to_dict, an internal utility that transforms query strings into Python dictionaries.
Kuaishou Video and Creator Parsing
The Kuaishou helper in media_platform/kuaishou/help.py mirrors the XiaoHongShu pattern with video-specific adaptations.
Parse a video URL with parse_video_info_from_url:
from media_platform.kuaishou.help import parse_video_info_from_url
vid = parse_video_info_from_url(
"https://www.kuaishou.com/short-video/3x3zxz4mjrsc8ke?authorId=3x84qugg4ch9zhs"
)
print(vid.video_id) # 3x3zxz4mjrsc8ke
print(vid.url_type) # normal
The return type is VideoUrlInfo, which captures the video ID and URL classification. The function handles both full URLs and plain video identifiers.
Parse a creator from raw ID:
from media_platform.kuaishou.help import parse_creator_info_from_url
creator = parse_creator_info_from_url("3x84qugg4ch9zhs")
print(creator.user_id) # 3x84qugg4ch9zhs
Zhihu HTML Content Extraction
Zhihu parsing differs because content resides in rendered HTML rather than URL parameters. The helper in media_platform/zhihu/help.py uses parsel.Selector to locate and extract embedded JSON data.
Extract answer content with extract_answer_content_from_html:
from media_platform.zhihu.help import extract_answer_content_from_html
with open("zhihu_answer_page.html", "r", encoding="utf-8") as f:
html = f.read()
answer = extract_answer_content_from_html(html)
print(answer.title) # Answer title from page
print(answer.content) # Cleaned markdown/text content
The function targets the js-initialData script block, deserializes it, and maps the result to a ZhihuContent dataclass. A parallel function, extract_article_content_from_html, handles Zhihu articles.
Douyin and Bilibili Parsers
Douyin and Bilibili implement identical patterns in media_platform/douyin/help.py and media_platform/bilibili/help.py:
parse_video_info_from_url— Extracts video IDs from Douyin share URLs or Bilibili BV/AV codesparse_creator_info_from_url— Parses user profiles from full URLs or raw sec_user_id values
Integrating Parsers into Crawler Workflows
Parsing functions integrate directly with platform crawlers. The KuaiShouCrawler class in media_platform/kuaishou/core.py expects pre-parsed VideoUrlInfo objects:
from media_platform.kuaishou.core import KuaiShouCrawler
from media_platform.kuaishou.help import parse_video_info_from_url
crawler = KuaiShouCrawler()
video_url = "3xf8enb8dbj6uig"
info = parse_video_info_from_url(video_url)
videos = crawler.get_videos([info])
This separation of concerns—parsing before crawling—allows validation and deduplication at the boundary of the system.
Key Source Files for MediaCrawler Parsing
| File | Purpose |
|---|---|
media_platform/xhs/help.py |
XiaoHongShu note/creator URL parsing |
media_platform/kuaishou/help.py |
Kuaishou video/creator extraction |
media_platform/douyin/help.py |
Douyin video/creator parsing |
media_platform/bilibili/help.py |
Bilibili video ID handling |
media_platform/zhihu/help.py |
Zhihu HTML/JSON content extraction |
base/base_crawler.py |
Abstract crawler orchestration |
tests/test_xhs_raw_response_errors.py |
Parser validation suite |
Summary
- MediaCrawler parsing lives in
media_platform/<platform>/help.py— one module per social platform - All parsers are pure functions — input string in, typed object out, no side effects
parse_note_info_from_note_urlandparse_creator_info_from_urlhandle XiaoHongShu URLsparse_video_info_from_urlextracts video metadata from Kuaishou, Douyin, and Bilibili linksextract_answer_content_from_htmlusesparsel.Selectorfor Zhihu HTML processing- Crawlers consume parsed objects directly — decouple extraction from network operations
Frequently Asked Questions
How does MediaCrawler handle malformed URLs?
Each parser in MediaCrawler validates input and returns None for unrecognizable formats. The extract_url_params_to_dict utility in media_platform/xhs/help.py safely handles missing query parameters, while platform-specific regex patterns filter invalid ID structures. Test suites under tests/ explicitly cover edge cases like empty strings and truncated URLs.
Can I use MediaCrawler's parsers without running the full crawler?
Yes. All helper functions are importable and executable independently. Import parse_note_info_from_note_url or extract_answer_content_from_html directly—these have no dependency on base/base_crawler.py or network operations. This design supports data pipeline integrations where raw content arrives through channels other than MediaCrawler's fetch layer.
What data types do MediaCrawler parsers return?
Parsers return frozen dataclass instances: NoteUrlInfo, CreatorUrlInfo, VideoUrlInfo, ZhihuContent, and platform-specific variants. These objects expose typed attributes (e.g., note_id: str, xsec_token: Optional[str]) and support conversion to dictionaries via .asdict() methods where implemented.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →