How to Parse Scraped Data with MediaCrawler: A Complete Guide

MediaCrawler parses scraped data through pure-function helpers in media_platform/<platform>/help.py that convert raw URLs, HTML, and API responses into strongly-typed data objects.

The NanmiCoder/MediaCrawler repository isolates all parsing logic in lightweight helper modules. Each platform—XiaoHongShu, Kuaishou, Douyin, Bilibili, Zhihu—ships its own help.py file containing data classes and transformation functions. This architecture keeps extraction logic predictable, testable, and reusable across different crawling workflows.

Understanding MediaCrawler's Parsing Architecture

MediaCrawler follows a pure-function design for all data extraction. The parsing layer never performs I/O or maintains state; it receives a raw string and returns a structured object.

The key design benefits include:

  • Unit-testing isolation — each parser validates independently in tests/
  • Cross-platform consistency — identical function signatures across all help.py modules
  • Crawler integration — platform-specific crawlers call helpers to normalize data before persistence

Core Parsing Operations by Platform

XiaoHongShu (XHS) URL Parsing

The XiaoHongShu helper in media_platform/xhs/help.py handles note and creator URLs through two primary functions.

Parse a note URL with parse_note_info_from_note_url:

from media_platform.xhs.help import parse_note_info_from_note_url

url = "https://www.xiaohongshu.com/explore/66fad51c000000001b0224b8?xsec_token=AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=&xsec_source=pc_search"
note_info = parse_note_info_from_note_url(url)

print(note_info.note_id)       # 66fad51c000000001b0224b8

print(note_info.xsec_token)    # AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=

print(note_info.xsec_source)   # pc_search

The function returns a NoteUrlInfo dataclass containing the note ID and optional security tokens required for downstream API calls.

Parse a creator URL with parse_creator_info_from_url:

from media_platform.xhs.help import parse_creator_info_from_url

creator = parse_creator_info_from_url(
    "https://www.xiaohongshu.com/user/profile/5eb8e1d400000000010075ae?xsec_token=abc&xsec_source=pc_feed"
)

print(creator.user_id)      # 5eb8e1d400000000010075ae

print(creator.xsec_token)   # abc

print(creator.xsec_source)  # pc_feed

This helper accepts either full profile URLs or raw user-ID strings, returning a CreatorUrlInfo instance.

Both functions rely on extract_url_params_to_dict, an internal utility that transforms query strings into Python dictionaries.

Kuaishou Video and Creator Parsing

The Kuaishou helper in media_platform/kuaishou/help.py mirrors the XiaoHongShu pattern with video-specific adaptations.

Parse a video URL with parse_video_info_from_url:

from media_platform.kuaishou.help import parse_video_info_from_url

vid = parse_video_info_from_url(
    "https://www.kuaishou.com/short-video/3x3zxz4mjrsc8ke?authorId=3x84qugg4ch9zhs"
)

print(vid.video_id)   # 3x3zxz4mjrsc8ke

print(vid.url_type)   # normal

The return type is VideoUrlInfo, which captures the video ID and URL classification. The function handles both full URLs and plain video identifiers.

Parse a creator from raw ID:

from media_platform.kuaishou.help import parse_creator_info_from_url

creator = parse_creator_info_from_url("3x84qugg4ch9zhs")
print(creator.user_id)  # 3x84qugg4ch9zhs

Zhihu HTML Content Extraction

Zhihu parsing differs because content resides in rendered HTML rather than URL parameters. The helper in media_platform/zhihu/help.py uses parsel.Selector to locate and extract embedded JSON data.

Extract answer content with extract_answer_content_from_html:

from media_platform.zhihu.help import extract_answer_content_from_html

with open("zhihu_answer_page.html", "r", encoding="utf-8") as f:
    html = f.read()

answer = extract_answer_content_from_html(html)
print(answer.title)     # Answer title from page

print(answer.content)   # Cleaned markdown/text content

The function targets the js-initialData script block, deserializes it, and maps the result to a ZhihuContent dataclass. A parallel function, extract_article_content_from_html, handles Zhihu articles.

Douyin and Bilibili Parsers

Douyin and Bilibili implement identical patterns in media_platform/douyin/help.py and media_platform/bilibili/help.py:

  • parse_video_info_from_url — Extracts video IDs from Douyin share URLs or Bilibili BV/AV codes
  • parse_creator_info_from_url — Parses user profiles from full URLs or raw sec_user_id values

Integrating Parsers into Crawler Workflows

Parsing functions integrate directly with platform crawlers. The KuaiShouCrawler class in media_platform/kuaishou/core.py expects pre-parsed VideoUrlInfo objects:

from media_platform.kuaishou.core import KuaiShouCrawler
from media_platform.kuaishou.help import parse_video_info_from_url

crawler = KuaiShouCrawler()
video_url = "3xf8enb8dbj6uig"
info = parse_video_info_from_url(video_url)

videos = crawler.get_videos([info])

This separation of concerns—parsing before crawling—allows validation and deduplication at the boundary of the system.

Key Source Files for MediaCrawler Parsing

File Purpose
media_platform/xhs/help.py XiaoHongShu note/creator URL parsing
media_platform/kuaishou/help.py Kuaishou video/creator extraction
media_platform/douyin/help.py Douyin video/creator parsing
media_platform/bilibili/help.py Bilibili video ID handling
media_platform/zhihu/help.py Zhihu HTML/JSON content extraction
base/base_crawler.py Abstract crawler orchestration
tests/test_xhs_raw_response_errors.py Parser validation suite

Summary

  • MediaCrawler parsing lives in media_platform/<platform>/help.py — one module per social platform
  • All parsers are pure functions — input string in, typed object out, no side effects
  • parse_note_info_from_note_url and parse_creator_info_from_url handle XiaoHongShu URLs
  • parse_video_info_from_url extracts video metadata from Kuaishou, Douyin, and Bilibili links
  • extract_answer_content_from_html uses parsel.Selector for Zhihu HTML processing
  • Crawlers consume parsed objects directly — decouple extraction from network operations

Frequently Asked Questions

How does MediaCrawler handle malformed URLs?

Each parser in MediaCrawler validates input and returns None for unrecognizable formats. The extract_url_params_to_dict utility in media_platform/xhs/help.py safely handles missing query parameters, while platform-specific regex patterns filter invalid ID structures. Test suites under tests/ explicitly cover edge cases like empty strings and truncated URLs.

Can I use MediaCrawler's parsers without running the full crawler?

Yes. All helper functions are importable and executable independently. Import parse_note_info_from_note_url or extract_answer_content_from_html directly—these have no dependency on base/base_crawler.py or network operations. This design supports data pipeline integrations where raw content arrives through channels other than MediaCrawler's fetch layer.

What data types do MediaCrawler parsers return?

Parsers return frozen dataclass instances: NoteUrlInfo, CreatorUrlInfo, VideoUrlInfo, ZhihuContent, and platform-specific variants. These objects expose typed attributes (e.g., note_id: str, xsec_token: Optional[str]) and support conversion to dictionaries via .asdict() methods where implemented.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →