# How to Parse Scraped Data with MediaCrawler: A Complete Guide

> Learn how to parse scraped data from MediaCrawler using pure-function helpers. Convert raw URLs, HTML, and API responses into strongly-typed data objects with this complete guide.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-12

---

**MediaCrawler parses scraped data through pure-function helpers in `media_platform/<platform>/help.py` that convert raw URLs, HTML, and API responses into strongly-typed data objects.**

The `NanmiCoder/MediaCrawler` repository isolates all parsing logic in lightweight helper modules. Each platform—XiaoHongShu, Kuaishou, Douyin, Bilibili, Zhihu—ships its own [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py) file containing data classes and transformation functions. This architecture keeps extraction logic predictable, testable, and reusable across different crawling workflows.

## Understanding MediaCrawler's Parsing Architecture

MediaCrawler follows a **pure-function design** for all data extraction. The parsing layer never performs I/O or maintains state; it receives a raw string and returns a structured object.

The key design benefits include:

- **Unit-testing isolation** — each parser validates independently in `tests/`
- **Cross-platform consistency** — identical function signatures across all [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py) modules
- **Crawler integration** — platform-specific crawlers call helpers to normalize data before persistence

## Core Parsing Operations by Platform

### XiaoHongShu (XHS) URL Parsing

The XiaoHongShu helper in [`media_platform/xhs/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/help.py) handles note and creator URLs through two primary functions.

**Parse a note URL with `parse_note_info_from_note_url`:**

```python
from media_platform.xhs.help import parse_note_info_from_note_url

url = "https://www.xiaohongshu.com/explore/66fad51c000000001b0224b8?xsec_token=AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=&xsec_source=pc_search"
note_info = parse_note_info_from_note_url(url)

print(note_info.note_id)       # 66fad51c000000001b0224b8

print(note_info.xsec_token)    # AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=

print(note_info.xsec_source)   # pc_search

```

The function returns a `NoteUrlInfo` dataclass containing the note ID and optional security tokens required for downstream API calls.

**Parse a creator URL with `parse_creator_info_from_url`:**

```python
from media_platform.xhs.help import parse_creator_info_from_url

creator = parse_creator_info_from_url(
    "https://www.xiaohongshu.com/user/profile/5eb8e1d400000000010075ae?xsec_token=abc&xsec_source=pc_feed"
)

print(creator.user_id)      # 5eb8e1d400000000010075ae

print(creator.xsec_token)   # abc

print(creator.xsec_source)  # pc_feed

```

This helper accepts either full profile URLs or raw user-ID strings, returning a `CreatorUrlInfo` instance.

Both functions rely on `extract_url_params_to_dict`, an internal utility that transforms query strings into Python dictionaries.

### Kuaishou Video and Creator Parsing

The Kuaishou helper in [`media_platform/kuaishou/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/help.py) mirrors the XiaoHongShu pattern with video-specific adaptations.

**Parse a video URL with `parse_video_info_from_url`:**

```python
from media_platform.kuaishou.help import parse_video_info_from_url

vid = parse_video_info_from_url(
    "https://www.kuaishou.com/short-video/3x3zxz4mjrsc8ke?authorId=3x84qugg4ch9zhs"
)

print(vid.video_id)   # 3x3zxz4mjrsc8ke

print(vid.url_type)   # normal

```

The return type is `VideoUrlInfo`, which captures the video ID and URL classification. The function handles both full URLs and plain video identifiers.

**Parse a creator from raw ID:**

```python
from media_platform.kuaishou.help import parse_creator_info_from_url

creator = parse_creator_info_from_url("3x84qugg4ch9zhs")
print(creator.user_id)  # 3x84qugg4ch9zhs

```

### Zhihu HTML Content Extraction

Zhihu parsing differs because content resides in rendered HTML rather than URL parameters. The helper in [`media_platform/zhihu/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/help.py) uses `parsel.Selector` to locate and extract embedded JSON data.

**Extract answer content with `extract_answer_content_from_html`:**

```python
from media_platform.zhihu.help import extract_answer_content_from_html

with open("zhihu_answer_page.html", "r", encoding="utf-8") as f:
    html = f.read()

answer = extract_answer_content_from_html(html)
print(answer.title)     # Answer title from page

print(answer.content)   # Cleaned markdown/text content

```

The function targets the `js-initialData` script block, deserializes it, and maps the result to a `ZhihuContent` dataclass. A parallel function, `extract_article_content_from_html`, handles Zhihu articles.

### Douyin and Bilibili Parsers

Douyin and Bilibili implement identical patterns in [`media_platform/douyin/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/help.py) and [`media_platform/bilibili/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/help.py):

- **`parse_video_info_from_url`** — Extracts video IDs from Douyin share URLs or Bilibili BV/AV codes
- **`parse_creator_info_from_url`** — Parses user profiles from full URLs or raw sec_user_id values

## Integrating Parsers into Crawler Workflows

Parsing functions integrate directly with platform crawlers. The `KuaiShouCrawler` class in [`media_platform/kuaishou/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/core.py) expects pre-parsed `VideoUrlInfo` objects:

```python
from media_platform.kuaishou.core import KuaiShouCrawler
from media_platform.kuaishou.help import parse_video_info_from_url

crawler = KuaiShouCrawler()
video_url = "3xf8enb8dbj6uig"
info = parse_video_info_from_url(video_url)

videos = crawler.get_videos([info])

```

This separation of concerns—parsing before crawling—allows validation and deduplication at the boundary of the system.

## Key Source Files for MediaCrawler Parsing

| File | Purpose |
|------|---------|
| [`media_platform/xhs/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/help.py) | XiaoHongShu note/creator URL parsing |
| [`media_platform/kuaishou/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/help.py) | Kuaishou video/creator extraction |
| [`media_platform/douyin/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/help.py) | Douyin video/creator parsing |
| [`media_platform/bilibili/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/help.py) | Bilibili video ID handling |
| [`media_platform/zhihu/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/help.py) | Zhihu HTML/JSON content extraction |
| [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) | Abstract crawler orchestration |
| [`tests/test_xhs_raw_response_errors.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_xhs_raw_response_errors.py) | Parser validation suite |

## Summary

- **MediaCrawler parsing lives in `media_platform/<platform>/help.py`** — one module per social platform
- **All parsers are pure functions** — input string in, typed object out, no side effects
- **`parse_note_info_from_note_url`** and **`parse_creator_info_from_url`** handle XiaoHongShu URLs
- **`parse_video_info_from_url`** extracts video metadata from Kuaishou, Douyin, and Bilibili links
- **`extract_answer_content_from_html`** uses `parsel.Selector` for Zhihu HTML processing
- **Crawlers consume parsed objects directly** — decouple extraction from network operations

## Frequently Asked Questions

### How does MediaCrawler handle malformed URLs?

Each parser in MediaCrawler validates input and returns `None` for unrecognizable formats. The `extract_url_params_to_dict` utility in [`media_platform/xhs/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/help.py) safely handles missing query parameters, while platform-specific regex patterns filter invalid ID structures. Test suites under `tests/` explicitly cover edge cases like empty strings and truncated URLs.

### Can I use MediaCrawler's parsers without running the full crawler?

Yes. All helper functions are importable and executable independently. Import `parse_note_info_from_note_url` or `extract_answer_content_from_html` directly—these have no dependency on [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) or network operations. This design supports data pipeline integrations where raw content arrives through channels other than MediaCrawler's fetch layer.

### What data types do MediaCrawler parsers return?

Parsers return frozen dataclass instances: `NoteUrlInfo`, `CreatorUrlInfo`, `VideoUrlInfo`, `ZhihuContent`, and platform-specific variants. These objects expose typed attributes (e.g., `note_id: str`, `xsec_token: Optional[str]`) and support conversion to dictionaries via `.asdict()` methods where implemented.