How to Parse the Data Extracted by MediaCrawler: A Complete Guide

MediaCrawler transforms raw HTML and JSON from Chinese social platforms into clean, typed Python objects through a three-layer pipeline of extraction utilities, platform-specific parsers, and Pydantic-style data models.

To parse the data extracted by MediaCrawler effectively, you need to understand how the repository structures its extraction workflow. The codebase processes content from platforms like Zhihu, Tieba, and Douyin by first sanitizing raw responses with low-level utilities, then mapping fields to platform-specific extractors, and finally validating against strict data models before persistence. This architecture ensures type safety and consistent formatting across MongoDB, Excel, and other backends.

Understanding the Three-Layer Parsing Architecture

MediaCrawler's parsing workflow is built around three distinct layers that transform raw network responses into database-ready records:

  1. Low-level extraction utilities – Generic functions in tools/crawler_util.py that strip HTML tags, decode URL parameters, and parse cookies.
  2. Platform-specific extractors – Classes like ZhihuHelp and TieBaExtractor that understand the JSON structure of each platform's API and invoke the low-level utilities.
  3. Store/Model layer – Typed data classes in model/*.py and persistence logic in store/*/_store_impl.py that handle the final validation and writing to configured backends.

Step 1: Clean Raw Content with Low-Level Utilities

The foundation of MediaCrawler's parsing pipeline lives in tools/crawler_util.py. This file contains the core extract_text_from_html function that removes script tags, style elements, and all remaining HTML markup to produce plain text.

def extract_text_from_html(html: str) -> str:
    """Extract text from HTML, removing all tags."""
    if not html:
        return ""
    # Remove script and style elements

    clean_html = re.sub(r'<(script|style)[^>]*>.*?</\1>', '',
                        html, flags=re.DOTALL)
    # Remove all other tags

    clean_text = re.sub(r'<[^>]+>', '', clean_html).strip()
    return clean_text

Source: tools/crawler_util.py#L15-L24

Additional helpers like extract_url_params_to_dict and convert_str_cookie_to_dict (lines 26-34) support the extractors when decoding query strings or authentication cookies.

Step 2: Transform Platform-Specific Responses

Each supported platform implements its own extractor class that knows how to navigate the platform's API response structure and apply the low-level cleaning utilities.

Zhihu Content Extraction

In media_platform/zhihu/help.py, the ZhihuHelp class processes JSON payloads from Zhihu's API. When extracting an answer, it sanitizes every text field using extract_text_from_html:

res.content_text = extract_text_from_html(answer.get("content", ""))
res.title = extract_text_from_html(answer.get("title", ""))
res.desc = extract_text_from_html(answer.get("description", "")
                                   or answer.get("excerpt", ""))

Source: media_platform/zhihu/help.py#L111-L115

Comment extraction follows the same pattern:

res.content = extract_text_from_html(comment.get("content"))

Source: media_platform/zhihu/help.py#L254-L256

Tieba Search Results

The TieBaExtractor class (demonstrated in tests/test_tieba_extractor.py) handles Baidu Tieba's HTML responses. It uses BeautifulSoup internally to parse the DOM, then relies on extract_text_from_html for final cleaning:

notes = TieBaExtractor.extract_search_note_list(read_fixture("search_keyword_notes.html"))

Source: tests/test_tieba_extractor.py#L16-L17

Douyin Media Processing

For Douyin, the extraction logic resides in store/douyin/_store_impl.py. Private helper functions prefixed with _extract_ parse the raw aweme_item JSON to assemble media URLs and image lists:

"video_download_url": _extract_video_download_url(aweme_item),
"pictures": ",".join(_extract_comment_image_list(comment_item)),

Source: store/douyin/_store_impl.py#L181-L184

These helpers extract specific fields from the nested JSON and handle any necessary HTML cleaning before the data reaches the model layer.

Step 3: Validate and Structure with Data Models

After extraction, data flows into strictly typed models defined in the model/ directory. These Pydantic-style classes ensure that all fields are present and correctly typed before persistence:

For example, the ZhihuContent model contains fields like content_text, title, desc, and author, which the extractor populates directly after cleaning.

Source: [model/m_zhihu.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_zhihu.py)

Step 4: Persist Clean Data via the Store Layer

The final parsing step involves the store implementations in store/*/_store_impl.py. These classes take the validated model objects and write them to the configured backend (MongoDB, Excel, etc.).

The StoreBase class in store/excel_store_base.py defines the interface, while platform-specific stores implement the concrete logic. For example, the Zhihu store uses the model's .dict() method to insert clean data:

await self.db.insert_one(content.dict())

Source: [store/zhihu/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py)

At this stage, no further HTML sanitization is required because the extractors have already produced clean text.

Complete Code Examples

Parse a Zhihu Answer

from media_platform.zhihu.help import ZhihuHelp
from tools.crawler_util import extract_text_from_html

# Assume raw_answer_json is the JSON payload from Zhihu API

zhihu = ZhihuHelp()
content = zhihu._extract_answer_content(raw_answer_json)

print("Title:", content.title)          # Clean text, no HTML tags

print("Body:", content.content_text)    # Plain text extracted by extract_text_from_html

Key calls: _extract_answer_content → extract_text_from_html at media_platform/zhihu/help.py#L111-L115

Parse a Tieba Note List from Saved HTML

from tests.test_tieba_extractor import TieBaExtractor
from pathlib import Path

html = Path("fixtures/search_keyword_notes.html").read_text()
notes = TieBaExtractor.extract_search_note_list(html)

for note in notes:
    print(note.title)   # Already stripped of HTML tags

Key call: extract_search_note_list internally uses extract_text_from_html to clean titles.

Store Parsed Douyin Media

from store.douyin._store_impl import DouyinStore, _extract_video_download_url, _extract_comment_image_list
from model.m_douyin import DouyinMedia

# aweme_detail is the raw JSON from Douyin API

media_dict = {
    "video_download_url": _extract_video_download_url(aweme_detail),
    "pictures": ",".join(_extract_comment_image_list(aweme_detail)),
    # ... other fields

}
media = DouyinMedia(**media_dict)

store = DouyinStore()
store.save(media)   # Persists clean data to the configured DB

Key helpers: _extract_video_download_url, _extract_comment_image_list defined in store/douyin/_store_impl.py#L181-L184

Summary

  • Low-level utilities in tools/crawler_util.py provide the extract_text_from_html function that strips HTML tags from raw content.
  • Platform extractors like ZhihuHelp and TieBaExtractor apply these utilities to platform-specific JSON/HTML responses at media_platform/zhihu/help.py and tests/test_tieba_extractor.py.
  • Data models in model/*.py validate the cleaned data into typed Python objects before persistence.
  • Store implementations in store/*/_store_impl.py receive validated models and write them to MongoDB, Excel, or other backends without additional parsing.

Frequently Asked Questions

How does MediaCrawler handle HTML sanitization for extracted content?

MediaCrawler uses the extract_text_from_html function in tools/crawler_util.py to remove script tags, style elements, and all remaining HTML markup using regular expressions. This utility is called by platform-specific extractors like ZhihuHelp before populating the data models, ensuring that stored content contains only plain text.

Can I customize the data models for different storage requirements?

Yes. The data models in model/*.py (such as m_zhihu.py and m_douyin.py) define the structure for each platform's content. You can modify these Pydantic-style classes to add fields, change types, or set default values. The store implementations in store/*/_store_impl.py will automatically use the updated model structure when persisting data.

Where should I add parsing logic for a new social platform?

To parse data for a new platform, create a new extractor class in media_platform/{platform}/help.py that uses the utilities from tools/crawler_util.py to clean HTML/JSON. Then define corresponding models in model/m_{platform}.py and implement the store logic in store/{platform}/_store_impl.py. Follow the existing patterns in the Zhihu and Tieba implementations for consistency.

What is the difference between the extractor and store layers?

The extractor layer (e.g., media_platform/zhihu/help.py) transforms raw API responses into cleaned Python dictionaries. The store layer (e.g., store/zhihu/_store_impl.py) takes these dictionaries, validates them against data models, and persists them to the configured backend. The extractor handles parsing and cleaning, while the store handles validation and database operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →