How to Parse the Data Extracted by MediaCrawler: A Complete Guide
MediaCrawler transforms raw HTML and JSON from Chinese social platforms into clean, typed Python objects through a three-layer pipeline of extraction utilities, platform-specific parsers, and Pydantic-style data models.
To parse the data extracted by MediaCrawler effectively, you need to understand how the repository structures its extraction workflow. The codebase processes content from platforms like Zhihu, Tieba, and Douyin by first sanitizing raw responses with low-level utilities, then mapping fields to platform-specific extractors, and finally validating against strict data models before persistence. This architecture ensures type safety and consistent formatting across MongoDB, Excel, and other backends.
Understanding the Three-Layer Parsing Architecture
MediaCrawler's parsing workflow is built around three distinct layers that transform raw network responses into database-ready records:
- Low-level extraction utilities – Generic functions in
tools/crawler_util.pythat strip HTML tags, decode URL parameters, and parse cookies. - Platform-specific extractors – Classes like
ZhihuHelpandTieBaExtractorthat understand the JSON structure of each platform's API and invoke the low-level utilities. - Store/Model layer – Typed data classes in
model/*.pyand persistence logic instore/*/_store_impl.pythat handle the final validation and writing to configured backends.
Step 1: Clean Raw Content with Low-Level Utilities
The foundation of MediaCrawler's parsing pipeline lives in tools/crawler_util.py. This file contains the core extract_text_from_html function that removes script tags, style elements, and all remaining HTML markup to produce plain text.
def extract_text_from_html(html: str) -> str:
"""Extract text from HTML, removing all tags."""
if not html:
return ""
# Remove script and style elements
clean_html = re.sub(r'<(script|style)[^>]*>.*?</\1>', '',
html, flags=re.DOTALL)
# Remove all other tags
clean_text = re.sub(r'<[^>]+>', '', clean_html).strip()
return clean_text
Source: tools/crawler_util.py#L15-L24
Additional helpers like extract_url_params_to_dict and convert_str_cookie_to_dict (lines 26-34) support the extractors when decoding query strings or authentication cookies.
Step 2: Transform Platform-Specific Responses
Each supported platform implements its own extractor class that knows how to navigate the platform's API response structure and apply the low-level cleaning utilities.
Zhihu Content Extraction
In media_platform/zhihu/help.py, the ZhihuHelp class processes JSON payloads from Zhihu's API. When extracting an answer, it sanitizes every text field using extract_text_from_html:
res.content_text = extract_text_from_html(answer.get("content", ""))
res.title = extract_text_from_html(answer.get("title", ""))
res.desc = extract_text_from_html(answer.get("description", "")
or answer.get("excerpt", ""))
Source: media_platform/zhihu/help.py#L111-L115
Comment extraction follows the same pattern:
res.content = extract_text_from_html(comment.get("content"))
Source: media_platform/zhihu/help.py#L254-L256
Tieba Search Results
The TieBaExtractor class (demonstrated in tests/test_tieba_extractor.py) handles Baidu Tieba's HTML responses. It uses BeautifulSoup internally to parse the DOM, then relies on extract_text_from_html for final cleaning:
notes = TieBaExtractor.extract_search_note_list(read_fixture("search_keyword_notes.html"))
Source: tests/test_tieba_extractor.py#L16-L17
Douyin Media Processing
For Douyin, the extraction logic resides in store/douyin/_store_impl.py. Private helper functions prefixed with _extract_ parse the raw aweme_item JSON to assemble media URLs and image lists:
"video_download_url": _extract_video_download_url(aweme_item),
"pictures": ",".join(_extract_comment_image_list(comment_item)),
Source: store/douyin/_store_impl.py#L181-L184
These helpers extract specific fields from the nested JSON and handle any necessary HTML cleaning before the data reaches the model layer.
Step 3: Validate and Structure with Data Models
After extraction, data flows into strictly typed models defined in the model/ directory. These Pydantic-style classes ensure that all fields are present and correctly typed before persistence:
- Zhihu:
model/m_zhihu.pydefinesZhihuContent,ZhihuComment, andZhihuCreator - Tieba:
model/m_baidu_tieba.pydefinesTiebaNoteandTiebaComment - Douyin:
model/m_douyin.pydefinesDouyinMediaandDouyinComment
For example, the ZhihuContent model contains fields like content_text, title, desc, and author, which the extractor populates directly after cleaning.
Source: [model/m_zhihu.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_zhihu.py)
Step 4: Persist Clean Data via the Store Layer
The final parsing step involves the store implementations in store/*/_store_impl.py. These classes take the validated model objects and write them to the configured backend (MongoDB, Excel, etc.).
The StoreBase class in store/excel_store_base.py defines the interface, while platform-specific stores implement the concrete logic. For example, the Zhihu store uses the model's .dict() method to insert clean data:
await self.db.insert_one(content.dict())
Source: [store/zhihu/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py)
At this stage, no further HTML sanitization is required because the extractors have already produced clean text.
Complete Code Examples
Parse a Zhihu Answer
from media_platform.zhihu.help import ZhihuHelp
from tools.crawler_util import extract_text_from_html
# Assume raw_answer_json is the JSON payload from Zhihu API
zhihu = ZhihuHelp()
content = zhihu._extract_answer_content(raw_answer_json)
print("Title:", content.title) # Clean text, no HTML tags
print("Body:", content.content_text) # Plain text extracted by extract_text_from_html
Key calls: _extract_answer_content → extract_text_from_html at media_platform/zhihu/help.py#L111-L115
Parse a Tieba Note List from Saved HTML
from tests.test_tieba_extractor import TieBaExtractor
from pathlib import Path
html = Path("fixtures/search_keyword_notes.html").read_text()
notes = TieBaExtractor.extract_search_note_list(html)
for note in notes:
print(note.title) # Already stripped of HTML tags
Key call: extract_search_note_list internally uses extract_text_from_html to clean titles.
Store Parsed Douyin Media
from store.douyin._store_impl import DouyinStore, _extract_video_download_url, _extract_comment_image_list
from model.m_douyin import DouyinMedia
# aweme_detail is the raw JSON from Douyin API
media_dict = {
"video_download_url": _extract_video_download_url(aweme_detail),
"pictures": ",".join(_extract_comment_image_list(aweme_detail)),
# ... other fields
}
media = DouyinMedia(**media_dict)
store = DouyinStore()
store.save(media) # Persists clean data to the configured DB
Key helpers: _extract_video_download_url, _extract_comment_image_list defined in store/douyin/_store_impl.py#L181-L184
Summary
- Low-level utilities in
tools/crawler_util.pyprovide theextract_text_from_htmlfunction that strips HTML tags from raw content. - Platform extractors like
ZhihuHelpandTieBaExtractorapply these utilities to platform-specific JSON/HTML responses atmedia_platform/zhihu/help.pyandtests/test_tieba_extractor.py. - Data models in
model/*.pyvalidate the cleaned data into typed Python objects before persistence. - Store implementations in
store/*/_store_impl.pyreceive validated models and write them to MongoDB, Excel, or other backends without additional parsing.
Frequently Asked Questions
How does MediaCrawler handle HTML sanitization for extracted content?
MediaCrawler uses the extract_text_from_html function in tools/crawler_util.py to remove script tags, style elements, and all remaining HTML markup using regular expressions. This utility is called by platform-specific extractors like ZhihuHelp before populating the data models, ensuring that stored content contains only plain text.
Can I customize the data models for different storage requirements?
Yes. The data models in model/*.py (such as m_zhihu.py and m_douyin.py) define the structure for each platform's content. You can modify these Pydantic-style classes to add fields, change types, or set default values. The store implementations in store/*/_store_impl.py will automatically use the updated model structure when persisting data.
Where should I add parsing logic for a new social platform?
To parse data for a new platform, create a new extractor class in media_platform/{platform}/help.py that uses the utilities from tools/crawler_util.py to clean HTML/JSON. Then define corresponding models in model/m_{platform}.py and implement the store logic in store/{platform}/_store_impl.py. Follow the existing patterns in the Zhihu and Tieba implementations for consistency.
What is the difference between the extractor and store layers?
The extractor layer (e.g., media_platform/zhihu/help.py) transforms raw API responses into cleaned Python dictionaries. The store layer (e.g., store/zhihu/_store_impl.py) takes these dictionaries, validates them against data models, and persists them to the configured backend. The extractor handles parsing and cleaning, while the store handles validation and database operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →