Enabling Second-Level Comment Crawling for All Platforms in MediaCrawler
MediaCrawler supports hierarchical comment extraction through a single global flag ENABLE_GET_SUB_COMMENTS that activates sub-comment retrieval logic across all platform-specific clients when enabled via CLI or configuration.
The MediaCrawler repository provides a unified framework for scraping content from Chinese social media platforms. While first-level (root) comments are collected by default, second-level (reply) comments require explicit activation through a centralized configuration mechanism that propagates to every supported platform client.
How the Configuration Flag Controls Sub-Comment Crawling
The feature centers on the ENABLE_GET_SUB_COMMENTS boolean defined in config/base_config.py at lines 16-18. By default, this value is set to False, ensuring that sub-comment crawling remains opt-in to conserve bandwidth and storage resources.
When you launch the crawler via command line, the argument parser in cmd_arg/arg.py (lines 13-18) exposes the --get_sub_comment option. This CLI flag accepts truthy values (yes, true, 1, y, t) and writes the parsed boolean directly to config.ENABLE_GET_SUB_COMMENTS.
Each platform client implements a guard clause that checks this flag before executing sub-comment retrieval logic. For example, in media_platform/zhihu/client.py (lines 51-53), the method returns an empty list immediately if the flag is disabled:
if not config.ENABLE_GET_SUB_COMMENTS:
return []
When the flag is True, clients enter a pagination loop that checks the root comment's sub_comment_count, repeatedly calls platform-specific APIs (such as get_child_comments or get_note_sub_comments), extracts comment objects via _extractor.extract_comments, and accumulates results into all_sub_comments for merging with first-level data.
Enabling Second-Level Comments via CLI
The simplest activation method is passing the --get_sub_comment flag when launching the crawler from your terminal:
python media_crawler.py --platform zhihu --get_sub_comment true
The argument parser accepts multiple truthy representations, allowing flexibility in shell scripts and automation pipelines. Valid positive values include yes, true, 1, y, and t.
Enabling Second-Level Comments Programmatically
For custom scripts or Jupyter notebooks that instantiate clients directly, import the configuration object and set the attribute before creating any client instances:
from config.base_config import config
# Enable sub-comment crawling globally
config.ENABLE_GET_SUB_COMMENTS = True
# Subsequent client instances will automatically fetch replies
from media_platform.zhihu.client import ZhihuClient
client = ZhihuClient()
This approach ensures that all subsequent platform client instantiations respect the updated configuration without requiring command-line arguments.
Platform-Specific Implementation Details
Each platform client implements the get_comments_all_sub_comments method with logic tailored to that platform's API structure and pagination mechanisms.
Zhihu
In media_platform/zhihu/client.py (lines 33-84), the client loops while parment_comment.sub_comment_count > 0, invoking get_child_comments for each pagination batch and extracting structured data via _extractor.extract_comments.
XiaoHongShu (XHS)
The XHS client in media_platform/xhs/client.py (lines 474-535) monitors sub_comment_has_more and sub_comment_cursor fields to manage pagination through the get_note_sub_comments endpoint.
Located in media_platform/weibo/client.py (lines 24-27 and 246), the implementation reads the nested comments field within comment items and forwards these replies to the callback function if present.
Tieba
The Tieba client at media_platform/tieba/client.py (lines 502-595 and 532) calculates the total number of sub-comment pages using integer division (sub_comment_count // 10 + 1) and generates sequential URLs to fetch each page.
Kuaishou
In media_platform/kuaishou/client.py (lines 277-305), the logic checks the hasSubComments boolean flag and utilizes cursor-based pagination through the pcursorV2 parameter to retrieve nested replies.
Douyin and Bilibili
These platforms handle sub-comments through the is_fetch_sub_comments parameter within their core classes rather than the global flag alone. See media_platform/douyin/core.py (line 259) and media_platform/bilibili/core.py (line 358) for the specific implementation details that integrate with the broader configuration system.
Data Storage and Schema Compatibility
When sub-comments are collected, they are treated as standard comment objects and persisted according to your selected storage backend—whether MongoDB, SQLite, JSONL, or others. The data models already include a sub_comment_count field (visible in model/m_zhihu.py and analogous platform models) that records the number of replies, ensuring backward compatibility with existing database schemas and downstream analytics pipelines.
Summary
- Global Flag: Set
ENABLE_GET_SUB_COMMENTS = Trueinconfig/base_config.pyor use the--get_sub_commentCLI argument to activate hierarchical comment crawling. - Platform Coverage: The flag controls sub-comment retrieval in Zhihu, XHS, Weibo, Tieba, and Kuaishou clients through guard clauses in their respective
get_comments_all_sub_commentsmethods. - CLI Integration: The
cmd_arg/arg.pymodule parses--get_sub_commentand accepts multiple truthy value formats for shell scripting convenience. - Pagination Logic: Each platform implements custom pagination handling (cursor-based, page-based, or count-based) to traverse the full reply hierarchy.
- Storage Ready: Sub-comments integrate seamlessly with existing storage backends and preserve the
sub_comment_countmetadata field.
Frequently Asked Questions
Does enabling sub-comment crawling significantly increase API rate limit usage?
Yes. Each second-level comment requires additional API calls proportional to the reply count. Platforms like Tieba calculate pages as sub_comment_count // 10 + 1, while cursor-based platforms like Kuaishou iterate until hasSubComments returns false. Monitor your request quotas when enabling this feature on high-engagement posts.
Which platforms support second-level comment extraction in MediaCrawler?
All major platforms in the repository support the feature: Zhihu, XiaoHongShu (XHS), Weibo, Tieba, Kuaishou, Douyin, and Bilibili. Each implements the logic in their respective client files (e.g., media_platform/zhihu/client.py) or core classes (media_platform/douyin/core.py).
How are sub-comments structured in the output data?
Sub-comments are extracted using the same _extractor.extract_comments method as first-level comments and merged into the final result set. They maintain identical schema fields including sub_comment_count, allowing recursive analysis of conversation depth without separate processing pipelines.
Can I disable sub-comment crawling mid-execution?
While the ENABLE_GET_SUB_COMMENTS flag is checked at the start of each get_comments_all_sub_comments invocation, changing the configuration value during runtime only affects subsequent comment retrieval calls. Already fetched sub-comments remain in memory, but new parent comments will skip sub-comment fetching if the flag is set to False before their processing begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →