How MediaCrawler Implements Second-Level Comment Crawling with ENABLE_GET_SUB_COMMENTS
Set ENABLE_GET_SUB_COMMENTS = True in config/base_config.py to enable automatic fetching of nested replies, which triggers the get_comments_all_sub_comments method to paginate through child comments across supported platforms like Zhihu, Douyin, and Bilibili.
MediaCrawler is an open-source multi-platform content crawling framework that supports optional second-level comment extraction via a centralized configuration flag. When enabled, the framework automatically detects parent comments containing nested replies and initiates a dedicated pagination workflow to retrieve all sub-comments while respecting platform rate limits. This implementation maintains API consistency across platforms including Zhihu, Douyin, Bilibili, Weibo, Tieba, and Kuaishou.
Configuration Flag in base_config.py
The feature is controlled by the global boolean ENABLE_GET_SUB_COMMENTS defined in the central configuration file. By default, sub-comment crawling is disabled to minimize API calls and processing time.
# config/base_config.py (lines 118-119)
ENABLE_GET_SUB_COMMENTS = False # ← toggle to enable sub‑comment crawling
You can activate this feature programmatically by setting config.ENABLE_GET_SUB_COMMENTS = True before initializing the client, or via the command line using --enable-sub-comments true, which is parsed in cmd_arg/arg.py and mapped to the same configuration variable.
Entry Point for Root and Sub-Comment Retrieval
For platforms like Zhihu, the primary entry point is get_note_all_comments in the platform-specific client. After retrieving batches of root comments, this method conditionally invokes the sub-comment crawler:
# media_platform/zhihu/client.py (lines 28-30)
await self.get_comments_all_sub_comments(content, comments,
crawl_interval=crawl_interval,
callback=callback)
This pattern is replicated across all supported platforms in their respective media_platform/{platform}/client.py files, ensuring uniform behavior whether you are crawling Zhihu, Douyin, or Bilibili.
The Sub-Comment Crawling Workflow
The method get_comments_all_sub_comments in media_platform/zhihu/client.py implements a seven-step workflow to retrieve nested replies. Each step includes specific guard clauses and pagination logic to handle platform API limitations.
Step 1: Configuration Guard Check
The method immediately returns an empty list if the feature is disabled, preventing unnecessary API calls:
# media_platform/zhihu/client.py (lines 51-53)
if not config.ENABLE_GET_SUB_COMMENTS:
return []
Step 2: Iterate Over Parent Comments
The crawler loops through each root comment, filtering for those with a non-zero sub_comment_count field:
# media_platform/zhihu/client.py (lines 55-57)
for parment_comment in comments:
if parment_comment.sub_comment_count == 0:
continue
Step 3: Pagination Loop
For each valid parent comment, the method enters a while loop that continues until the API reports is_end or the pagination offset stops changing:
# media_platform/zhihu/client.py (lines 63-70)
while not is_end:
# ... API call logic ...
if offset == child_comment_res.get("offset", offset):
break
is_end = child_comment_res.get("is_end", True)
Step 4: Child Comment API Request
Inside the pagination loop, the method calls get_child_comments, which maps to the platform's child comment endpoint (e.g., Zhihu's /child_comment):
# media_platform/zhihu/client.py (lines 65-66)
child_comment_res = await self.get_child_comments(parment_comment.comment_id, offset, limit)
Step 5: Data Extraction and Aggregation
Raw JSON responses are transformed into structured ZhihuComment objects using the platform's extractor class:
# media_platform/zhihu/client.py (lines 71-72)
sub_comments = self._extractor.extract_comments(content, child_comment_res.get("data"))
all_sub_comments.extend(sub_comments)
Step 6: Optional Callback Execution
If a callback function was provided (useful for streaming results to databases or real-time processing), it is invoked after each batch:
# media_platform/zhihu/client.py (lines 78-79)
if callback:
await callback(sub_comments)
Step 7: Rate Limiting
The method respects the crawl_interval parameter to avoid triggering platform anti-bot measures:
# media_platform/zhihu/client.py (line 82)
await asyncio.sleep(crawl_interval)
Practical Implementation Examples
Enable via Command Line
When running MediaCrawler from the terminal, pass the flag to activate sub-comment crawling:
python main.py --enable-sub-comments true --platform zhihu
The argument parser in cmd_arg/arg.py automatically maps this to config.ENABLE_GET_SUB_COMMENTS.
Programmatic Usage with Callback
For custom pipelines, enable the flag in Python and provide a callback to process sub-comments incrementally:
import asyncio
import config
from media_platform.zhihu.client import ZhiHuClient
async def sub_comment_handler(sub_comments):
"""Process each batch of sub-comments as they arrive."""
print(f"Received {len(sub_comments)} nested replies")
# Insert into database or stream to analytics here
async def main():
# Enable second-level comment crawling
config.ENABLE_GET_SUB_COMMENTS = True
client = ZhiHuClient(
timeout=15,
proxy=None,
headers={"User-Agent": "MediaCrawler"},
playwright_page=None,
cookie_dict={"d_c0": "..."}
)
# Fetch content and all comments (root + sub)
content = await client.get_note_by_keyword("Python", page=1)
comments = await client.get_note_all_comments(
content[0],
crawl_interval=0.5,
callback=sub_comment_handler
)
asyncio.run(main())
In this example, get_note_all_comments automatically invokes get_comments_all_sub_comments because the configuration flag is enabled, returning both root and nested comments in a unified list.
Summary
- Configuration: Set
ENABLE_GET_SUB_COMMENTS = Trueinconfig/base_config.pyto activate the feature globally. - Entry Point: Platform clients use
get_note_all_comments, which internally callsget_comments_all_sub_commentsafter fetching root comments. - Workflow: The sub-comment crawler checks the configuration flag, iterates parent comments with replies, paginates through child API endpoints, extracts data using platform-specific extractors, and supports optional callbacks for streaming.
- Rate Limiting: Built-in
crawl_intervalsleeping prevents API abuse across all platforms. - Consistency: The same
ENABLE_GET_SUB_COMMENTSflag and method patterns work identically across Zhihu, Douyin, Bilibili, Weibo, Tieba, and Kuaishou implementations.
Frequently Asked Questions
What happens if ENABLE_GET_SUB_COMMENTS is False?
When ENABLE_GET_SUB_COMMENTS remains False (the default), the get_comments_all_sub_comments method returns an empty list immediately without making any API calls to fetch child comments. Only root-level comments are retrieved and processed, significantly reducing API quota usage and execution time.
Which platforms support second-level comment crawling?
The implementation in media_platform/zhihu/client.py serves as the reference pattern, but the same configuration flag and workflow are implemented across all supported platforms including Douyin, Bilibili, Weibo, Tieba, and Kuaishou. Each platform provides its own get_child_comments method and extractor, but the control logic remains consistent.
How does the callback function work in sub-comment crawling?
The optional callback parameter in get_comments_all_sub_comments accepts an async function that receives a list of ZhihuComment objects (or platform-specific equivalents) after each pagination batch. This enables real-time processing, such as writing to databases or updating progress bars, without waiting for the entire sub-comment tree to load.
Can I adjust the delay between sub-comment API requests?
Yes, the crawl_interval parameter passed to get_note_all_comments or get_comments_all_sub_comments controls the delay between pagination requests. The method executes await asyncio.sleep(crawl_interval) after each batch, allowing you to respect platform rate limits by setting values typically between 0.5 and 2.0 seconds.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →