How to Implement Second-Level Comment Crawling (Nested Replies) in MediaCrawler
MediaCrawler implements second-level comment crawling through a unified get_comments_all_sub_comments coroutine that each platform-specific client overrides to detect, paginate, fetch, and extract nested replies beneath parent comments.
The MediaCrawler open-source framework provides a consistent abstraction for crawling nested comment threads across multiple Chinese social platforms. This guide breaks down the implementation pattern, platform-specific details, and practical configuration based on the source code in NanmiCoder/MediaCrawler.
How Sub-Comment Crawling Works
MediaCrawler treats nested replies as sub-comments with a standardized seven-step flow implemented in every platform client:
- Detect availability – check
sub_comment_countor equivalent field on the parent comment - Calculate pagination – determine total pages from count and page size
- Build request URL – construct platform-specific endpoint or GraphQL query
- Fetch page – execute HTTP request via
session.getor Playwright navigation - Extract data – delegate to
self._page_extractor.extract_*_sub_comments - Invoke callback – pass results to user-supplied handler for real-time processing
- Aggregate results – collect all sub-comments into a master list
This design cleanly separates first-level and nested crawling, controlled by the sub_comment_mode configuration flag in each platform's config.
Platform-Specific Implementations
Zhihu: Cursor-Based Pagination with Extractor Delegation
In media_platform/zhihu/client.py (lines 329–384), ZhihuClient.get_comments_all_sub_comments checks parment_comment.sub_comment_count, loops through child pages, and calls _extractor.extract_comments for each response.
# media_platform/zhihu/client.py#L329-L384
async def get_comments_all_sub_comments(
self,
note_id: str,
comment: Any,
callback: Optional[Callable] = None
) -> List[Dict]:
"""Fetch all sub-comments for a Zhihu parent comment."""
if not comment.sub_comment_count:
return []
sub_comments = []
# Pagination logic using page offsets
for page in range(1, (comment.sub_comment_count // 10) + 2):
# Build and fetch sub-comment page URL
...
if callback:
await callback(note_id, batch)
return sub_comments
XiaoHongShu: Cursor Pagination with has_more Flags
The XiaoHongShu implementation in media_platform/xhs/client.py (lines 517–576) validates sub_comment_mode, retrieves sub_comments directly when embedded, and handles pagination via sub_comment_cursor and has_more flags for API cursor-based traversal.
Weibo: Embedded Sub-Comments Array
Weibo's lightweight implementation in media_platform/weibo/client.py (lines 247–256) reads comment.get("comments") directly from the parent comment object, invokes the optional callback, and returns the aggregated list without additional API calls for basic cases.
# media_platform/weibo/client.py#L247-L256
async def get_comments_all_sub_comments(
self,
note_id: str,
comment: Dict,
callback: Optional[Callable] = None
) -> List[Dict]:
sub_comments = comment.get("comments", [])
if callback and sub_comments:
await callback(note_id, sub_comments)
return sub_comments
Baidu Tieba: Playwright Navigation with URL Construction
Tieba's implementation in media_platform/tieba/client.py (lines 516–595) calculates max sub-pages from sub_comment_count, constructs sub-comment URLs using /comment/page?tid=… endpoints, uses Playwright for browser navigation, and extracts sub-comments from rendered pages.
KuaiShou: GraphQL Query Execution
KuaiShou leverages GraphQL in media_platform/kuaishou/client.py (lines 337–363). It checks hasSubComments flag, queries the vision_sub_comment_list.graphql operation, and processes paginated GraphQL responses.
End-to-End Integration Flow
When get_note_all_comments is invoked, the execution flow spans core and client layers:
# Conceptual flow across ZhihuCore / ZhihuClient
async def get_note_all_comments(self, note_id: str) -> List[Comment]:
# 1. Fetch first-level comments
first_level = await self.client.get_note_comments(note_id)
all_comments = []
for comment in first_level:
all_comments.append(comment)
# 2. Conditionally fetch sub-comments
if self.config.sub_comment_mode and comment.sub_comment_count:
sub_comments = await self.client.get_comments_all_sub_comments(
note_id, comment, callback=self._on_sub_comments
)
comment.sub_comments = sub_comments
return all_comments
The callback mechanism enables real-time streaming—each batch of sub-comments can be persisted or processed without waiting for full completion.
Practical Implementation Examples
Enabling Sub-Comment Mode in Zhihu
from media_platform.zhihu.core import ZhihuCore
from media_platform.zhihu.config import zhihu_config
core = ZhihuCore(zhihu_config)
note_id = "1234567890"
# Enable nested reply crawling (default: True)
zhihu_config.sub_comment_mode = True
# Fetch complete comment hierarchy
comments = await core.get_note_all_comments(note_id)
# Traverse nested structure
for comment in comments:
print(f"Comment: {comment.content}")
for sub in comment.sub_comments or []:
print(f" ↳ Reply: {sub.content}")
Manual Sub-Comment Fetching in Tieba
from media_platform.tieba.core import TiebaCore
from media_platform.tieba.config import tieba_config
core = TiebaCore(tieba_config)
note_id = "987654321"
# Get first-level comments
first_level = await core.get_note_all_comments(note_id)
# Selectively drill into specific parent
parent = first_level[0]
sub_comments = await core.client.get_comments_all_sub_comments(
note_id=note_id,
comment=parent,
callback=None # Optional: provide async callback for streaming
)
print(f"Parent [{parent.username}]: {parent.content}")
for sub in sub_comments:
print(f" ↳ {sub.username}: {sub.content}")
Custom Callback for Real-Time Processing
async def persist_sub_comments(note_id: str, sub_comments: List[Dict]):
"""Custom callback invoked per sub-comment batch."""
for sub in sub_comments:
await database.insert({
"note_id": note_id,
"parent_id": sub.get("parent_id"),
"content": sub.get("content"),
"create_time": sub.get("create_time")
})
# Usage during crawl
await client.get_comments_all_sub_comments(
note_id, parent_comment, callback=persist_sub_comments
)
Configuration and Performance Considerations
| Setting | Location | Effect |
|---|---|---|
sub_comment_mode |
media_platform/{platform}/config.py |
Master toggle enabling/disabling all sub-comment fetching |
| Page size constants | Platform client files | Controls sub-comments per API request (typically 10) |
| Playwright timeout | tieba/client.py |
Adjust for slow-rendering sub-comment pages |
Performance note: Sub-comment crawling multiplies request volume proportionally to reply density. For high-traffic posts, implement rate limiting and consider callback-based streaming to reduce memory pressure.
Key Source Files
| File | Purpose | Line Range |
|---|---|---|
media_platform/zhihu/client.py |
Zhihu nested comment implementation | 329–384 |
media_platform/xhs/client.py |
XiaoHongShu cursor-based pagination | 517–576 |
media_platform/weibo/client.py |
Weibo embedded sub-comment extraction | 247–256 |
media_platform/tieba/client.py |
Tieba Playwright-based crawling | 516–595 |
media_platform/kuaishou/client.py |
KuaiShou GraphQL queries | 337–363 |
media_platform/zhihu/core.py |
Orchestration layer integrating sub-comments | Full file |
Summary
- Unified interface: Every platform implements
get_comments_all_sub_commentswith identical signature:(note_id, comment, callback) -> List[Dict] - Detection pattern: Check
sub_comment_count,hasSubComments, or equivalent field before pagination - Pagination strategies: Offset-based (Zhihu, Tieba), cursor-based (XiaoHongShu), GraphQL pagination (KuaiShou), or embedded arrays (Weibo)
- Enable via config: Set
sub_comment_mode = Truein platform config to activate automatic nested crawling duringget_note_all_comments - Extensibility: Callback parameter supports real-time processing without blocking full aggregation
Frequently Asked Questions
How do I disable second-level comment crawling to save API quota?
Set sub_comment_mode = False in your platform configuration before initializing the core. In media_platform/zhihu/config.py or equivalent, this prevents get_note_all_comments from invoking get_comments_all_sub_comments entirely.
Why does Tieba use Playwright while other platforms use direct HTTP requests?
Baidu Tieba's sub-comment endpoints require JavaScript execution to render comment data. The implementation in media_platform/tieba/client.py navigates to constructed URLs via self.playwright_page.goto and extracts from the rendered DOM, whereas Zhihu, XiaoHongShu, and KuaiShou expose sub-comments through JSON APIs.
Can I fetch sub-comments for a specific parent comment without crawling all comments?
Yes. Bypass the core orchestration and call client.get_comments_all_sub_comments directly with the parent comment object and optional callback. This is useful for targeted backfill or incremental updates on high-engagement threads.
How does the callback parameter improve performance for large threads?
The callback receives each batch of sub-comments immediately after extraction, enabling streaming persistence to databases or message queues. Without a callback, all sub-comments accumulate in memory until the full parent comment's replies are fetched—potentially problematic for viral posts with thousands of nested replies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →