How to Implement Second-Level Comment Crawling (Nested Replies) in MediaCrawler

MediaCrawler implements second-level comment crawling through a unified get_comments_all_sub_comments coroutine that each platform-specific client overrides to detect, paginate, fetch, and extract nested replies beneath parent comments.

The MediaCrawler open-source framework provides a consistent abstraction for crawling nested comment threads across multiple Chinese social platforms. This guide breaks down the implementation pattern, platform-specific details, and practical configuration based on the source code in NanmiCoder/MediaCrawler.

How Sub-Comment Crawling Works

MediaCrawler treats nested replies as sub-comments with a standardized seven-step flow implemented in every platform client:

  1. Detect availability – check sub_comment_count or equivalent field on the parent comment
  2. Calculate pagination – determine total pages from count and page size
  3. Build request URL – construct platform-specific endpoint or GraphQL query
  4. Fetch page – execute HTTP request via session.get or Playwright navigation
  5. Extract data – delegate to self._page_extractor.extract_*_sub_comments
  6. Invoke callback – pass results to user-supplied handler for real-time processing
  7. Aggregate results – collect all sub-comments into a master list

This design cleanly separates first-level and nested crawling, controlled by the sub_comment_mode configuration flag in each platform's config.

Platform-Specific Implementations

Zhihu: Cursor-Based Pagination with Extractor Delegation

In media_platform/zhihu/client.py (lines 329–384), ZhihuClient.get_comments_all_sub_comments checks parment_comment.sub_comment_count, loops through child pages, and calls _extractor.extract_comments for each response.


# media_platform/zhihu/client.py#L329-L384

async def get_comments_all_sub_comments(
    self, 
    note_id: str, 
    comment: Any, 
    callback: Optional[Callable] = None
) -> List[Dict]:
    """Fetch all sub-comments for a Zhihu parent comment."""
    if not comment.sub_comment_count:
        return []
    
    sub_comments = []
    # Pagination logic using page offsets

    for page in range(1, (comment.sub_comment_count // 10) + 2):
        # Build and fetch sub-comment page URL

        ...
        if callback:
            await callback(note_id, batch)
    
    return sub_comments

XiaoHongShu: Cursor Pagination with has_more Flags

The XiaoHongShu implementation in media_platform/xhs/client.py (lines 517–576) validates sub_comment_mode, retrieves sub_comments directly when embedded, and handles pagination via sub_comment_cursor and has_more flags for API cursor-based traversal.

Weibo: Embedded Sub-Comments Array

Weibo's lightweight implementation in media_platform/weibo/client.py (lines 247–256) reads comment.get("comments") directly from the parent comment object, invokes the optional callback, and returns the aggregated list without additional API calls for basic cases.


# media_platform/weibo/client.py#L247-L256

async def get_comments_all_sub_comments(
    self,
    note_id: str,
    comment: Dict,
    callback: Optional[Callable] = None
) -> List[Dict]:
    sub_comments = comment.get("comments", [])
    if callback and sub_comments:
        await callback(note_id, sub_comments)
    return sub_comments

Baidu Tieba: Playwright Navigation with URL Construction

Tieba's implementation in media_platform/tieba/client.py (lines 516–595) calculates max sub-pages from sub_comment_count, constructs sub-comment URLs using /comment/page?tid=… endpoints, uses Playwright for browser navigation, and extracts sub-comments from rendered pages.

KuaiShou: GraphQL Query Execution

KuaiShou leverages GraphQL in media_platform/kuaishou/client.py (lines 337–363). It checks hasSubComments flag, queries the vision_sub_comment_list.graphql operation, and processes paginated GraphQL responses.

End-to-End Integration Flow

When get_note_all_comments is invoked, the execution flow spans core and client layers:


# Conceptual flow across ZhihuCore / ZhihuClient

async def get_note_all_comments(self, note_id: str) -> List[Comment]:
    # 1. Fetch first-level comments

    first_level = await self.client.get_note_comments(note_id)
    
    all_comments = []
    for comment in first_level:
        all_comments.append(comment)
        
        # 2. Conditionally fetch sub-comments

        if self.config.sub_comment_mode and comment.sub_comment_count:
            sub_comments = await self.client.get_comments_all_sub_comments(
                note_id, comment, callback=self._on_sub_comments
            )
            comment.sub_comments = sub_comments
    
    return all_comments

The callback mechanism enables real-time streaming—each batch of sub-comments can be persisted or processed without waiting for full completion.

Practical Implementation Examples

Enabling Sub-Comment Mode in Zhihu

from media_platform.zhihu.core import ZhihuCore
from media_platform.zhihu.config import zhihu_config

core = ZhihuCore(zhihu_config)
note_id = "1234567890"

# Enable nested reply crawling (default: True)

zhihu_config.sub_comment_mode = True

# Fetch complete comment hierarchy

comments = await core.get_note_all_comments(note_id)

# Traverse nested structure

for comment in comments:
    print(f"Comment: {comment.content}")
    for sub in comment.sub_comments or []:
        print(f"  ↳ Reply: {sub.content}")

Manual Sub-Comment Fetching in Tieba

from media_platform.tieba.core import TiebaCore
from media_platform.tieba.config import tieba_config

core = TiebaCore(tieba_config)
note_id = "987654321"

# Get first-level comments

first_level = await core.get_note_all_comments(note_id)

# Selectively drill into specific parent

parent = first_level[0]
sub_comments = await core.client.get_comments_all_sub_comments(
    note_id=note_id,
    comment=parent,
    callback=None  # Optional: provide async callback for streaming

)

print(f"Parent [{parent.username}]: {parent.content}")
for sub in sub_comments:
    print(f"  ↳ {sub.username}: {sub.content}")

Custom Callback for Real-Time Processing

async def persist_sub_comments(note_id: str, sub_comments: List[Dict]):
    """Custom callback invoked per sub-comment batch."""
    for sub in sub_comments:
        await database.insert({
            "note_id": note_id,
            "parent_id": sub.get("parent_id"),
            "content": sub.get("content"),
            "create_time": sub.get("create_time")
        })

# Usage during crawl

await client.get_comments_all_sub_comments(
    note_id, parent_comment, callback=persist_sub_comments
)

Configuration and Performance Considerations

Setting Location Effect
sub_comment_mode media_platform/{platform}/config.py Master toggle enabling/disabling all sub-comment fetching
Page size constants Platform client files Controls sub-comments per API request (typically 10)
Playwright timeout tieba/client.py Adjust for slow-rendering sub-comment pages

Performance note: Sub-comment crawling multiplies request volume proportionally to reply density. For high-traffic posts, implement rate limiting and consider callback-based streaming to reduce memory pressure.

Key Source Files

File Purpose Line Range
media_platform/zhihu/client.py Zhihu nested comment implementation 329–384
media_platform/xhs/client.py XiaoHongShu cursor-based pagination 517–576
media_platform/weibo/client.py Weibo embedded sub-comment extraction 247–256
media_platform/tieba/client.py Tieba Playwright-based crawling 516–595
media_platform/kuaishou/client.py KuaiShou GraphQL queries 337–363
media_platform/zhihu/core.py Orchestration layer integrating sub-comments Full file

Summary

  • Unified interface: Every platform implements get_comments_all_sub_comments with identical signature: (note_id, comment, callback) -> List[Dict]
  • Detection pattern: Check sub_comment_count, hasSubComments, or equivalent field before pagination
  • Pagination strategies: Offset-based (Zhihu, Tieba), cursor-based (XiaoHongShu), GraphQL pagination (KuaiShou), or embedded arrays (Weibo)
  • Enable via config: Set sub_comment_mode = True in platform config to activate automatic nested crawling during get_note_all_comments
  • Extensibility: Callback parameter supports real-time processing without blocking full aggregation

Frequently Asked Questions

How do I disable second-level comment crawling to save API quota?

Set sub_comment_mode = False in your platform configuration before initializing the core. In media_platform/zhihu/config.py or equivalent, this prevents get_note_all_comments from invoking get_comments_all_sub_comments entirely.

Why does Tieba use Playwright while other platforms use direct HTTP requests?

Baidu Tieba's sub-comment endpoints require JavaScript execution to render comment data. The implementation in media_platform/tieba/client.py navigates to constructed URLs via self.playwright_page.goto and extracts from the rendered DOM, whereas Zhihu, XiaoHongShu, and KuaiShou expose sub-comments through JSON APIs.

Can I fetch sub-comments for a specific parent comment without crawling all comments?

Yes. Bypass the core orchestration and call client.get_comments_all_sub_comments directly with the parent comment object and optional callback. This is useful for targeted backfill or incremental updates on high-engagement threads.

How does the callback parameter improve performance for large threads?

The callback receives each batch of sub-comments immediately after extraction, enabling streaming persistence to databases or message queues. Without a callback, all sub-comments accumulate in memory until the full parent comment's replies are fetched—potentially problematic for viral posts with thousands of nested replies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →