# How to Implement Second-Level Comment Crawling (Nested Replies) in MediaCrawler

> Learn how to implement second-level comment crawling in MediaCrawler using the get_comments_all_sub_comments coroutine. Discover how to detect, paginate, and fetch nested replies efficiently.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-14

---

**MediaCrawler implements second-level comment crawling through a unified `get_comments_all_sub_comments` coroutine that each platform-specific client overrides to detect, paginate, fetch, and extract nested replies beneath parent comments.**

The MediaCrawler open-source framework provides a consistent abstraction for crawling nested comment threads across multiple Chinese social platforms. This guide breaks down the implementation pattern, platform-specific details, and practical configuration based on the source code in `NanmiCoder/MediaCrawler`.

## How Sub-Comment Crawling Works

MediaCrawler treats nested replies as **sub-comments** with a standardized seven-step flow implemented in every platform client:

1. **Detect availability** – check `sub_comment_count` or equivalent field on the parent comment
2. **Calculate pagination** – determine total pages from count and page size
3. **Build request URL** – construct platform-specific endpoint or GraphQL query
4. **Fetch page** – execute HTTP request via `session.get` or Playwright navigation
5. **Extract data** – delegate to `self._page_extractor.extract_*_sub_comments`
6. **Invoke callback** – pass results to user-supplied handler for real-time processing
7. **Aggregate results** – collect all sub-comments into a master list

This design cleanly separates first-level and nested crawling, controlled by the `sub_comment_mode` configuration flag in each platform's config.

## Platform-Specific Implementations

### Zhihu: Cursor-Based Pagination with Extractor Delegation

In [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) (lines 329–384), `ZhihuClient.get_comments_all_sub_comments` checks `parment_comment.sub_comment_count`, loops through child pages, and calls `_extractor.extract_comments` for each response.

```python

# media_platform/zhihu/client.py#L329-L384

async def get_comments_all_sub_comments(
    self, 
    note_id: str, 
    comment: Any, 
    callback: Optional[Callable] = None
) -> List[Dict]:
    """Fetch all sub-comments for a Zhihu parent comment."""
    if not comment.sub_comment_count:
        return []
    
    sub_comments = []
    # Pagination logic using page offsets

    for page in range(1, (comment.sub_comment_count // 10) + 2):
        # Build and fetch sub-comment page URL

        ...
        if callback:
            await callback(note_id, batch)
    
    return sub_comments

```

### XiaoHongShu: Cursor Pagination with `has_more` Flags

The XiaoHongShu implementation in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) (lines 517–576) validates `sub_comment_mode`, retrieves `sub_comments` directly when embedded, and handles pagination via `sub_comment_cursor` and `has_more` flags for API cursor-based traversal.

### Weibo: Embedded Sub-Comments Array

Weibo's lightweight implementation in [`media_platform/weibo/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/client.py) (lines 247–256) reads `comment.get("comments")` directly from the parent comment object, invokes the optional callback, and returns the aggregated list without additional API calls for basic cases.

```python

# media_platform/weibo/client.py#L247-L256

async def get_comments_all_sub_comments(
    self,
    note_id: str,
    comment: Dict,
    callback: Optional[Callable] = None
) -> List[Dict]:
    sub_comments = comment.get("comments", [])
    if callback and sub_comments:
        await callback(note_id, sub_comments)
    return sub_comments

```

### Baidu Tieba: Playwright Navigation with URL Construction

Tieba's implementation in [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) (lines 516–595) calculates max sub-pages from `sub_comment_count`, constructs sub-comment URLs using `/comment/page?tid=…` endpoints, uses **Playwright** for browser navigation, and extracts sub-comments from rendered pages.

### KuaiShou: GraphQL Query Execution

KuaiShou leverages GraphQL in [`media_platform/kuaishou/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/client.py) (lines 337–363). It checks `hasSubComments` flag, queries the `vision_sub_comment_list.graphql` operation, and processes paginated GraphQL responses.

## End-to-End Integration Flow

When `get_note_all_comments` is invoked, the execution flow spans core and client layers:

```python

# Conceptual flow across ZhihuCore / ZhihuClient

async def get_note_all_comments(self, note_id: str) -> List[Comment]:
    # 1. Fetch first-level comments

    first_level = await self.client.get_note_comments(note_id)
    
    all_comments = []
    for comment in first_level:
        all_comments.append(comment)
        
        # 2. Conditionally fetch sub-comments

        if self.config.sub_comment_mode and comment.sub_comment_count:
            sub_comments = await self.client.get_comments_all_sub_comments(
                note_id, comment, callback=self._on_sub_comments
            )
            comment.sub_comments = sub_comments
    
    return all_comments

```

The **callback mechanism** enables real-time streaming—each batch of sub-comments can be persisted or processed without waiting for full completion.

## Practical Implementation Examples

### Enabling Sub-Comment Mode in Zhihu

```python
from media_platform.zhihu.core import ZhihuCore
from media_platform.zhihu.config import zhihu_config

core = ZhihuCore(zhihu_config)
note_id = "1234567890"

# Enable nested reply crawling (default: True)

zhihu_config.sub_comment_mode = True

# Fetch complete comment hierarchy

comments = await core.get_note_all_comments(note_id)

# Traverse nested structure

for comment in comments:
    print(f"Comment: {comment.content}")
    for sub in comment.sub_comments or []:
        print(f"  ↳ Reply: {sub.content}")

```

### Manual Sub-Comment Fetching in Tieba

```python
from media_platform.tieba.core import TiebaCore
from media_platform.tieba.config import tieba_config

core = TiebaCore(tieba_config)
note_id = "987654321"

# Get first-level comments

first_level = await core.get_note_all_comments(note_id)

# Selectively drill into specific parent

parent = first_level[0]
sub_comments = await core.client.get_comments_all_sub_comments(
    note_id=note_id,
    comment=parent,
    callback=None  # Optional: provide async callback for streaming

)

print(f"Parent [{parent.username}]: {parent.content}")
for sub in sub_comments:
    print(f"  ↳ {sub.username}: {sub.content}")

```

### Custom Callback for Real-Time Processing

```python
async def persist_sub_comments(note_id: str, sub_comments: List[Dict]):
    """Custom callback invoked per sub-comment batch."""
    for sub in sub_comments:
        await database.insert({
            "note_id": note_id,
            "parent_id": sub.get("parent_id"),
            "content": sub.get("content"),
            "create_time": sub.get("create_time")
        })

# Usage during crawl

await client.get_comments_all_sub_comments(
    note_id, parent_comment, callback=persist_sub_comments
)

```

## Configuration and Performance Considerations

| Setting | Location | Effect |
|--------|----------|--------|
| `sub_comment_mode` | `media_platform/{platform}/config.py` | Master toggle enabling/disabling all sub-comment fetching |
| Page size constants | Platform client files | Controls sub-comments per API request (typically 10) |
| Playwright timeout | [`tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tieba/client.py) | Adjust for slow-rendering sub-comment pages |

**Performance note:** Sub-comment crawling multiplies request volume proportionally to reply density. For high-traffic posts, implement rate limiting and consider callback-based streaming to reduce memory pressure.

## Key Source Files

| File | Purpose | Line Range |
|------|---------|-----------|
| [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) | Zhihu nested comment implementation | 329–384 |
| [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) | XiaoHongShu cursor-based pagination | 517–576 |
| [`media_platform/weibo/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/client.py) | Weibo embedded sub-comment extraction | 247–256 |
| [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) | Tieba Playwright-based crawling | 516–595 |
| [`media_platform/kuaishou/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/client.py) | KuaiShou GraphQL queries | 337–363 |
| [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) | Orchestration layer integrating sub-comments | Full file |

## Summary

- **Unified interface:** Every platform implements `get_comments_all_sub_comments` with identical signature: `(note_id, comment, callback) -> List[Dict]`
- **Detection pattern:** Check `sub_comment_count`, `hasSubComments`, or equivalent field before pagination
- **Pagination strategies:** Offset-based (Zhihu, Tieba), cursor-based (XiaoHongShu), GraphQL pagination (KuaiShou), or embedded arrays (Weibo)
- **Enable via config:** Set `sub_comment_mode = True` in platform config to activate automatic nested crawling during `get_note_all_comments`
- **Extensibility:** Callback parameter supports real-time processing without blocking full aggregation

## Frequently Asked Questions

### How do I disable second-level comment crawling to save API quota?

Set `sub_comment_mode = False` in your platform configuration before initializing the core. In [`media_platform/zhihu/config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/config.py) or equivalent, this prevents `get_note_all_comments` from invoking `get_comments_all_sub_comments` entirely.

### Why does Tieba use Playwright while other platforms use direct HTTP requests?

Baidu Tieba's sub-comment endpoints require JavaScript execution to render comment data. The implementation in [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) navigates to constructed URLs via `self.playwright_page.goto` and extracts from the rendered DOM, whereas Zhihu, XiaoHongShu, and KuaiShou expose sub-comments through JSON APIs.

### Can I fetch sub-comments for a specific parent comment without crawling all comments?

Yes. Bypass the core orchestration and call `client.get_comments_all_sub_comments` directly with the parent comment object and optional callback. This is useful for targeted backfill or incremental updates on high-engagement threads.

### How does the callback parameter improve performance for large threads?

The callback receives each batch of sub-comments immediately after extraction, enabling streaming persistence to databases or message queues. Without a callback, all sub-comments accumulate in memory until the full parent comment's replies are fetched—potentially problematic for viral posts with thousands of nested replies.