# How MediaCrawler Handles Comment Pagination with Nested Comments in Tieba

> Discover how MediaCrawler tackles comment pagination and nested comments in Tieba. Learn about its two-stage system using total replay page and sub comment count for efficient data retrieval.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-02

---

**MediaCrawler implements a two-stage pagination system that first iterates through top-level comment pages using `total_replay_page`, then recursively fetches nested sub-comments calculated from `sub_comment_count`, with configurable delays and callback hooks.**

MediaCrawler is an open-source social media crawling framework that supports multiple platforms including Baidu Tieba. Understanding how it handles comment pagination with nested comments reveals a robust architecture designed to respect rate limits while capturing complete conversation threads.

## First-Level Comment Pagination

The primary pagination logic resides in [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) within the `BaiduTieBaClient.get_note_all_comments` method. This method orchestrates the crawl of top-level comments using a `while` loop that terminates when either the maximum page count is reached or the caller-specified `max_count` threshold is exceeded.

### The Pagination Loop

The loop control relies on two critical variables: `note_detail.total_replay_page` (provided by the note's metadata) and the `max_count` parameter (defaulting to 10). For each iteration, the client increments a `current_page` counter and validates the stopping condition:

```python
while note_detail.total_replay_page >= current_page and len(result) < max_count:
    api_data = await self._get_pc_page_data(note_id=note_detail.note_id, page=current_page)
    comments = self._page_extractor.extract_tieba_note_parent_comments_from_api(
        api_data, note_detail=note_detail
    )
    result.extend(comments)
    await self.get_comments_all_sub_comments(comments, crawl_interval, callback)
    await asyncio.sleep(crawl_interval)
    current_page += 1

```

This implementation appears in [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) at lines 52-78.

### API Integration and Callback Hooks

Each page request invokes `_get_pc_page_data` to fetch raw API data, which is then processed through `extract_tieba_note_parent_comments_from_api` to generate structured comment objects. The method supports an optional `callback` function invoked after each page fetch, enabling real-time processing or storage without waiting for the entire crawl to complete.

## Nested Sub-Comment Pagination

After retrieving each batch of parent comments, MediaCrawler automatically triggers `get_comments_all_sub_comments` to handle nested replies. This method iterates through every parent comment where `sub_comment_count > 0` and calculates the required number of sub-pages.

### Calculating Sub-Comment Pages

The maximum number of sub-pages derives from the reply count using integer division: `sub_comment_count // 10 + 1`. This assumes 10 sub-comments per page, a pattern consistent with Tieba's interface:

```python
for parment_comment in comments:
    if parment_comment.sub_comment_count == 0:
        continue
    current_page = 1
    max_sub_page_num = parment_comment.sub_comment_count // 10 + 1
    while max_sub_page_num >= current_page:
        # Fetching logic continues here

        current_page += 1

```

### Browser-Based Extraction

Unlike the parent comments that use direct API calls, sub-comments are fetched via Playwright to avoid API throttling. The client constructs a specific URL pattern (`/p/comment?tid=...&pid=...&fid=...&pn=...`) and navigates using `self.playwright_page.goto`:

```python
sub_comment_url = (
    f"{self._host}/p/comment?tid={parment_comment.note_id}"
    f"&pid={parment_comment.comment_id}&fid={parment_comment.tieba_id}"
    f"&pn={current_page}"
)
await self.playwright_page.goto(sub_comment_url, wait_until="domcontentloaded")
sub_comments = self._page_extractor.extract_tieba_note_sub_comments(
    page_content, parent_comment=parment_comment
)

```

This browser fallback strategy appears in [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) at lines 112-150. Both pagination stages respect the `crawl_interval` delay (default 1 second) between requests to prevent rate limiting.

## Configuration and Error Handling

The pagination system exposes several configuration points:

- **`max_count`**: Limits the total number of top-level comments retrieved
- **`crawl_interval`**: Controls the delay between page requests (applies to both parent and sub-comment fetches)
- **`callback`**: Optional function receiving `(note_id, comments)` for custom processing after each page

Error handling is implemented through exception catching that logs failures and breaks the loop, preventing infinite crawls when encountering malformed responses or network issues.

## Complete Implementation Example

The following example demonstrates fetching 50 top-level comments with a 0.5-second delay, including automatic nested comment retrieval:

```python
import asyncio
from media_platform.tieba.client import BaiduTieBaClient

async def print_comments(note_id: str, comments):
    print(f"Fetched {len(comments)} comments for note {note_id}")

async def main():
    client = BaiduTieBaClient()
    note = await client.get_note_by_id("123456789")
    
    # Crawl all top-level comments (up to 50) with 0.5s delay

    all_comments = await client.get_note_all_comments(
        note_detail=note,
        crawl_interval=0.5,
        max_count=50,
        callback=print_comments,
    )
    
    print(f"Total top-level comments: {len(all_comments)}")
    
    # Sub-comments are already included via internal calls

    # To fetch separately:

    sub_comments = await client.get_comments_all_sub_comments(
        all_comments, 
        crawl_interval=0.5,
        callback=print_comments
    )
    print(f"Total sub-comments: {len(sub_comments)}")

asyncio.run(main())

```

## Summary

- **MediaCrawler uses a two-stage approach**: first paginating through `total_replay_page` for top-level comments, then calculating sub-pages via `sub_comment_count // 10 + 1` for nested replies.
- **Top-level comments** are extracted via `extract_tieba_note_parent_comments_from_api` after calling `_get_pc_page_data` in the [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) implementation.
- **Nested comments** require Playwright browser automation to fetch sub-pages from `/p/comment` endpoints, processed through `extract_tieba_note_sub_comments`.
- **Configurable throttling** via `crawl_interval` and optional `callback` hooks allow real-time processing while respecting rate limits.
- **Automatic recursion** occurs within `get_note_all_comments`, which internally calls `get_comments_all_sub_comments` for each parent comment batch.

## Frequently Asked Questions

### How does MediaCrawler determine when to stop paginating comments?

The pagination stops when the `current_page` exceeds `note_detail.total_replay_page` for top-level comments, or when the total comment count reaches the caller-specified `max_count` parameter. For sub-comments, the loop terminates after processing `sub_comment_count // 10 + 1` pages.

### Why does MediaCrawler use Playwright for sub-comments but API calls for parent comments?

Sub-comments are fetched via Playwright (`self.playwright_page.goto`) to avoid API throttling and anti-bot mechanisms, while parent comments use direct API calls (`_get_pc_page_data`) for efficiency. This hybrid approach balances speed with reliability when crawling nested conversation threads.

### Can I customize the delay between comment page requests?

Yes, the `crawl_interval` parameter (default 1.0 seconds) controls the delay between requests for both top-level and sub-comment pagination. Pass this parameter to `get_note_all_comments` or `get_comments_all_sub_comments` to adjust the throttling behavior according to your rate limit requirements.

### Where is the nested comment pagination logic implemented in the source code?

The nested comment pagination is implemented in [`media_platform/tieba/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/client.py) at lines 112-150 within the `get_comments_all_sub_comments` method, while the parent comment pagination resides in lines 52-78 within `get_note_all_comments`.