How MediaCrawler Handles Comment Pagination with Nested Comments in Tieba

MediaCrawler implements a two-stage pagination system that first iterates through top-level comment pages using total_replay_page, then recursively fetches nested sub-comments calculated from sub_comment_count, with configurable delays and callback hooks.

MediaCrawler is an open-source social media crawling framework that supports multiple platforms including Baidu Tieba. Understanding how it handles comment pagination with nested comments reveals a robust architecture designed to respect rate limits while capturing complete conversation threads.

First-Level Comment Pagination

The primary pagination logic resides in media_platform/tieba/client.py within the BaiduTieBaClient.get_note_all_comments method. This method orchestrates the crawl of top-level comments using a while loop that terminates when either the maximum page count is reached or the caller-specified max_count threshold is exceeded.

The Pagination Loop

The loop control relies on two critical variables: note_detail.total_replay_page (provided by the note's metadata) and the max_count parameter (defaulting to 10). For each iteration, the client increments a current_page counter and validates the stopping condition:

while note_detail.total_replay_page >= current_page and len(result) < max_count:
    api_data = await self._get_pc_page_data(note_id=note_detail.note_id, page=current_page)
    comments = self._page_extractor.extract_tieba_note_parent_comments_from_api(
        api_data, note_detail=note_detail
    )
    result.extend(comments)
    await self.get_comments_all_sub_comments(comments, crawl_interval, callback)
    await asyncio.sleep(crawl_interval)
    current_page += 1

This implementation appears in media_platform/tieba/client.py at lines 52-78.

API Integration and Callback Hooks

Each page request invokes _get_pc_page_data to fetch raw API data, which is then processed through extract_tieba_note_parent_comments_from_api to generate structured comment objects. The method supports an optional callback function invoked after each page fetch, enabling real-time processing or storage without waiting for the entire crawl to complete.

Nested Sub-Comment Pagination

After retrieving each batch of parent comments, MediaCrawler automatically triggers get_comments_all_sub_comments to handle nested replies. This method iterates through every parent comment where sub_comment_count > 0 and calculates the required number of sub-pages.

Calculating Sub-Comment Pages

The maximum number of sub-pages derives from the reply count using integer division: sub_comment_count // 10 + 1. This assumes 10 sub-comments per page, a pattern consistent with Tieba's interface:

for parment_comment in comments:
    if parment_comment.sub_comment_count == 0:
        continue
    current_page = 1
    max_sub_page_num = parment_comment.sub_comment_count // 10 + 1
    while max_sub_page_num >= current_page:
        # Fetching logic continues here

        current_page += 1

Browser-Based Extraction

Unlike the parent comments that use direct API calls, sub-comments are fetched via Playwright to avoid API throttling. The client constructs a specific URL pattern (/p/comment?tid=...&pid=...&fid=...&pn=...) and navigates using self.playwright_page.goto:

sub_comment_url = (
    f"{self._host}/p/comment?tid={parment_comment.note_id}"
    f"&pid={parment_comment.comment_id}&fid={parment_comment.tieba_id}"
    f"&pn={current_page}"
)
await self.playwright_page.goto(sub_comment_url, wait_until="domcontentloaded")
sub_comments = self._page_extractor.extract_tieba_note_sub_comments(
    page_content, parent_comment=parment_comment
)

This browser fallback strategy appears in media_platform/tieba/client.py at lines 112-150. Both pagination stages respect the crawl_interval delay (default 1 second) between requests to prevent rate limiting.

Configuration and Error Handling

The pagination system exposes several configuration points:

  • max_count: Limits the total number of top-level comments retrieved
  • crawl_interval: Controls the delay between page requests (applies to both parent and sub-comment fetches)
  • callback: Optional function receiving (note_id, comments) for custom processing after each page

Error handling is implemented through exception catching that logs failures and breaks the loop, preventing infinite crawls when encountering malformed responses or network issues.

Complete Implementation Example

The following example demonstrates fetching 50 top-level comments with a 0.5-second delay, including automatic nested comment retrieval:

import asyncio
from media_platform.tieba.client import BaiduTieBaClient

async def print_comments(note_id: str, comments):
    print(f"Fetched {len(comments)} comments for note {note_id}")

async def main():
    client = BaiduTieBaClient()
    note = await client.get_note_by_id("123456789")
    
    # Crawl all top-level comments (up to 50) with 0.5s delay

    all_comments = await client.get_note_all_comments(
        note_detail=note,
        crawl_interval=0.5,
        max_count=50,
        callback=print_comments,
    )
    
    print(f"Total top-level comments: {len(all_comments)}")
    
    # Sub-comments are already included via internal calls

    # To fetch separately:

    sub_comments = await client.get_comments_all_sub_comments(
        all_comments, 
        crawl_interval=0.5,
        callback=print_comments
    )
    print(f"Total sub-comments: {len(sub_comments)}")

asyncio.run(main())

Summary

  • MediaCrawler uses a two-stage approach: first paginating through total_replay_page for top-level comments, then calculating sub-pages via sub_comment_count // 10 + 1 for nested replies.
  • Top-level comments are extracted via extract_tieba_note_parent_comments_from_api after calling _get_pc_page_data in the media_platform/tieba/client.py implementation.
  • Nested comments require Playwright browser automation to fetch sub-pages from /p/comment endpoints, processed through extract_tieba_note_sub_comments.
  • Configurable throttling via crawl_interval and optional callback hooks allow real-time processing while respecting rate limits.
  • Automatic recursion occurs within get_note_all_comments, which internally calls get_comments_all_sub_comments for each parent comment batch.

Frequently Asked Questions

How does MediaCrawler determine when to stop paginating comments?

The pagination stops when the current_page exceeds note_detail.total_replay_page for top-level comments, or when the total comment count reaches the caller-specified max_count parameter. For sub-comments, the loop terminates after processing sub_comment_count // 10 + 1 pages.

Why does MediaCrawler use Playwright for sub-comments but API calls for parent comments?

Sub-comments are fetched via Playwright (self.playwright_page.goto) to avoid API throttling and anti-bot mechanisms, while parent comments use direct API calls (_get_pc_page_data) for efficiency. This hybrid approach balances speed with reliability when crawling nested conversation threads.

Can I customize the delay between comment page requests?

Yes, the crawl_interval parameter (default 1.0 seconds) controls the delay between requests for both top-level and sub-comment pagination. Pass this parameter to get_note_all_comments or get_comments_all_sub_comments to adjust the throttling behavior according to your rate limit requirements.

Where is the nested comment pagination logic implemented in the source code?

The nested comment pagination is implemented in media_platform/tieba/client.py at lines 112-150 within the get_comments_all_sub_comments method, while the parent comment pagination resides in lines 52-78 within get_note_all_comments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →