How MediaCrawler Handles Comment Pagination with Nested Comments in Tieba
MediaCrawler implements a two-stage pagination system that first iterates through top-level comment pages using total_replay_page, then recursively fetches nested sub-comments calculated from sub_comment_count, with configurable delays and callback hooks.
MediaCrawler is an open-source social media crawling framework that supports multiple platforms including Baidu Tieba. Understanding how it handles comment pagination with nested comments reveals a robust architecture designed to respect rate limits while capturing complete conversation threads.
First-Level Comment Pagination
The primary pagination logic resides in media_platform/tieba/client.py within the BaiduTieBaClient.get_note_all_comments method. This method orchestrates the crawl of top-level comments using a while loop that terminates when either the maximum page count is reached or the caller-specified max_count threshold is exceeded.
The Pagination Loop
The loop control relies on two critical variables: note_detail.total_replay_page (provided by the note's metadata) and the max_count parameter (defaulting to 10). For each iteration, the client increments a current_page counter and validates the stopping condition:
while note_detail.total_replay_page >= current_page and len(result) < max_count:
api_data = await self._get_pc_page_data(note_id=note_detail.note_id, page=current_page)
comments = self._page_extractor.extract_tieba_note_parent_comments_from_api(
api_data, note_detail=note_detail
)
result.extend(comments)
await self.get_comments_all_sub_comments(comments, crawl_interval, callback)
await asyncio.sleep(crawl_interval)
current_page += 1
This implementation appears in media_platform/tieba/client.py at lines 52-78.
API Integration and Callback Hooks
Each page request invokes _get_pc_page_data to fetch raw API data, which is then processed through extract_tieba_note_parent_comments_from_api to generate structured comment objects. The method supports an optional callback function invoked after each page fetch, enabling real-time processing or storage without waiting for the entire crawl to complete.
Nested Sub-Comment Pagination
After retrieving each batch of parent comments, MediaCrawler automatically triggers get_comments_all_sub_comments to handle nested replies. This method iterates through every parent comment where sub_comment_count > 0 and calculates the required number of sub-pages.
Calculating Sub-Comment Pages
The maximum number of sub-pages derives from the reply count using integer division: sub_comment_count // 10 + 1. This assumes 10 sub-comments per page, a pattern consistent with Tieba's interface:
for parment_comment in comments:
if parment_comment.sub_comment_count == 0:
continue
current_page = 1
max_sub_page_num = parment_comment.sub_comment_count // 10 + 1
while max_sub_page_num >= current_page:
# Fetching logic continues here
current_page += 1
Browser-Based Extraction
Unlike the parent comments that use direct API calls, sub-comments are fetched via Playwright to avoid API throttling. The client constructs a specific URL pattern (/p/comment?tid=...&pid=...&fid=...&pn=...) and navigates using self.playwright_page.goto:
sub_comment_url = (
f"{self._host}/p/comment?tid={parment_comment.note_id}"
f"&pid={parment_comment.comment_id}&fid={parment_comment.tieba_id}"
f"&pn={current_page}"
)
await self.playwright_page.goto(sub_comment_url, wait_until="domcontentloaded")
sub_comments = self._page_extractor.extract_tieba_note_sub_comments(
page_content, parent_comment=parment_comment
)
This browser fallback strategy appears in media_platform/tieba/client.py at lines 112-150. Both pagination stages respect the crawl_interval delay (default 1 second) between requests to prevent rate limiting.
Configuration and Error Handling
The pagination system exposes several configuration points:
max_count: Limits the total number of top-level comments retrievedcrawl_interval: Controls the delay between page requests (applies to both parent and sub-comment fetches)callback: Optional function receiving(note_id, comments)for custom processing after each page
Error handling is implemented through exception catching that logs failures and breaks the loop, preventing infinite crawls when encountering malformed responses or network issues.
Complete Implementation Example
The following example demonstrates fetching 50 top-level comments with a 0.5-second delay, including automatic nested comment retrieval:
import asyncio
from media_platform.tieba.client import BaiduTieBaClient
async def print_comments(note_id: str, comments):
print(f"Fetched {len(comments)} comments for note {note_id}")
async def main():
client = BaiduTieBaClient()
note = await client.get_note_by_id("123456789")
# Crawl all top-level comments (up to 50) with 0.5s delay
all_comments = await client.get_note_all_comments(
note_detail=note,
crawl_interval=0.5,
max_count=50,
callback=print_comments,
)
print(f"Total top-level comments: {len(all_comments)}")
# Sub-comments are already included via internal calls
# To fetch separately:
sub_comments = await client.get_comments_all_sub_comments(
all_comments,
crawl_interval=0.5,
callback=print_comments
)
print(f"Total sub-comments: {len(sub_comments)}")
asyncio.run(main())
Summary
- MediaCrawler uses a two-stage approach: first paginating through
total_replay_pagefor top-level comments, then calculating sub-pages viasub_comment_count // 10 + 1for nested replies. - Top-level comments are extracted via
extract_tieba_note_parent_comments_from_apiafter calling_get_pc_page_datain themedia_platform/tieba/client.pyimplementation. - Nested comments require Playwright browser automation to fetch sub-pages from
/p/commentendpoints, processed throughextract_tieba_note_sub_comments. - Configurable throttling via
crawl_intervaland optionalcallbackhooks allow real-time processing while respecting rate limits. - Automatic recursion occurs within
get_note_all_comments, which internally callsget_comments_all_sub_commentsfor each parent comment batch.
Frequently Asked Questions
How does MediaCrawler determine when to stop paginating comments?
The pagination stops when the current_page exceeds note_detail.total_replay_page for top-level comments, or when the total comment count reaches the caller-specified max_count parameter. For sub-comments, the loop terminates after processing sub_comment_count // 10 + 1 pages.
Why does MediaCrawler use Playwright for sub-comments but API calls for parent comments?
Sub-comments are fetched via Playwright (self.playwright_page.goto) to avoid API throttling and anti-bot mechanisms, while parent comments use direct API calls (_get_pc_page_data) for efficiency. This hybrid approach balances speed with reliability when crawling nested conversation threads.
Can I customize the delay between comment page requests?
Yes, the crawl_interval parameter (default 1.0 seconds) controls the delay between requests for both top-level and sub-comment pagination. Pass this parameter to get_note_all_comments or get_comments_all_sub_comments to adjust the throttling behavior according to your rate limit requirements.
Where is the nested comment pagination logic implemented in the source code?
The nested comment pagination is implemented in media_platform/tieba/client.py at lines 112-150 within the get_comments_all_sub_comments method, while the parent comment pagination resides in lines 52-78 within get_note_all_comments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →