How MediaCrawler Implements Rate Limiting: Pagination, Sleep Intervals, and Jitter

MediaCrawler controls request rates through a combination of per-platform pagination limits, configurable global sleep intervals, random jitter delays, and cumulative note caps defined in config/base_config.py.

MediaCrawler is an open-source multi-platform content scraper supporting Weibo, Zhihu, Xiaohongshu, Douyin, Bilibili, and others. To avoid triggering anti-bot protections and respect API boundaries, the project implements a defensive rate limiting strategy that blends static page-size constraints with dynamic sleep controls. This article examines the specific implementation details found in the NanmiCoder/MediaCrawler repository.

Per-Platform Pagination Limits

Each platform crawler defines a hardcoded constant representing the maximum number of items the target service returns per API page. These values act as the fundamental throttle for batch retrieval.

These limits are consumed within pagination loops that compare the product of page * limit against CRAWLER_MAX_NOTES_COUNT to ensure the crawler never exceeds the configured total item count.

Configurable Sleep Intervals

The global configuration file config/base_config.py defines CRAWLER_MAX_SLEEP_SEC, which specifies the mandatory pause duration between requests. After each network call, the crawler executes await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) to throttle the overall request rate.

In the Weibo crawler at line 452, this appears with an explicit comment: "Sleep after request to avoid rate limiting" → media_platform/weibo/core.py#L452. The same pattern appears in Zhihu’s core loop at line 189 → media_platform/zhihu/core.py#L189, and consistently across other platform implementations including Tieba, Kuaishou, Douyin, and Bilibili.


# Example: Applying global sleep after each page fetch

page = start_page
while (page - start_page + 1) * zhihu_limit_count <= config.CRAWLER_MAX_NOTES_COUNT:
    notes = await client.get_notes(..., limit=zhihu_limit_count)
    await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)  # Global rate limit

    page += 1

Random Jitter for Burst Mitigation

Certain platforms introduce micro-delays to break up bursty traffic patterns that might trigger behavioral detection. Specifically, the Xiaohongshu crawler injects a random sub-second pause using await asyncio.sleep(random.random()) at line 493 → media_platform/xhs/core.py#L493. This jitter adds entropy to request timing, making the traffic appear less automated.


# Example: Random jitter to avoid pattern detection

await client.fetch_note_detail(...)
await asyncio.sleep(random.random())  # Tiny random delay (0.0-1.0s)

Global Note Count Caps

The CRAWLER_MAX_NOTES_COUNT constant in config/base_config.py serves as a hard ceiling for the total number of items to retrieve. The pagination logic ensures that if the per-page limit exceeds the global cap, the cap is adjusted accordingly. For example, in the Zhihu implementation, the code explicitly checks if config.CRAWLER_MAX_NOTES_COUNT < zhihu_limit_count: to prevent requesting more data than permitted by the user’s configuration.

This cumulative limit works in tandem with the per-page constraints to create a bounded request volume, reducing the absolute number of API calls regardless of frequency.

Summary

  • Fixed pagination limits (10–20 items per page) defined per platform constrain batch sizes at the source.
  • Global sleep intervals (CRAWLER_MAX_SLEEP_SEC) enforce mandatory delays after every request.
  • Random jitter (random.random()) on platforms like Xiaohongshu breaks predictable request cadences.
  • Cumulative caps (CRAWLER_MAX_NOTES_COUNT) provide an absolute ceiling on total data volume.
  • All configuration is centralized in config/base_config.py while platform-specific logic resides in media_platform/{platform}/core.py.

Frequently Asked Questions

Where is the rate limiting configuration defined in MediaCrawler?

The primary configuration is located in config/base_config.py, which defines CRAWLER_MAX_SLEEP_SEC for sleep intervals and CRAWLER_MAX_NOTES_COUNT for total item limits. Platform-specific pagination limits are hardcoded as constants (e.g., zhihu_limit_count) in their respective media_platform/{platform}/core.py files.

How does MediaCrawler avoid being blocked by Weibo?

The Weibo crawler implements a weibo_limit_count = 10 to restrict page sizes and explicitly calls await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) after each request at line 452. The inline comment confirms this is designed specifically to avoid rate limiting by inserting a configurable delay between API calls.

What is the purpose of the random jitter in Xiaohongshu crawling?

The await asyncio.sleep(random.random()) call at line 493 in media_platform/xhs/core.py introduces a random sub-second delay. This jitter prevents the formation of rhythmic request patterns that automated detection systems often flag as bot behavior, thereby reducing the risk of IP throttling.

Can I adjust the crawl speed without modifying source code?

Yes. By changing CRAWLER_MAX_SLEEP_SEC and CRAWLER_MAX_NOTES_COUNT in config/base_config.py, you can increase or decrease the delay between requests and the total volume of data fetched. These are global settings that affect all platform crawlers without requiring changes to individual platform implementation files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →