How to Handle Rate Limiting in MediaCrawler: Configuration and Implementation Guide

MediaCrawler prevents throttling and IP bans through a centralized sleep-based throttling system driven by the CRAWLER_MAX_SLEEP_SEC configuration constant and explicit asyncio.sleep calls after every network request.

MediaCrawler is an open-source asynchronous crawler for Chinese social media platforms that implements a robust, configuration-driven approach to handle rate limiting. By centralizing throttle controls in config/base_config.py and consistently applying sleep intervals across all platform implementations, the tool enables developers to crawl Weibo, Zhihu, Douyin, and other sites without triggering anti-bot protections.

Configure the Global Sleep Interval

The foundation of MediaCrawler's rate limiting strategy resides in the CRAWLER_MAX_SLEEP_SEC constant defined in config/base_config.py. This single configuration value controls the number of seconds each coroutine pauses after performing network requests or page navigations. By adjusting this constant, developers tune the crawler's aggressiveness without modifying platform-specific code.

The configuration also defines CRAWLER_MAX_NOTES_COUNT, which sets a global cap on the total number of items to retrieve per run. Platform implementations reference this limit to prevent exhausting request quotas in a single session.

Implement Platform-Specific Throttling

Every supported platform inserts explicit await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) calls at critical execution points to ensure consistent throttling across the codebase.

In media_platform/weibo/core.py, the crawler invokes this sleep pattern in two key locations:

  • After each page fetch in the search method (lines 86-88)
  • After retrieving a note's full text (lines 52-54)

Similarly, media_platform/zhihu/core.py applies identical sleep logic after page requests and comment pagination operations. The pattern extends to media_platform/tieba/core.py, media_platform/douyin/core.py, media_platform/bilibili/core.py, and other platform modules, ensuring uniform rate limiting regardless of the target site's specific API constraints.

Some implementations also implement adaptive limits. For example, the Weibo crawler defines weibo_limit_count = 10 to establish a minimum threshold based on the platform's page size, automatically updating the global configuration if the user-provided maximum is lower than this value.

Control Concurrency to Reduce Burst Traffic

MediaCrawler complements sleep-based throttling with concurrency controls defined by MAX_CONCURRENCY_NUM in the configuration module. This setting limits the number of simultaneous tasks executing at any given time, preventing burst traffic patterns that platforms might interpret as abusive behavior.

The crawler creates semaphores using asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) before launching parallel fetch operations, commonly utilized in methods like get_specified_notes. This semaphore-based approach ensures that even when crawling multiple pages or items concurrently, the total number of in-flight requests never exceeds the configured threshold.

Practical Implementation Steps

To effectively handle rate limiting in your MediaCrawler deployment, follow these configuration steps:

  1. Set a safe sleep interval – Adjust config.CRAWLER_MAX_SLEEP_SEC (default approximately 2 seconds) to comply with your target platform's documented rate limits. Increase this value for more aggressive anti-bot protections.

  2. Respect per-page limits – Ensure config.CRAWLER_MAX_NOTES_COUNT meets or exceeds the platform's native page size (e.g., weibo_limit_count = 10 for Weibo) to prevent configuration conflicts.

  3. Limit concurrency – Set config.MAX_CONCURRENCY_NUM to bound simultaneous coroutines, typically between 1-5 for conservative crawling or higher for permissive targets.


# Override configuration before launching the crawler

import config
from media_platform.weibo.core import WeiboCrawler

# Conservative 3-second pause between requests

config.CRAWLER_MAX_SLEEP_SEC = 3

# Reduce concurrent tasks to minimize burst traffic

config.MAX_CONCURRENCY_NUM = 2

# Initialize and run

crawler = WeiboCrawler()
await crawler.start()

Platform implementations consistently apply the sleep pattern after network operations:


# Illustrative snippet from platform core implementations

async def fetch_page(self, page_num: int):
    # Execute HTTP request or Playwright navigation

    await self.client.get_page(page_num)
    
    # Mandatory rate-limiting pause

    await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
    utils.logger.info(f"Paused {config.CRAWLER_MAX_SLEEP_SEC}s after page {page_num}")

Summary

  • Centralized configuration – Rate limiting is controlled through CRAWLER_MAX_SLEEP_SEC and MAX_CONCURRENCY_NUM in config/base_config.py, enabling global adjustments without code changes.
  • Explicit sleep implementation – Every platform module (Weibo, Zhihu, Douyin, Bilibili, Tieba) inserts await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) after network requests to prevent throttling.
  • Concurrency management – asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) bounds simultaneous tasks to eliminate burst traffic patterns.
  • Adaptive limits – The crawler adjusts CRAWLER_MAX_NOTES_COUNT based on platform-specific page sizes (e.g., Weibo's 10-item limit) to prevent quota exhaustion.

Frequently Asked Questions

What is the default sleep interval in MediaCrawler?

The default CRAWLER_MAX_SLEEP_SEC value is approximately 2 seconds, though this may vary by platform implementation. You can verify and modify this constant in config/base_config.py to match your target site's specific rate limiting requirements.

How does MediaCrawler handle different rate limits across platforms?

MediaCrawler uses a unified sleep mechanism (CRAWLER_MAX_SLEEP_SEC) across all platforms, but individual implementations in media_platform/*/core.py files can adjust behavior through adaptive limits like weibo_limit_count. For strict platform-specific compliance, modify the global sleep interval before initializing the specific crawler class.

Can I disable rate limiting in MediaCrawler?

While technically possible by setting CRAWLER_MAX_SLEEP_SEC to 0, this is strongly discouraged as it will likely trigger immediate IP bans or CAPTCHA challenges from target platforms. The configuration is designed to be tunable, not bypassable, to maintain sustainable long-term crawling operations.

Where is the concurrency limit configured?

The concurrency limit is defined by MAX_CONCURRENCY_NUM in config/base_config.py and enforced through asyncio.Semaphore() instances in crawler methods such as get_specified_notes. Reducing this value from its default decreases simultaneous connections and further reduces the risk of rate limiting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →