Implementing Request Rate Limiting in MediaCrawler: A Complete Configuration Guide

MediaCrawler implements request rate limiting through a centralized configuration system using CRAWLER_MAX_SLEEP_SEC for mandatory request delays and MAX_CONCURRENCY_NUM for concurrency restrictions, preventing IP throttling across all supported platforms.

Implementing request rate limiting in MediaCrawler is essential for maintaining stable, long-running crawling sessions without triggering anti-bot protections on target platforms. The project employs a configurable, sleep-based throttling mechanism combined with concurrency controls to manage request velocity across Weibo, Zhihu, Douyin, and other platforms. All rate-limiting parameters are centralized in the configuration module, allowing developers to tune crawler aggressiveness without modifying platform-specific implementation code.

Centralized Configuration for Sleep Intervals

The foundation of MediaCrawler's rate limiting strategy resides in config/base_config.py, where global throttling constants define the crawler's behavior across all supported media platforms.

The CRAWLER_MAX_SLEEP_SEC Constant

The primary rate-limiting mechanism relies on CRAWLER_MAX_SLEEP_SEC, a configuration constant that defines the number of seconds each coroutine pauses after network requests or page navigations. According to the source code in config/base_config.py, this value defaults to approximately 2 seconds, though developers should adjust it based on target platform documentation and tolerance thresholds.

This centralized approach ensures that sleep duration remains consistent across Weibo, Zhihu, Tieba, Douyin, Bilibili, and Kuaishou implementations without requiring platform-specific hardcoding.

Request Quota Management with CRAWLER_MAX_NOTES_COUNT

MediaCrawler also implements CRAWLER_MAX_NOTES_COUNT to prevent excessive data extraction that might trigger rate limits. Platform implementations enforce minimum thresholds based on pagination limits. For example, in media_platform/weibo/core.py, the crawler establishes a weibo_limit_count = 10 (the platform's per-page default) and automatically updates the global configuration if the user-provided maximum is lower than this baseline.

Platform-Specific Implementation Patterns

Each platform crawler inserts explicit await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) calls at strategic points in the execution flow to ensure compliance with rate limits.

Weibo Implementation

In media_platform/weibo/core.py, the rate limiting appears in multiple critical paths:

  • Search pagination: Lines 86-88 insert await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) immediately after each page fetch operation
  • Detail fetching: Lines 52-54 apply the same sleep delay after retrieving a note's full text content

This pattern ensures that both high-volume search operations and individual content retrieval respect the configured throttling intervals.

Other Platform Implementations

The same architectural pattern extends across all supported platforms:

Each implementation references the same centralized configuration, maintaining consistency while preventing platform-specific rate limit violations.

Concurrency Control with Semaphores

Beyond sleep-based throttling, MediaCrawler limits burst traffic through MAX_CONCURRENCY_NUM, which restricts simultaneous asynchronous operations. The crawler creates semaphores using asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) before launching parallel fetches.

In methods like get_specified_notes, the semaphore ensures that only the configured number of requests execute concurrently, further reducing the risk of being flagged for abusive traffic patterns. This concurrency limitation works in tandem with sleep intervals to create a two-layered defense against rate limiting.

Practical Configuration Examples

To implement conservative rate limiting for production crawling, override the default configuration before initializing the crawler:

import config
from media_platform.weibo.core import WeiboCrawler
import asyncio

# Configure a 3-second pause between requests for safer crawling

config.CRAWLER_MAX_SLEEP_SEC = 3

# Reduce concurrent operations to 2 simultaneous tasks

config.MAX_CONCURRENCY_NUM = 2

# Ensure the crawler respects Weibo's pagination limits

config.CRAWLER_MAX_NOTES_COUNT = max(config.CRAWLER_MAX_NOTES_COUNT, 10)

# Initialize and run the crawler

crawler = WeiboCrawler()
await crawler.start()

When extending the crawler or adding new platforms, implement the standard sleep pattern:

async def fetch_page(self, page_num: int):
    # Execute the HTTP request or Playwright navigation

    response = await self.client.get_page(page_num)
    
    # Mandatory rate-limiting pause

    await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
    utils.logger.info(f"Paused {config.CRAWLER_MAX_SLEEP_SEC}s after fetching page {page_num}")
    
    return response

Summary

  • Centralized configuration: All rate-limiting parameters (CRAWLER_MAX_SLEEP_SEC, MAX_CONCURRENCY_NUM, CRAWLER_MAX_NOTES_COUNT) reside in config/base_config.py, enabling global adjustments without code changes
  • Explicit sleep implementation: Every platform crawler inserts await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) after network operations, as seen in media_platform/weibo/core.py (lines 52-54 and 86-88) and equivalent files for Zhihu, Douyin, Bilibili, and Tieba
  • Dual-layer protection: The combination of sleep intervals and concurrency semaphores prevents both rapid sequential requests and burst traffic patterns
  • Adaptive limits: Platform-specific minimums (like weibo_limit_count = 10) ensure configuration values respect pagination constraints

Frequently Asked Questions

Where are the rate limiting settings configured in MediaCrawler?

Rate limiting settings are centralized in config/base_config.py. This file contains CRAWLER_MAX_SLEEP_SEC (controlling delay between requests), MAX_CONCURRENCY_NUM (limiting simultaneous connections), and CRAWLER_MAX_NOTES_COUNT (restricting total items per run). Modifying these values affects all platform crawlers without requiring changes to individual implementation files.

How does MediaCrawler prevent IP blocking during high-volume crawling?

The framework implements a two-layered defense: first, CRAWLER_MAX_SLEEP_SEC forces explicit delays via asyncio.sleep() after every network request in platform core files like media_platform/weibo/core.py and media_platform/zhihu/core.py. Second, MAX_CONCURRENCY_NUM creates semaphores that cap simultaneous requests, preventing burst traffic patterns that trigger rate limits.

Can I adjust crawling speed without modifying source code?

Yes. Since CRAWLER_MAX_SLEEP_SEC and MAX_CONCURRENCY_NUM are imported from the config module, you can override these values at runtime before initializing your crawler instance. Alternatively, you can expose these settings through environment variables or command-line arguments that modify the config module attributes, allowing dynamic speed adjustment per execution.

What happens if CRAWLER_MAX_NOTES_COUNT is set below a platform's page size?

Platform implementations automatically adjust the configuration to respect pagination constraints. For example, in media_platform/weibo/core.py, the code sets weibo_limit_count = 10 and updates the global configuration if the user-provided maximum is lower, ensuring the crawler doesn't attempt invalid API calls while still respecting the rate limiting boundaries established by CRAWLER_MAX_SLEEP_SEC.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →