# Implementing Request Rate Limiting in MediaCrawler: A Complete Configuration Guide

> Configure request rate limiting in MediaCrawler using CRAWLER_MAX_SLEEP_SEC and MAX_CONCURRENCY_NUM. Prevent IP throttling and optimize crawling performance with this essential guide.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**MediaCrawler implements request rate limiting through a centralized configuration system using `CRAWLER_MAX_SLEEP_SEC` for mandatory request delays and `MAX_CONCURRENCY_NUM` for concurrency restrictions, preventing IP throttling across all supported platforms.**

Implementing request rate limiting in MediaCrawler is essential for maintaining stable, long-running crawling sessions without triggering anti-bot protections on target platforms. The project employs a configurable, sleep-based throttling mechanism combined with concurrency controls to manage request velocity across Weibo, Zhihu, Douyin, and other platforms. All rate-limiting parameters are centralized in the configuration module, allowing developers to tune crawler aggressiveness without modifying platform-specific implementation code.

## Centralized Configuration for Sleep Intervals

The foundation of MediaCrawler's rate limiting strategy resides in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), where global throttling constants define the crawler's behavior across all supported media platforms.

### The CRAWLER_MAX_SLEEP_SEC Constant

The primary rate-limiting mechanism relies on **`CRAWLER_MAX_SLEEP_SEC`**, a configuration constant that defines the number of seconds each coroutine pauses after network requests or page navigations. According to the source code in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), this value defaults to approximately 2 seconds, though developers should adjust it based on target platform documentation and tolerance thresholds.

This centralized approach ensures that sleep duration remains consistent across Weibo, Zhihu, Tieba, Douyin, Bilibili, and Kuaishou implementations without requiring platform-specific hardcoding.

### Request Quota Management with CRAWLER_MAX_NOTES_COUNT

MediaCrawler also implements **`CRAWLER_MAX_NOTES_COUNT`** to prevent excessive data extraction that might trigger rate limits. Platform implementations enforce minimum thresholds based on pagination limits. For example, in [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py), the crawler establishes a `weibo_limit_count = 10` (the platform's per-page default) and automatically updates the global configuration if the user-provided maximum is lower than this baseline.

## Platform-Specific Implementation Patterns

Each platform crawler inserts explicit `await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` calls at strategic points in the execution flow to ensure compliance with rate limits.

### Weibo Implementation

In [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py), the rate limiting appears in multiple critical paths:

- **Search pagination**: Lines 86-88 insert `await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` immediately after each page fetch operation
- **Detail fetching**: Lines 52-54 apply the same sleep delay after retrieving a note's full text content

This pattern ensures that both high-volume search operations and individual content retrieval respect the configured throttling intervals.

### Other Platform Implementations

The same architectural pattern extends across all supported platforms:

- **Zhihu**: [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) applies sleep delays after each page request and comment pagination operation
- **Tieba**: [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py) follows identical throttling logic for forum crawling
- **Douyin**: [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) inserts sleep calls after pagination and comment fetching sequences
- **Bilibili**: [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) manages rate limiting during search and video detail retrieval

Each implementation references the same centralized configuration, maintaining consistency while preventing platform-specific rate limit violations.

## Concurrency Control with Semaphores

Beyond sleep-based throttling, MediaCrawler limits burst traffic through **`MAX_CONCURRENCY_NUM`**, which restricts simultaneous asynchronous operations. The crawler creates semaphores using `asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)` before launching parallel fetches.

In methods like `get_specified_notes`, the semaphore ensures that only the configured number of requests execute concurrently, further reducing the risk of being flagged for abusive traffic patterns. This concurrency limitation works in tandem with sleep intervals to create a two-layered defense against rate limiting.

## Practical Configuration Examples

To implement conservative rate limiting for production crawling, override the default configuration before initializing the crawler:

```python
import config
from media_platform.weibo.core import WeiboCrawler
import asyncio

# Configure a 3-second pause between requests for safer crawling

config.CRAWLER_MAX_SLEEP_SEC = 3

# Reduce concurrent operations to 2 simultaneous tasks

config.MAX_CONCURRENCY_NUM = 2

# Ensure the crawler respects Weibo's pagination limits

config.CRAWLER_MAX_NOTES_COUNT = max(config.CRAWLER_MAX_NOTES_COUNT, 10)

# Initialize and run the crawler

crawler = WeiboCrawler()
await crawler.start()

```

When extending the crawler or adding new platforms, implement the standard sleep pattern:

```python
async def fetch_page(self, page_num: int):
    # Execute the HTTP request or Playwright navigation

    response = await self.client.get_page(page_num)
    
    # Mandatory rate-limiting pause

    await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
    utils.logger.info(f"Paused {config.CRAWLER_MAX_SLEEP_SEC}s after fetching page {page_num}")
    
    return response

```

## Summary

- **Centralized configuration**: All rate-limiting parameters (`CRAWLER_MAX_SLEEP_SEC`, `MAX_CONCURRENCY_NUM`, `CRAWLER_MAX_NOTES_COUNT`) reside in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), enabling global adjustments without code changes
- **Explicit sleep implementation**: Every platform crawler inserts `await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` after network operations, as seen in [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py) (lines 52-54 and 86-88) and equivalent files for Zhihu, Douyin, Bilibili, and Tieba
- **Dual-layer protection**: The combination of sleep intervals and concurrency semaphores prevents both rapid sequential requests and burst traffic patterns
- **Adaptive limits**: Platform-specific minimums (like `weibo_limit_count = 10`) ensure configuration values respect pagination constraints

## Frequently Asked Questions

### Where are the rate limiting settings configured in MediaCrawler?

Rate limiting settings are centralized in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). This file contains **`CRAWLER_MAX_SLEEP_SEC`** (controlling delay between requests), **`MAX_CONCURRENCY_NUM`** (limiting simultaneous connections), and **`CRAWLER_MAX_NOTES_COUNT`** (restricting total items per run). Modifying these values affects all platform crawlers without requiring changes to individual implementation files.

### How does MediaCrawler prevent IP blocking during high-volume crawling?

The framework implements a two-layered defense: first, **`CRAWLER_MAX_SLEEP_SEC`** forces explicit delays via `asyncio.sleep()` after every network request in platform core files like [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py) and [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py). Second, **`MAX_CONCURRENCY_NUM`** creates semaphores that cap simultaneous requests, preventing burst traffic patterns that trigger rate limits.

### Can I adjust crawling speed without modifying source code?

Yes. Since `CRAWLER_MAX_SLEEP_SEC` and `MAX_CONCURRENCY_NUM` are imported from the config module, you can override these values at runtime before initializing your crawler instance. Alternatively, you can expose these settings through environment variables or command-line arguments that modify the config module attributes, allowing dynamic speed adjustment per execution.

### What happens if CRAWLER_MAX_NOTES_COUNT is set below a platform's page size?

Platform implementations automatically adjust the configuration to respect pagination constraints. For example, in [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py), the code sets `weibo_limit_count = 10` and updates the global configuration if the user-provided maximum is lower, ensuring the crawler doesn't attempt invalid API calls while still respecting the rate limiting boundaries established by `CRAWLER_MAX_SLEEP_SEC`.