# How to Handle Rate Limiting in MediaCrawler: Configuration and Implementation Guide

> Learn to handle rate limiting in MediaCrawler. Discover how CRAWLER_MAX_SLEEP_SEC and asyncio.sleep prevent throttling and IP bans for smooth data crawling.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**MediaCrawler prevents throttling and IP bans through a centralized sleep-based throttling system driven by the `CRAWLER_MAX_SLEEP_SEC` configuration constant and explicit `asyncio.sleep` calls after every network request.**

MediaCrawler is an open-source asynchronous crawler for Chinese social media platforms that implements a robust, configuration-driven approach to handle rate limiting. By centralizing throttle controls in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and consistently applying sleep intervals across all platform implementations, the tool enables developers to crawl Weibo, Zhihu, Douyin, and other sites without triggering anti-bot protections.

## Configure the Global Sleep Interval

The foundation of MediaCrawler's rate limiting strategy resides in the `CRAWLER_MAX_SLEEP_SEC` constant defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). This single configuration value controls the number of seconds each coroutine pauses after performing network requests or page navigations. By adjusting this constant, developers tune the crawler's aggressiveness without modifying platform-specific code.

The configuration also defines `CRAWLER_MAX_NOTES_COUNT`, which sets a global cap on the total number of items to retrieve per run. Platform implementations reference this limit to prevent exhausting request quotas in a single session.

## Implement Platform-Specific Throttling

Every supported platform inserts explicit `await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` calls at critical execution points to ensure consistent throttling across the codebase.

In [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py), the crawler invokes this sleep pattern in two key locations:
- After each page fetch in the `search` method (lines 86-88)
- After retrieving a note's full text (lines 52-54)

Similarly, [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) applies identical sleep logic after page requests and comment pagination operations. The pattern extends to [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py), [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py), [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py), and other platform modules, ensuring uniform rate limiting regardless of the target site's specific API constraints.

Some implementations also implement adaptive limits. For example, the Weibo crawler defines `weibo_limit_count = 10` to establish a minimum threshold based on the platform's page size, automatically updating the global configuration if the user-provided maximum is lower than this value.

## Control Concurrency to Reduce Burst Traffic

MediaCrawler complements sleep-based throttling with concurrency controls defined by `MAX_CONCURRENCY_NUM` in the configuration module. This setting limits the number of simultaneous tasks executing at any given time, preventing burst traffic patterns that platforms might interpret as abusive behavior.

The crawler creates semaphores using `asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)` before launching parallel fetch operations, commonly utilized in methods like `get_specified_notes`. This semaphore-based approach ensures that even when crawling multiple pages or items concurrently, the total number of in-flight requests never exceeds the configured threshold.

## Practical Implementation Steps

To effectively handle rate limiting in your MediaCrawler deployment, follow these configuration steps:

1. **Set a safe sleep interval** – Adjust `config.CRAWLER_MAX_SLEEP_SEC` (default approximately 2 seconds) to comply with your target platform's documented rate limits. Increase this value for more aggressive anti-bot protections.

2. **Respect per-page limits** – Ensure `config.CRAWLER_MAX_NOTES_COUNT` meets or exceeds the platform's native page size (e.g., `weibo_limit_count = 10` for Weibo) to prevent configuration conflicts.

3. **Limit concurrency** – Set `config.MAX_CONCURRENCY_NUM` to bound simultaneous coroutines, typically between 1-5 for conservative crawling or higher for permissive targets.

```python

# Override configuration before launching the crawler

import config
from media_platform.weibo.core import WeiboCrawler

# Conservative 3-second pause between requests

config.CRAWLER_MAX_SLEEP_SEC = 3

# Reduce concurrent tasks to minimize burst traffic

config.MAX_CONCURRENCY_NUM = 2

# Initialize and run

crawler = WeiboCrawler()
await crawler.start()

```

Platform implementations consistently apply the sleep pattern after network operations:

```python

# Illustrative snippet from platform core implementations

async def fetch_page(self, page_num: int):
    # Execute HTTP request or Playwright navigation

    await self.client.get_page(page_num)
    
    # Mandatory rate-limiting pause

    await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
    utils.logger.info(f"Paused {config.CRAWLER_MAX_SLEEP_SEC}s after page {page_num}")

```

## Summary

- **Centralized configuration** – Rate limiting is controlled through `CRAWLER_MAX_SLEEP_SEC` and `MAX_CONCURRENCY_NUM` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), enabling global adjustments without code changes.
- **Explicit sleep implementation** – Every platform module (Weibo, Zhihu, Douyin, Bilibili, Tieba) inserts `await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` after network requests to prevent throttling.
- **Concurrency management** – `asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)` bounds simultaneous tasks to eliminate burst traffic patterns.
- **Adaptive limits** – The crawler adjusts `CRAWLER_MAX_NOTES_COUNT` based on platform-specific page sizes (e.g., Weibo's 10-item limit) to prevent quota exhaustion.

## Frequently Asked Questions

### What is the default sleep interval in MediaCrawler?

The default `CRAWLER_MAX_SLEEP_SEC` value is approximately 2 seconds, though this may vary by platform implementation. You can verify and modify this constant in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to match your target site's specific rate limiting requirements.

### How does MediaCrawler handle different rate limits across platforms?

MediaCrawler uses a unified sleep mechanism (`CRAWLER_MAX_SLEEP_SEC`) across all platforms, but individual implementations in `media_platform/*/core.py` files can adjust behavior through adaptive limits like `weibo_limit_count`. For strict platform-specific compliance, modify the global sleep interval before initializing the specific crawler class.

### Can I disable rate limiting in MediaCrawler?

While technically possible by setting `CRAWLER_MAX_SLEEP_SEC` to 0, this is strongly discouraged as it will likely trigger immediate IP bans or CAPTCHA challenges from target platforms. The configuration is designed to be tunable, not bypassable, to maintain sustainable long-term crawling operations.

### Where is the concurrency limit configured?

The concurrency limit is defined by `MAX_CONCURRENCY_NUM` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and enforced through `asyncio.Semaphore()` instances in crawler methods such as `get_specified_notes`. Reducing this value from its default decreases simultaneous connections and further reduces the risk of rate limiting.