# How to Configure Rate Limiting and Crawl Intervals in MediaCrawler to Avoid Platform Bans

> Learn to configure rate limiting and crawl intervals in MediaCrawler to prevent platform bans. Set global or custom intervals to manage request pacing effectively and ensure smooth crawling.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: best-practices
- Published: 2026-08-14

---

**Set `config.CRAWLER_MAX_SLEEP_SEC` in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) for a global crawl interval, or pass a custom `crawl_interval` parameter to any client method to fine-tune request pacing per platform.**

MediaCrawler, the open-source Python framework for scraping Chinese social media platforms (Xiaohongshu, Zhihu, Weibo, Douyin, Bilibili, KuaiShou, and Tieba), implements built-in throttling mechanisms to help users avoid rate limits and account bans. Understanding how to configure these controls—both globally and per-call—is essential for reliable, long-running data collection.

---

## How MediaCrawler Handles Rate Limiting

The framework uses a **two-layer throttling strategy**: a configurable global default that applies to all platforms, plus optional per-method overrides for granular control.

### Global Crawl Interval via `CRAWLER_MAX_SLEEP_SEC`

Every platform crawler imports the shared configuration module and uses `config.CRAWLER_MAX_SLEEP_SEC` as its baseline delay between HTTP requests. This constant is defined in [`main/config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/config/base_config.py) and propagated across all core implementations:

- **Zhihu**: `crawl_interval=config.CRAWLER_MAX_SLEEP_SEC`
- **Xiaohongshu (XHS)**: `crawl_interval=config.CRAWLER_MAX_SLEEP_SEC`
- **Weibo**: `crawl_interval=config.CRAWLER_MAX_SLEEP_SEC`
- **Tieba**: `crawl_interval=config.CRAWLER_MAX_SLEEP_SEC`
- **KuaiShou**: `crawl_interval=config.CRAWLER_MAX_SLEEP_SEC`
- **Douyin**: `crawl_interval=config.CRAWLER_MAX_SLEEP_SEC`
- **Bilibili**: `crawl_interval=config.CRAWLER_MAX_SLEEP_SEC`

The core implementations apply this value as a fixed `asyncio.sleep()` pause after each request, ensuring consistent spacing regardless of response time.

### Per-Call Override with `crawl_interval` Parameter

Every public client method that performs paginated fetching exposes a `crawl_interval: float = 1.0` parameter. This allows runtime adjustment without modifying configuration files.

For example, in `ZhihuClient.get_all_notes_by_creator_url()`:

```python
async def get_all_notes_by_creator_url(
    self,
    url_token: str,
    crawl_interval: float = 1.0,
    callback: Optional[Callable] = None
) -> List[Dict]:
    """
    Fetch all notes from a Zhihu creator with configurable throttling.
    
    Args:
        crawl_interval: Seconds to sleep between requests.
                         Defaults to 1.0; override to increase safety margin.
    """
    return await self._fetch_paginated(
        endpoint=f"/creators/{url_token}/articles",
        crawl_interval=crawl_interval,
        callback=callback
    )

```

The value flows through to the underlying core routine:

```python

# In zhihu/core.py - the actual sleep implementation

await asyncio.sleep(crawl_interval)

```

---

## Platform-Specific Throttling Behaviors

### KuaiShou: Random Jitter for Pattern Obfuscation

The KuaiShou client adds stochastic variation to avoid predictable request timing. In [`media_platform/kuaishou/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/client.py):

```python
import random

async def get_notes_by_creator(
    self,
    creator_id: str,
    crawl_interval: float = 1.0
) -> List[Dict]:
    notes = []
    for page in self._paginate(creator_id):
        notes.extend(page)
        # Random jitter: 1-3 second additional delay

        jitter = random.uniform(1, 3)
        await asyncio.sleep(crawl_interval + jitter)
    return notes

```

This makes traffic signatures harder to fingerprint. **Keep jitter enabled for production deployments.**

### Rate-Limit Detection and Retry Backoff

MediaCrawler's HTTP utilities raise `RateLimitError` (defined in [`media_platform/xhs/exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/exception.py) and similar modules) when platforms signal throttling. The `async_retry` decorator implements exponential backoff:

| Attempt | Delay |
|---------|-------|
| 1st retry | 1 second |
| 2nd retry | 2 seconds |
| 3rd retry | 4 seconds |
| 4th+ retry | 8 seconds (cap) |

The crawler logs each retry with platform-specific context:

```

[2024-01-15 09:23:41] WARNING  [XiaoHongShuCrawler] Rate limit hit (429), backing off 4.0s
[2024-01-15 09:23:45] INFO     [XiaoHongShuCrawler] Resuming crawl with interval=2.5

```

---

## Configuration Guide: Avoiding Platform Bans

### Step 1: Tune the Global Default

Edit [`main/config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/config/base_config.py):

```python

# main/config/base_config.py

# Base crawl interval in seconds (float)

# Increase for aggressive platforms (Weibo, Zhihu): 2.0 - 5.0

# Decrease for tolerant platforms only with proxy rotation: 0.5 - 1.0

CRAWLER_MAX_SLEEP_SEC = 2.5

```

Recommended starting values by platform:

| Platform | Recommended `CRAWLER_MAX_SLEEP_SEC` | Notes |
|----------|-------------------------------------|-------|
| Zhihu | 2.5 - 4.0 | Strict rate limiting; frequent 429s below 2s |
| Xiaohongshu | 1.5 - 3.0 | Moderate limits; CSRF protection active |
| Weibo | 3.0 - 5.0 | Very aggressive blocking; prioritize accounts with age |
| Tieba | 2.0 - 3.0 | IP-based limits; combine with proxy pools |
| KuaiShou | 1.5 - 2.5 | Jitter adds 1-3s automatically |
| Douyin | 2.0 - 4.0 | Signature-based detection; interval helps evasion |
| Bilibili | 1.0 - 2.0 | Most tolerant; can reduce if using rotating IPs |

### Step 2: Apply Per-Call Overrides for Sensitive Operations

When fetching high-value targets or running concurrent jobs, increase spacing explicitly:

```python
from media_platform.zhihu.client import ZhihuClient
from media_platform.xhs.client import XHSClient
import asyncio

async def safe_batch_crawl():
    """Crawl multiple creators with conservative delays."""
    zhihu = ZhihuClient()
    xhs = XHSClient()
    
    # High-priority Zhihu creator: extra cautious

    zhihu_data = await zhihu.get_all_notes_by_creator_url(
        url_token="zhangsan-123",
        crawl_interval=4.0,  # 4 seconds between requests

        callback=process_note
    )
    
    # Xiaohongshu with moderate pace

    xhs_data = await xhs.get_notes_by_keyword(
        keyword="skincare",
        crawl_interval=2.5,
        max_notes=500
    )
    
    return zhihu_data, xhs_data

asyncio.run(safe_batch_crawl())

```

### Step 3: Calculate Intervals from Platform Limits

Some platform APIs document explicit rate limits. Convert requests-per-minute to crawl intervals:

```python
def interval_from_rpm(max_requests_per_minute: int, safety_factor: float = 0.8) -> float:
    """
    Calculate safe crawl interval from documented rate limit.
    
    safety_factor: Use 80% of limit to leave headroom.
    """
    seconds_per_request = 60.0 / (max_requests_per_minute * safety_factor)
    return round(seconds_per_request, 2)

# Example: Platform allows 30 requests/minute

# Safe interval: 60 / (30 * 0.8) = 2.5 seconds

crawl_interval = interval_from_rpm(30)  # → 2.5

```

### Step 4: Monitor and Iterate

Enable verbose logging to observe actual sleep durations and rate-limit events:

```python

# In your entry script or config

import logging
logging.getLogger("MediaCrawler").setLevel(logging.DEBUG)

```

Watch for these patterns in logs:

- `"Sleeping for X.XXs"` — normal operation, verify interval matches config
- `"Rate limit hit (429)"` — increase interval or add proxy rotation
- `"Sleeping for X.XXs + jitter"` — KuaiShou jitter active, monitor total delay

---

## Complete Configuration Example

**File: [`main/config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/config/base_config.py)**

```python
"""
MediaCrawler global configuration.
Adjust these values before starting production crawls.
"""

import os

# Core throttling parameter

# Default: 1.0s, Recommended production: 2.0-5.0s depending on platform

CRAWLER_MAX_SLEEP_SEC = float(os.getenv("CRAWLER_SLEEP", "2.5"))

# Retry behavior

MAX_RETRIES = 3                    # Exponential backoff attempts

RETRY_BACKOFF_BASE = 1.0           # First retry delay in seconds

# Proxy configuration (recommended for high-volume crawling)

ENABLE_PROXY = False
PROXY_POOL_URL = "http://proxy-provider:5010/get/"

# Platform-specific overrides (optional)

PLATFORM_OVERRIDES = {
    "weibo": {"crawl_interval": 4.0},
    "zhihu": {"crawl_interval": 3.0},
    "xhs": {"crawl_interval": 2.0},
}

```

**File: [`production_crawl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/production_crawl.py)**

```python
#!/usr/bin/env python3
"""
Production crawl script with environment-aware throttling.
"""
import os
import asyncio
from media_platform.zhihu.client import ZhihuClient
from media_platform.xhs.client import XHSClient
from main.config import base_config


async def main():
    # Environment override takes precedence

    interval = float(os.getenv("CRAWLER_INTERVAL", base_config.CRAWLER_MAX_SLEEP_SEC))
    
    print(f"Starting crawl with interval={interval}s")
    
    zhihu = ZhihuClient()
    xhs = XHSClient()
    
    # Parallel execution with same interval

    results = await asyncio.gather(
        zhihu.get_all_notes_by_creator_url(
            url_token="example-creator",
            crawl_interval=interval
        ),
        xhs.get_notes_by_keyword(
            keyword="product-review",
            crawl_interval=interval,
            max_notes=1000
        ),
        return_exceptions=True
    )
    
    # Handle any rate-limit exceptions

    for platform, result in zip(["zhihu", "xhs"], results):
        if isinstance(result, Exception):
            print(f"{platform} failed: {result}")
        else:
            print(f"{platform}: collected {len(result)} items")

if __name__ == "__main__":
    asyncio.run(main())

```

**Run with elevated safety:**

```bash
export CRAWLER_INTERVAL=3.5
export CRAWLER_SLEEP=3.5  # Fallback if interval not set

python production_crawl.py

```

---

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`main/config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/config/base_config.py) | Global configuration including `CRAWLER_MAX_SLEEP_SEC` |
| [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) | Zhihu crawling logic with interval enforcement |
| [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) | Xiaohongshu crawling implementation |
| [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py) | Weibo crawling with rate-limit handling |
| [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py) | Tieba crawling implementation |
| [`media_platform/kuaishou/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/core.py) | KuaiShou crawling with jitter |
| [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) | Douyin crawling implementation |
| [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) | Bilibili crawling implementation |
| [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) | Zhihu client with `crawl_interval` parameter |
| [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) | XHS client with per-call interval override |
| [`media_platform/kuaishou/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/client.py) | KuaiShou client with random jitter |
| [`media_platform/xhs/exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/exception.py) | `RateLimitError` definition |

---

## Summary

- **Global control**: Modify `CRAWLER_MAX_SLEEP_SEC` in [`main/config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/config/base_config.py) to set a project-wide default crawl interval.

- **Per-call precision**: Pass `crawl_interval` to any client method (`get_all_notes_by_creator_url`, `get_notes_by_keyword`, etc.) for targeted adjustments.

- **Platform awareness**: Weibo and Zhihu require conservative settings (3-5s); Bilibili tolerates more aggressive pacing (1-2s).

- **Obfuscation**: KuaiShou automatically adds 1-3s random jitter—maintain this behavior in production.

- **Resilience**: Exponential backoff on `RateLimitError` provides automatic recovery; monitor logs to iteratively optimize intervals.

---

## Frequently Asked Questions

### How do I find the optimal crawl interval for a specific platform?

Start with the documented rate limit or recommended values in this guide, then monitor logs for 429 responses. If you see rate-limit errors within the first 100 requests, increase the interval by 0.5-1.0 seconds and retry. For undocumented platforms, begin at 3.0 seconds and decrease gradually while observing ban patterns.

### Can I disable rate limiting entirely for faster testing?

Setting `crawl_interval=0` or `CRAWLER_MAX_SLEEP_SEC=0` removes artificial delays, but this will trigger immediate platform blocks. For local testing with cached responses, use `crawl_interval=0.1` minimum. Never deploy zero-delay configurations against live platforms.

### Does MediaCrawler support dynamic interval adjustment based on response headers?

Currently, the framework does not parse `Retry-After` headers or adaptive rate-limit responses. The `crawl_interval` and `CRAWLER_MAX_SLEEP_SEC` values are static once set. Implement custom logic in your callback functions to read response metadata and adjust subsequent calls if needed.

### What's the difference between `crawl_interval` and the retry backoff delay?

`crawl_interval` controls **intentional spacing between successful requests**—your primary anti-ban mechanism. The retry backoff activates only **after rate-limit errors**, using exponential delays (1s, 2s, 4s, 8s) to recover from temporary blocks without manual intervention.