How to Configure Rate Limiting and Crawl Intervals in MediaCrawler to Avoid Platform Bans

Set config.CRAWLER_MAX_SLEEP_SEC in base_config.py for a global crawl interval, or pass a custom crawl_interval parameter to any client method to fine-tune request pacing per platform.

MediaCrawler, the open-source Python framework for scraping Chinese social media platforms (Xiaohongshu, Zhihu, Weibo, Douyin, Bilibili, KuaiShou, and Tieba), implements built-in throttling mechanisms to help users avoid rate limits and account bans. Understanding how to configure these controls—both globally and per-call—is essential for reliable, long-running data collection.


How MediaCrawler Handles Rate Limiting

The framework uses a two-layer throttling strategy: a configurable global default that applies to all platforms, plus optional per-method overrides for granular control.

Global Crawl Interval via CRAWLER_MAX_SLEEP_SEC

Every platform crawler imports the shared configuration module and uses config.CRAWLER_MAX_SLEEP_SEC as its baseline delay between HTTP requests. This constant is defined in main/config/base_config.py and propagated across all core implementations:

  • Zhihu: crawl_interval=config.CRAWLER_MAX_SLEEP_SEC
  • Xiaohongshu (XHS): crawl_interval=config.CRAWLER_MAX_SLEEP_SEC
  • Weibo: crawl_interval=config.CRAWLER_MAX_SLEEP_SEC
  • Tieba: crawl_interval=config.CRAWLER_MAX_SLEEP_SEC
  • KuaiShou: crawl_interval=config.CRAWLER_MAX_SLEEP_SEC
  • Douyin: crawl_interval=config.CRAWLER_MAX_SLEEP_SEC
  • Bilibili: crawl_interval=config.CRAWLER_MAX_SLEEP_SEC

The core implementations apply this value as a fixed asyncio.sleep() pause after each request, ensuring consistent spacing regardless of response time.

Per-Call Override with crawl_interval Parameter

Every public client method that performs paginated fetching exposes a crawl_interval: float = 1.0 parameter. This allows runtime adjustment without modifying configuration files.

For example, in ZhihuClient.get_all_notes_by_creator_url():

async def get_all_notes_by_creator_url(
    self,
    url_token: str,
    crawl_interval: float = 1.0,
    callback: Optional[Callable] = None
) -> List[Dict]:
    """
    Fetch all notes from a Zhihu creator with configurable throttling.
    
    Args:
        crawl_interval: Seconds to sleep between requests.
                         Defaults to 1.0; override to increase safety margin.
    """
    return await self._fetch_paginated(
        endpoint=f"/creators/{url_token}/articles",
        crawl_interval=crawl_interval,
        callback=callback
    )

The value flows through to the underlying core routine:


# In zhihu/core.py - the actual sleep implementation

await asyncio.sleep(crawl_interval)

Platform-Specific Throttling Behaviors

KuaiShou: Random Jitter for Pattern Obfuscation

The KuaiShou client adds stochastic variation to avoid predictable request timing. In media_platform/kuaishou/client.py:

import random

async def get_notes_by_creator(
    self,
    creator_id: str,
    crawl_interval: float = 1.0
) -> List[Dict]:
    notes = []
    for page in self._paginate(creator_id):
        notes.extend(page)
        # Random jitter: 1-3 second additional delay

        jitter = random.uniform(1, 3)
        await asyncio.sleep(crawl_interval + jitter)
    return notes

This makes traffic signatures harder to fingerprint. Keep jitter enabled for production deployments.

Rate-Limit Detection and Retry Backoff

MediaCrawler's HTTP utilities raise RateLimitError (defined in media_platform/xhs/exception.py and similar modules) when platforms signal throttling. The async_retry decorator implements exponential backoff:

Attempt Delay
1st retry 1 second
2nd retry 2 seconds
3rd retry 4 seconds
4th+ retry 8 seconds (cap)

The crawler logs each retry with platform-specific context:


[2024-01-15 09:23:41] WARNING  [XiaoHongShuCrawler] Rate limit hit (429), backing off 4.0s
[2024-01-15 09:23:45] INFO     [XiaoHongShuCrawler] Resuming crawl with interval=2.5


Configuration Guide: Avoiding Platform Bans

Step 1: Tune the Global Default

Edit main/config/base_config.py:


# main/config/base_config.py

# Base crawl interval in seconds (float)

# Increase for aggressive platforms (Weibo, Zhihu): 2.0 - 5.0

# Decrease for tolerant platforms only with proxy rotation: 0.5 - 1.0

CRAWLER_MAX_SLEEP_SEC = 2.5

Recommended starting values by platform:

Platform Recommended CRAWLER_MAX_SLEEP_SEC Notes
Zhihu 2.5 - 4.0 Strict rate limiting; frequent 429s below 2s
Xiaohongshu 1.5 - 3.0 Moderate limits; CSRF protection active
Weibo 3.0 - 5.0 Very aggressive blocking; prioritize accounts with age
Tieba 2.0 - 3.0 IP-based limits; combine with proxy pools
KuaiShou 1.5 - 2.5 Jitter adds 1-3s automatically
Douyin 2.0 - 4.0 Signature-based detection; interval helps evasion
Bilibili 1.0 - 2.0 Most tolerant; can reduce if using rotating IPs

Step 2: Apply Per-Call Overrides for Sensitive Operations

When fetching high-value targets or running concurrent jobs, increase spacing explicitly:

from media_platform.zhihu.client import ZhihuClient
from media_platform.xhs.client import XHSClient
import asyncio

async def safe_batch_crawl():
    """Crawl multiple creators with conservative delays."""
    zhihu = ZhihuClient()
    xhs = XHSClient()
    
    # High-priority Zhihu creator: extra cautious

    zhihu_data = await zhihu.get_all_notes_by_creator_url(
        url_token="zhangsan-123",
        crawl_interval=4.0,  # 4 seconds between requests

        callback=process_note
    )
    
    # Xiaohongshu with moderate pace

    xhs_data = await xhs.get_notes_by_keyword(
        keyword="skincare",
        crawl_interval=2.5,
        max_notes=500
    )
    
    return zhihu_data, xhs_data

asyncio.run(safe_batch_crawl())

Step 3: Calculate Intervals from Platform Limits

Some platform APIs document explicit rate limits. Convert requests-per-minute to crawl intervals:

def interval_from_rpm(max_requests_per_minute: int, safety_factor: float = 0.8) -> float:
    """
    Calculate safe crawl interval from documented rate limit.
    
    safety_factor: Use 80% of limit to leave headroom.
    """
    seconds_per_request = 60.0 / (max_requests_per_minute * safety_factor)
    return round(seconds_per_request, 2)

# Example: Platform allows 30 requests/minute

# Safe interval: 60 / (30 * 0.8) = 2.5 seconds

crawl_interval = interval_from_rpm(30)  # → 2.5

Step 4: Monitor and Iterate

Enable verbose logging to observe actual sleep durations and rate-limit events:


# In your entry script or config

import logging
logging.getLogger("MediaCrawler").setLevel(logging.DEBUG)

Watch for these patterns in logs:

  • "Sleeping for X.XXs" — normal operation, verify interval matches config
  • "Rate limit hit (429)" — increase interval or add proxy rotation
  • "Sleeping for X.XXs + jitter" — KuaiShou jitter active, monitor total delay

Complete Configuration Example

File: main/config/base_config.py

"""
MediaCrawler global configuration.
Adjust these values before starting production crawls.
"""

import os

# Core throttling parameter

# Default: 1.0s, Recommended production: 2.0-5.0s depending on platform

CRAWLER_MAX_SLEEP_SEC = float(os.getenv("CRAWLER_SLEEP", "2.5"))

# Retry behavior

MAX_RETRIES = 3                    # Exponential backoff attempts

RETRY_BACKOFF_BASE = 1.0           # First retry delay in seconds

# Proxy configuration (recommended for high-volume crawling)

ENABLE_PROXY = False
PROXY_POOL_URL = "http://proxy-provider:5010/get/"

# Platform-specific overrides (optional)

PLATFORM_OVERRIDES = {
    "weibo": {"crawl_interval": 4.0},
    "zhihu": {"crawl_interval": 3.0},
    "xhs": {"crawl_interval": 2.0},
}

File: production_crawl.py

#!/usr/bin/env python3
"""
Production crawl script with environment-aware throttling.
"""
import os
import asyncio
from media_platform.zhihu.client import ZhihuClient
from media_platform.xhs.client import XHSClient
from main.config import base_config


async def main():
    # Environment override takes precedence

    interval = float(os.getenv("CRAWLER_INTERVAL", base_config.CRAWLER_MAX_SLEEP_SEC))
    
    print(f"Starting crawl with interval={interval}s")
    
    zhihu = ZhihuClient()
    xhs = XHSClient()
    
    # Parallel execution with same interval

    results = await asyncio.gather(
        zhihu.get_all_notes_by_creator_url(
            url_token="example-creator",
            crawl_interval=interval
        ),
        xhs.get_notes_by_keyword(
            keyword="product-review",
            crawl_interval=interval,
            max_notes=1000
        ),
        return_exceptions=True
    )
    
    # Handle any rate-limit exceptions

    for platform, result in zip(["zhihu", "xhs"], results):
        if isinstance(result, Exception):
            print(f"{platform} failed: {result}")
        else:
            print(f"{platform}: collected {len(result)} items")

if __name__ == "__main__":
    asyncio.run(main())

Run with elevated safety:

export CRAWLER_INTERVAL=3.5
export CRAWLER_SLEEP=3.5  # Fallback if interval not set

python production_crawl.py

Key Source Files Reference

File Purpose
main/config/base_config.py Global configuration including CRAWLER_MAX_SLEEP_SEC
media_platform/zhihu/core.py Zhihu crawling logic with interval enforcement
media_platform/xhs/core.py Xiaohongshu crawling implementation
media_platform/weibo/core.py Weibo crawling with rate-limit handling
media_platform/tieba/core.py Tieba crawling implementation
media_platform/kuaishou/core.py KuaiShou crawling with jitter
media_platform/douyin/core.py Douyin crawling implementation
media_platform/bilibili/core.py Bilibili crawling implementation
media_platform/zhihu/client.py Zhihu client with crawl_interval parameter
media_platform/xhs/client.py XHS client with per-call interval override
media_platform/kuaishou/client.py KuaiShou client with random jitter
media_platform/xhs/exception.py RateLimitError definition

Summary

  • Global control: Modify CRAWLER_MAX_SLEEP_SEC in main/config/base_config.py to set a project-wide default crawl interval.

  • Per-call precision: Pass crawl_interval to any client method (get_all_notes_by_creator_url, get_notes_by_keyword, etc.) for targeted adjustments.

  • Platform awareness: Weibo and Zhihu require conservative settings (3-5s); Bilibili tolerates more aggressive pacing (1-2s).

  • Obfuscation: KuaiShou automatically adds 1-3s random jitter—maintain this behavior in production.

  • Resilience: Exponential backoff on RateLimitError provides automatic recovery; monitor logs to iteratively optimize intervals.


Frequently Asked Questions

How do I find the optimal crawl interval for a specific platform?

Start with the documented rate limit or recommended values in this guide, then monitor logs for 429 responses. If you see rate-limit errors within the first 100 requests, increase the interval by 0.5-1.0 seconds and retry. For undocumented platforms, begin at 3.0 seconds and decrease gradually while observing ban patterns.

Can I disable rate limiting entirely for faster testing?

Setting crawl_interval=0 or CRAWLER_MAX_SLEEP_SEC=0 removes artificial delays, but this will trigger immediate platform blocks. For local testing with cached responses, use crawl_interval=0.1 minimum. Never deploy zero-delay configurations against live platforms.

Does MediaCrawler support dynamic interval adjustment based on response headers?

Currently, the framework does not parse Retry-After headers or adaptive rate-limit responses. The crawl_interval and CRAWLER_MAX_SLEEP_SEC values are static once set. Implement custom logic in your callback functions to read response metadata and adjust subsequent calls if needed.

What's the difference between crawl_interval and the retry backoff delay?

crawl_interval controls intentional spacing between successful requests—your primary anti-ban mechanism. The retry backoff activates only after rate-limit errors, using exponential delays (1s, 2s, 4s, 8s) to recover from temporary blocks without manual intervention.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →