How to Configure Rate Limiting and Crawl Intervals in MediaCrawler to Avoid Platform Bans
Set config.CRAWLER_MAX_SLEEP_SEC in base_config.py for a global crawl interval, or pass a custom crawl_interval parameter to any client method to fine-tune request pacing per platform.
MediaCrawler, the open-source Python framework for scraping Chinese social media platforms (Xiaohongshu, Zhihu, Weibo, Douyin, Bilibili, KuaiShou, and Tieba), implements built-in throttling mechanisms to help users avoid rate limits and account bans. Understanding how to configure these controls—both globally and per-call—is essential for reliable, long-running data collection.
How MediaCrawler Handles Rate Limiting
The framework uses a two-layer throttling strategy: a configurable global default that applies to all platforms, plus optional per-method overrides for granular control.
Global Crawl Interval via CRAWLER_MAX_SLEEP_SEC
Every platform crawler imports the shared configuration module and uses config.CRAWLER_MAX_SLEEP_SEC as its baseline delay between HTTP requests. This constant is defined in main/config/base_config.py and propagated across all core implementations:
- Zhihu:
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC - Xiaohongshu (XHS):
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC - Weibo:
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC - Tieba:
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC - KuaiShou:
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC - Douyin:
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC - Bilibili:
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC
The core implementations apply this value as a fixed asyncio.sleep() pause after each request, ensuring consistent spacing regardless of response time.
Per-Call Override with crawl_interval Parameter
Every public client method that performs paginated fetching exposes a crawl_interval: float = 1.0 parameter. This allows runtime adjustment without modifying configuration files.
For example, in ZhihuClient.get_all_notes_by_creator_url():
async def get_all_notes_by_creator_url(
self,
url_token: str,
crawl_interval: float = 1.0,
callback: Optional[Callable] = None
) -> List[Dict]:
"""
Fetch all notes from a Zhihu creator with configurable throttling.
Args:
crawl_interval: Seconds to sleep between requests.
Defaults to 1.0; override to increase safety margin.
"""
return await self._fetch_paginated(
endpoint=f"/creators/{url_token}/articles",
crawl_interval=crawl_interval,
callback=callback
)
The value flows through to the underlying core routine:
# In zhihu/core.py - the actual sleep implementation
await asyncio.sleep(crawl_interval)
Platform-Specific Throttling Behaviors
KuaiShou: Random Jitter for Pattern Obfuscation
The KuaiShou client adds stochastic variation to avoid predictable request timing. In media_platform/kuaishou/client.py:
import random
async def get_notes_by_creator(
self,
creator_id: str,
crawl_interval: float = 1.0
) -> List[Dict]:
notes = []
for page in self._paginate(creator_id):
notes.extend(page)
# Random jitter: 1-3 second additional delay
jitter = random.uniform(1, 3)
await asyncio.sleep(crawl_interval + jitter)
return notes
This makes traffic signatures harder to fingerprint. Keep jitter enabled for production deployments.
Rate-Limit Detection and Retry Backoff
MediaCrawler's HTTP utilities raise RateLimitError (defined in media_platform/xhs/exception.py and similar modules) when platforms signal throttling. The async_retry decorator implements exponential backoff:
| Attempt | Delay |
|---|---|
| 1st retry | 1 second |
| 2nd retry | 2 seconds |
| 3rd retry | 4 seconds |
| 4th+ retry | 8 seconds (cap) |
The crawler logs each retry with platform-specific context:
[2024-01-15 09:23:41] WARNING [XiaoHongShuCrawler] Rate limit hit (429), backing off 4.0s
[2024-01-15 09:23:45] INFO [XiaoHongShuCrawler] Resuming crawl with interval=2.5
Configuration Guide: Avoiding Platform Bans
Step 1: Tune the Global Default
Edit main/config/base_config.py:
# main/config/base_config.py
# Base crawl interval in seconds (float)
# Increase for aggressive platforms (Weibo, Zhihu): 2.0 - 5.0
# Decrease for tolerant platforms only with proxy rotation: 0.5 - 1.0
CRAWLER_MAX_SLEEP_SEC = 2.5
Recommended starting values by platform:
| Platform | Recommended CRAWLER_MAX_SLEEP_SEC |
Notes |
|---|---|---|
| Zhihu | 2.5 - 4.0 | Strict rate limiting; frequent 429s below 2s |
| Xiaohongshu | 1.5 - 3.0 | Moderate limits; CSRF protection active |
| 3.0 - 5.0 | Very aggressive blocking; prioritize accounts with age | |
| Tieba | 2.0 - 3.0 | IP-based limits; combine with proxy pools |
| KuaiShou | 1.5 - 2.5 | Jitter adds 1-3s automatically |
| Douyin | 2.0 - 4.0 | Signature-based detection; interval helps evasion |
| Bilibili | 1.0 - 2.0 | Most tolerant; can reduce if using rotating IPs |
Step 2: Apply Per-Call Overrides for Sensitive Operations
When fetching high-value targets or running concurrent jobs, increase spacing explicitly:
from media_platform.zhihu.client import ZhihuClient
from media_platform.xhs.client import XHSClient
import asyncio
async def safe_batch_crawl():
"""Crawl multiple creators with conservative delays."""
zhihu = ZhihuClient()
xhs = XHSClient()
# High-priority Zhihu creator: extra cautious
zhihu_data = await zhihu.get_all_notes_by_creator_url(
url_token="zhangsan-123",
crawl_interval=4.0, # 4 seconds between requests
callback=process_note
)
# Xiaohongshu with moderate pace
xhs_data = await xhs.get_notes_by_keyword(
keyword="skincare",
crawl_interval=2.5,
max_notes=500
)
return zhihu_data, xhs_data
asyncio.run(safe_batch_crawl())
Step 3: Calculate Intervals from Platform Limits
Some platform APIs document explicit rate limits. Convert requests-per-minute to crawl intervals:
def interval_from_rpm(max_requests_per_minute: int, safety_factor: float = 0.8) -> float:
"""
Calculate safe crawl interval from documented rate limit.
safety_factor: Use 80% of limit to leave headroom.
"""
seconds_per_request = 60.0 / (max_requests_per_minute * safety_factor)
return round(seconds_per_request, 2)
# Example: Platform allows 30 requests/minute
# Safe interval: 60 / (30 * 0.8) = 2.5 seconds
crawl_interval = interval_from_rpm(30) # → 2.5
Step 4: Monitor and Iterate
Enable verbose logging to observe actual sleep durations and rate-limit events:
# In your entry script or config
import logging
logging.getLogger("MediaCrawler").setLevel(logging.DEBUG)
Watch for these patterns in logs:
"Sleeping for X.XXs"— normal operation, verify interval matches config"Rate limit hit (429)"— increase interval or add proxy rotation"Sleeping for X.XXs + jitter"— KuaiShou jitter active, monitor total delay
Complete Configuration Example
File: main/config/base_config.py
"""
MediaCrawler global configuration.
Adjust these values before starting production crawls.
"""
import os
# Core throttling parameter
# Default: 1.0s, Recommended production: 2.0-5.0s depending on platform
CRAWLER_MAX_SLEEP_SEC = float(os.getenv("CRAWLER_SLEEP", "2.5"))
# Retry behavior
MAX_RETRIES = 3 # Exponential backoff attempts
RETRY_BACKOFF_BASE = 1.0 # First retry delay in seconds
# Proxy configuration (recommended for high-volume crawling)
ENABLE_PROXY = False
PROXY_POOL_URL = "http://proxy-provider:5010/get/"
# Platform-specific overrides (optional)
PLATFORM_OVERRIDES = {
"weibo": {"crawl_interval": 4.0},
"zhihu": {"crawl_interval": 3.0},
"xhs": {"crawl_interval": 2.0},
}
File: production_crawl.py
#!/usr/bin/env python3
"""
Production crawl script with environment-aware throttling.
"""
import os
import asyncio
from media_platform.zhihu.client import ZhihuClient
from media_platform.xhs.client import XHSClient
from main.config import base_config
async def main():
# Environment override takes precedence
interval = float(os.getenv("CRAWLER_INTERVAL", base_config.CRAWLER_MAX_SLEEP_SEC))
print(f"Starting crawl with interval={interval}s")
zhihu = ZhihuClient()
xhs = XHSClient()
# Parallel execution with same interval
results = await asyncio.gather(
zhihu.get_all_notes_by_creator_url(
url_token="example-creator",
crawl_interval=interval
),
xhs.get_notes_by_keyword(
keyword="product-review",
crawl_interval=interval,
max_notes=1000
),
return_exceptions=True
)
# Handle any rate-limit exceptions
for platform, result in zip(["zhihu", "xhs"], results):
if isinstance(result, Exception):
print(f"{platform} failed: {result}")
else:
print(f"{platform}: collected {len(result)} items")
if __name__ == "__main__":
asyncio.run(main())
Run with elevated safety:
export CRAWLER_INTERVAL=3.5
export CRAWLER_SLEEP=3.5 # Fallback if interval not set
python production_crawl.py
Key Source Files Reference
| File | Purpose |
|---|---|
main/config/base_config.py |
Global configuration including CRAWLER_MAX_SLEEP_SEC |
media_platform/zhihu/core.py |
Zhihu crawling logic with interval enforcement |
media_platform/xhs/core.py |
Xiaohongshu crawling implementation |
media_platform/weibo/core.py |
Weibo crawling with rate-limit handling |
media_platform/tieba/core.py |
Tieba crawling implementation |
media_platform/kuaishou/core.py |
KuaiShou crawling with jitter |
media_platform/douyin/core.py |
Douyin crawling implementation |
media_platform/bilibili/core.py |
Bilibili crawling implementation |
media_platform/zhihu/client.py |
Zhihu client with crawl_interval parameter |
media_platform/xhs/client.py |
XHS client with per-call interval override |
media_platform/kuaishou/client.py |
KuaiShou client with random jitter |
media_platform/xhs/exception.py |
RateLimitError definition |
Summary
-
Global control: Modify
CRAWLER_MAX_SLEEP_SECinmain/config/base_config.pyto set a project-wide default crawl interval. -
Per-call precision: Pass
crawl_intervalto any client method (get_all_notes_by_creator_url,get_notes_by_keyword, etc.) for targeted adjustments. -
Platform awareness: Weibo and Zhihu require conservative settings (3-5s); Bilibili tolerates more aggressive pacing (1-2s).
-
Obfuscation: KuaiShou automatically adds 1-3s random jitter—maintain this behavior in production.
-
Resilience: Exponential backoff on
RateLimitErrorprovides automatic recovery; monitor logs to iteratively optimize intervals.
Frequently Asked Questions
How do I find the optimal crawl interval for a specific platform?
Start with the documented rate limit or recommended values in this guide, then monitor logs for 429 responses. If you see rate-limit errors within the first 100 requests, increase the interval by 0.5-1.0 seconds and retry. For undocumented platforms, begin at 3.0 seconds and decrease gradually while observing ban patterns.
Can I disable rate limiting entirely for faster testing?
Setting crawl_interval=0 or CRAWLER_MAX_SLEEP_SEC=0 removes artificial delays, but this will trigger immediate platform blocks. For local testing with cached responses, use crawl_interval=0.1 minimum. Never deploy zero-delay configurations against live platforms.
Does MediaCrawler support dynamic interval adjustment based on response headers?
Currently, the framework does not parse Retry-After headers or adaptive rate-limit responses. The crawl_interval and CRAWLER_MAX_SLEEP_SEC values are static once set. Implement custom logic in your callback functions to read response metadata and adjust subsequent calls if needed.
What's the difference between crawl_interval and the retry backoff delay?
crawl_interval controls intentional spacing between successful requests—your primary anti-ban mechanism. The retry backoff activates only after rate-limit errors, using exponential delays (1s, 2s, 4s, 8s) to recover from temporary blocks without manual intervention.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →