Does MediaCrawler Support Rate Limiting? Yes—Here's How It Works
Yes, MediaCrawler implements built-in rate limiting through configurable sleep intervals, concurrency caps, automatic proxy rotation, and platform-specific retry logic with exponential back-off.
MediaCrawler enforces polite crawling across all supported platforms to prevent IP blocks and service disruptions. The NanmiCoder/MediaCrawler repository provides multiple layers of rate-limit protection that you can customize via config/base_config.py or override at runtime.
How Rate Limiting Works in MediaCrawler
The crawler combines preemptive throttling (sleeping between requests) with reactive handling (detecting rate-limit responses and retrying). Here is how each platform implements these strategies.
Xiaohongshu (XHS): Sleep After Every Request
In media_platform/xhs/core.py, the crawler pauses execution after fetching note details or comments:
# media_platform/xhs/core.py (excerpt)
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
When the platform raises IPBlockError or PlatformAccessError, the code logs guidance to reduce frequency, switch IPs, or verify account status.
Weibo: Consistent Sleep Intervals
The Weibo module in media_platform/weibo/core.py applies the same global sleep constant after every page and detail fetch:
# media_platform/weibo/core.py (excerpt)
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
Kuaishou: Exponential Back-Off on Rate-Limit Detection
The Kuaishou client in media_platform/kuaishou/client.py detects server-side throttling explicitly. When result == 2, it triggers exponential back-off:
# media_platform/kuaishou/client.py (excerpt)
if result.get("result") == 2:
delay = 5 * (2**attempt) + random.uniform(0, 2)
utils.logger.warning(
f"[KuaiShouClient.request_rest_v2_signed] rate limited (result:2) on {uri}, "
f"retry in {delay:.1f}s, attempt {attempt + 1}/{max_retry}"
)
await asyncio.sleep(delay)
continue
The delay scales with each retry attempt, adding random jitter to prevent synchronized retry storms.
Configuring Rate-Limit Settings
All rate-limit controls live in config/base_config.py. Adjust these three parameters to tune crawler behavior:
| Parameter | Default | Purpose |
|---|---|---|
CRAWLER_MAX_SLEEP_SEC |
2 |
Seconds to pause between individual requests |
MAX_CONCURRENCY_NUM |
5 |
Maximum parallel async tasks |
ENABLE_IP_PROXY |
False |
Toggle automatic proxy pool rotation |
Reduce Request Frequency
Increase the sleep interval to slow down your crawler:
import config
config.CRAWLER_MAX_SLEEP_SEC = 5 # 5-second pause between requests
Limit Parallelism
Lower concurrency to reduce simultaneous connection load:
config.MAX_CONCURRENCY_NUM = 3 # cap at 3 concurrent tasks
Enable Proxy Rotation
Rotate IPs automatically to distribute request volume:
# config/base_config.py
ENABLE_IP_PROXY = True
IP_PROXY_POOL_COUNT = 10
Platform-Specific Rate-Limit Behaviors
| Platform | File | Strategy |
|---|---|---|
| Xiaohongshu | media_platform/xhs/core.py |
Fixed sleep after each fetch; exception handling for blocks |
media_platform/weibo/core.py |
Fixed sleep in search and detail loops | |
| Kuaishou | media_platform/kuaishou/client.py |
Detect result: 2, retry with exponential back-off |
The docs/excel_export_guide.md file explicitly warns users to "check IP/rate limits" when exporting data, reinforcing the importance of these controls.
Handling Rate-Limit Exceptions
When platforms enforce hard limits, MediaCrawler surfaces actionable errors:
- IPBlockError: Your IP has been temporarily banned. Solution: enable
ENABLE_IP_PROXYor wait. - PlatformAccessError: Account-level restriction. Solution: verify credentials or reduce
MAX_CONCURRENCY_NUM.
Both exceptions log remediation suggestions directly in the console output.
Summary
- MediaCrawler supports rate limiting across all platforms via
CRAWLER_MAX_SLEEP_SEC,MAX_CONCURRENCY_NUM, andENABLE_IP_PROXY. - Xiaohongshu and Weibo use fixed sleep intervals between requests.
- Kuaishou detects rate-limit responses and applies exponential back-off with jitter.
- Central configuration in
config/base_config.pycontrols global behavior. - Proxy rotation bypasses IP-based rate limits when enabled.
Frequently Asked Questions
How do I make MediaCrawler crawl slower to avoid blocks?
Increase CRAWLER_MAX_SLEEP_SEC in config/base_config.py. For aggressive platforms, values of 5–10 seconds are common. Reduce MAX_CONCURRENCY_NUM to 1–3 for single-threaded, sequential crawling.
Does MediaCrawler automatically retry when rate limited?
Only Kuaishou implements automatic retry with exponential back-off in media_platform/kuaishou/client.py. For Xiaohongshu and Weibo, the crawler sleeps preemptively; if a block occurs, you must adjust configuration and restart.
Can I use rotating proxies with MediaCrawler?
Yes. Set ENABLE_IP_PROXY = True and configure IP_PROXY_POOL_COUNT in config/base_config.py. The proxy pool rotates automatically, distributing requests across multiple IPs to evade rate limits.
What happens if my IP gets blocked while crawling?
MediaCrawler raises IPBlockError (Xiaohongshu) or logs platform-specific warnings. Enable proxy rotation or increase CRAWLER_MAX_SLEEP_SEC, then restart the crawler. The docs/excel_export_guide.md documentation specifically recommends checking "IP/rate limits" when troubleshooting blocked requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →