Does MediaCrawler Support Rate Limiting? Yes—Here's How It Works

Yes, MediaCrawler implements built-in rate limiting through configurable sleep intervals, concurrency caps, automatic proxy rotation, and platform-specific retry logic with exponential back-off.

MediaCrawler enforces polite crawling across all supported platforms to prevent IP blocks and service disruptions. The NanmiCoder/MediaCrawler repository provides multiple layers of rate-limit protection that you can customize via config/base_config.py or override at runtime.


How Rate Limiting Works in MediaCrawler

The crawler combines preemptive throttling (sleeping between requests) with reactive handling (detecting rate-limit responses and retrying). Here is how each platform implements these strategies.

Xiaohongshu (XHS): Sleep After Every Request

In media_platform/xhs/core.py, the crawler pauses execution after fetching note details or comments:


# media_platform/xhs/core.py (excerpt)

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

When the platform raises IPBlockError or PlatformAccessError, the code logs guidance to reduce frequency, switch IPs, or verify account status.

Weibo: Consistent Sleep Intervals

The Weibo module in media_platform/weibo/core.py applies the same global sleep constant after every page and detail fetch:


# media_platform/weibo/core.py (excerpt)

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

Kuaishou: Exponential Back-Off on Rate-Limit Detection

The Kuaishou client in media_platform/kuaishou/client.py detects server-side throttling explicitly. When result == 2, it triggers exponential back-off:


# media_platform/kuaishou/client.py (excerpt)

if result.get("result") == 2:
    delay = 5 * (2**attempt) + random.uniform(0, 2)
    utils.logger.warning(
        f"[KuaiShouClient.request_rest_v2_signed] rate limited (result:2) on {uri}, "
        f"retry in {delay:.1f}s, attempt {attempt + 1}/{max_retry}"
    )
    await asyncio.sleep(delay)
    continue

The delay scales with each retry attempt, adding random jitter to prevent synchronized retry storms.


Configuring Rate-Limit Settings

All rate-limit controls live in config/base_config.py. Adjust these three parameters to tune crawler behavior:

Parameter Default Purpose
CRAWLER_MAX_SLEEP_SEC 2 Seconds to pause between individual requests
MAX_CONCURRENCY_NUM 5 Maximum parallel async tasks
ENABLE_IP_PROXY False Toggle automatic proxy pool rotation

Reduce Request Frequency

Increase the sleep interval to slow down your crawler:

import config
config.CRAWLER_MAX_SLEEP_SEC = 5   # 5-second pause between requests

Limit Parallelism

Lower concurrency to reduce simultaneous connection load:

config.MAX_CONCURRENCY_NUM = 3     # cap at 3 concurrent tasks

Enable Proxy Rotation

Rotate IPs automatically to distribute request volume:


# config/base_config.py

ENABLE_IP_PROXY = True
IP_PROXY_POOL_COUNT = 10

Platform-Specific Rate-Limit Behaviors

Platform File Strategy
Xiaohongshu media_platform/xhs/core.py Fixed sleep after each fetch; exception handling for blocks
Weibo media_platform/weibo/core.py Fixed sleep in search and detail loops
Kuaishou media_platform/kuaishou/client.py Detect result: 2, retry with exponential back-off

The docs/excel_export_guide.md file explicitly warns users to "check IP/rate limits" when exporting data, reinforcing the importance of these controls.


Handling Rate-Limit Exceptions

When platforms enforce hard limits, MediaCrawler surfaces actionable errors:

  • IPBlockError: Your IP has been temporarily banned. Solution: enable ENABLE_IP_PROXY or wait.
  • PlatformAccessError: Account-level restriction. Solution: verify credentials or reduce MAX_CONCURRENCY_NUM.

Both exceptions log remediation suggestions directly in the console output.


Summary

  • MediaCrawler supports rate limiting across all platforms via CRAWLER_MAX_SLEEP_SEC, MAX_CONCURRENCY_NUM, and ENABLE_IP_PROXY.
  • Xiaohongshu and Weibo use fixed sleep intervals between requests.
  • Kuaishou detects rate-limit responses and applies exponential back-off with jitter.
  • Central configuration in config/base_config.py controls global behavior.
  • Proxy rotation bypasses IP-based rate limits when enabled.

Frequently Asked Questions

How do I make MediaCrawler crawl slower to avoid blocks?

Increase CRAWLER_MAX_SLEEP_SEC in config/base_config.py. For aggressive platforms, values of 5–10 seconds are common. Reduce MAX_CONCURRENCY_NUM to 1–3 for single-threaded, sequential crawling.

Does MediaCrawler automatically retry when rate limited?

Only Kuaishou implements automatic retry with exponential back-off in media_platform/kuaishou/client.py. For Xiaohongshu and Weibo, the crawler sleeps preemptively; if a block occurs, you must adjust configuration and restart.

Can I use rotating proxies with MediaCrawler?

Yes. Set ENABLE_IP_PROXY = True and configure IP_PROXY_POOL_COUNT in config/base_config.py. The proxy pool rotates automatically, distributing requests across multiple IPs to evade rate limits.

What happens if my IP gets blocked while crawling?

MediaCrawler raises IPBlockError (Xiaohongshu) or logs platform-specific warnings. Enable proxy rotation or increase CRAWLER_MAX_SLEEP_SEC, then restart the crawler. The docs/excel_export_guide.md documentation specifically recommends checking "IP/rate limits" when troubleshooting blocked requests.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →