How to Configure Rate Limiting and Sleep Intervals Using CRAWLER_MAX_SLEEP_SEC in MediaCrawler

Set the CRAWLER_MAX_SLEEP_SEC constant in config/base_config.py to control the pause duration between HTTP requests across all crawler platforms.

MediaCrawler implements request throttling through a unified sleep mechanism that prevents API rate limit violations and mimics human browsing patterns. The CRAWLER_MAX_SLEEP_SEC configuration value serves as the single source of truth for pacing across every supported platform including Zhihu, XiaoHongShu, Tieba, and Weibo. Understanding how to configure this parameter ensures optimal scraping speeds while minimizing the risk of IP bans or service blocks.

Understanding CRAWLER_MAX_SLEEP_SEC

The global constant CRAWLER_MAX_SLEEP_SEC defines the default interval (in seconds) that the crawler waits between consecutive HTTP requests. Defined at line 137 in config/base_config.py, this value propagates to all platform-specific implementations through a shared configuration object.

When the crawler executes, it calls await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) after each network operation. This asynchronous yield returns control to the event loop, allowing other coroutines to process while enforcing the mandatory delay. The implementation guarantees consistent pacing behavior regardless of which social media platform the crawler targets.

How Rate Limiting Works Under the Hood

MediaCrawler's rate limiting operates through three distinct phases:

  1. Configuration Loading – At import time, Python executes config/base_config.py, binding CRAWLER_MAX_SLEEP_SEC to the global config namespace.
  2. Request Execution – Platform crawlers perform HTTP operations using aiohttp or similar async libraries.
  3. Throttling Application – Immediately following each response, the crawler invokes asyncio.sleep() using the configured interval before processing the next item.

This pattern appears consistently across the codebase. For example, in media_platform/zhihu/core.py at lines 188-191, the Zhihu crawler explicitly logs the sleep duration:

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
utils.logger.info(f"[ZhihuCrawler.search] Sleeping for "
                  f"{config.CRAWLER_MAX_SLEEP_SEC} seconds after page {page-1}")

Identical implementations exist in media_platform/tieba/core.py (lines 200-201) and media_platform/weibo/core.py (lines 187-188), ensuring uniform behavior across platforms.

Configuring Sleep Intervals

Modifying the Global Configuration

The simplest method to adjust rate limiting involves editing the source configuration file directly. Open config/base_config.py and modify line 137:


# config/base_config.py

CRAWLER_MAX_SLEEP_SEC = 5  # Pauses 5 seconds between requests

This change applies universally to all crawler instances upon the next execution. Values typically range from 1 (aggressive) to 10+ seconds (conservative), depending on the target platform's tolerance.

Runtime Override via CLI Arguments

For dynamic adjustment without code modification, override the configuration at startup through the main entry point. The following pattern allows command-line specification of sleep intervals:


# main.py (entry point implementation)

import argparse
import config

parser = argparse.ArgumentParser()
parser.add_argument("--sleep-sec", type=int, 
                    help="Override the default crawler sleep interval")
args = parser.parse_args()

if args.sleep_sec:
    config.CRAWLER_MAX_SLEEP_SEC = args.sleep_sec

Execute the crawler with custom timing:

python main.py --sleep-sec 3

This approach assigns 3 to config.CRAWLER_MAX_SLEEP_SEC, overriding the default value before any crawler tasks initialize.

Platform-Specific Implementation Examples

Each media platform follows the same throttling pattern while maintaining separate logging contexts. The following examples demonstrate the consistency:

Zhihu Implementation (media_platform/zhihu/core.py):

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

Tieba Implementation (media_platform/tieba/core.py):

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

Weibo Implementation (media_platform/weibo/core.py):

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

All platforms reference the identical config object, ensuring that changing CRAWLER_MAX_SLEEP_SEC in one location affects the entire application uniformly.

Balancing Throughput and API Safety

The sleep interval interacts with other configuration limits to shape overall system throughput. MediaCrawler respects maximum content boundaries defined by CRAWLER_MAX_NOTES_COUNT and CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES.

When combined with higher sleep values, these limits create natural rate caps that prevent excessive API consumption. For instance, crawling 100 notes with a 5-second sleep interval generates a minimum total execution time of 500 seconds (approximately 8.3 minutes), keeping request volumes well within typical platform rate limits.

The logging system records every sleep event with the exact duration, enabling audit trails for debugging throttling behavior or verifying compliance with target service terms.

Summary

  • CRAWLER_MAX_SLEEP_SEC in config/base_config.py controls the global delay between HTTP requests
  • The crawler uses await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) after each network call across all platforms
  • Modify the constant directly in base_config.py or override it programmatically at runtime
  • Sleep intervals work in conjunction with content limits (CRAWLER_MAX_NOTES_COUNT) to manage total throughput
  • All sleep events are logged with duration details for operational transparency

Frequently Asked Questions

What is the default value for CRAWLER_MAX_SLEEP_SEC in MediaCrawler?

The default value varies by version, but the constant is defined at line 137 of config/base_config.py. Users should inspect this file directly, as the repository maintainers may adjust the baseline to reflect current platform tolerances. Typical defaults range between 1 and 3 seconds.

Can I set different sleep intervals for different platforms?

While CRAWLER_MAX_SLEEP_SEC is global, you can implement platform-specific overrides by conditionally checking the crawler type before the sleep call. However, the stock MediaCrawler implementation uses a unified constant to ensure consistent behavior across all media_platform/ modules. Forking the repository allows customization of individual platform files like zhihu/core.py or tieba/core.py to use distinct timing variables.

How does CRAWLER_MAX_SLEEP_SEC interact with randomization or jitter?

The base implementation uses a fixed interval. To add randomization (recommended for avoiding detection patterns), wrap the sleep call with random.uniform():

import random
await asyncio.sleep(random.uniform(1, config.CRAWLER_MAX_SLEEP_SEC))

This modification requires editing the specific platform core files in media_platform/, as the stock configuration does not include built-in jitter support.

Will reducing CRAWLER_MAX_SLEEP_SEC to 0 improve performance?

Setting the value to 0 eliminates intentional delays, significantly increasing request rates and substantially raising the risk of IP bans or account suspension. Most platforms implement automatic rate limiting that returns HTTP 429 (Too Many Requests) or temporary blocks when detecting rapid sequential requests. Maintain intervals of at least 1-2 seconds for production scraping tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →