How to Configure Rate Limiting and Sleep Intervals Using CRAWLER_MAX_SLEEP_SEC in MediaCrawler
Set the CRAWLER_MAX_SLEEP_SEC constant in config/base_config.py to control the pause duration between HTTP requests across all crawler platforms.
MediaCrawler implements request throttling through a unified sleep mechanism that prevents API rate limit violations and mimics human browsing patterns. The CRAWLER_MAX_SLEEP_SEC configuration value serves as the single source of truth for pacing across every supported platform including Zhihu, XiaoHongShu, Tieba, and Weibo. Understanding how to configure this parameter ensures optimal scraping speeds while minimizing the risk of IP bans or service blocks.
Understanding CRAWLER_MAX_SLEEP_SEC
The global constant CRAWLER_MAX_SLEEP_SEC defines the default interval (in seconds) that the crawler waits between consecutive HTTP requests. Defined at line 137 in config/base_config.py, this value propagates to all platform-specific implementations through a shared configuration object.
When the crawler executes, it calls await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) after each network operation. This asynchronous yield returns control to the event loop, allowing other coroutines to process while enforcing the mandatory delay. The implementation guarantees consistent pacing behavior regardless of which social media platform the crawler targets.
How Rate Limiting Works Under the Hood
MediaCrawler's rate limiting operates through three distinct phases:
- Configuration Loading – At import time, Python executes
config/base_config.py, bindingCRAWLER_MAX_SLEEP_SECto the global config namespace. - Request Execution – Platform crawlers perform HTTP operations using
aiohttpor similar async libraries. - Throttling Application – Immediately following each response, the crawler invokes
asyncio.sleep()using the configured interval before processing the next item.
This pattern appears consistently across the codebase. For example, in media_platform/zhihu/core.py at lines 188-191, the Zhihu crawler explicitly logs the sleep duration:
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
utils.logger.info(f"[ZhihuCrawler.search] Sleeping for "
f"{config.CRAWLER_MAX_SLEEP_SEC} seconds after page {page-1}")
Identical implementations exist in media_platform/tieba/core.py (lines 200-201) and media_platform/weibo/core.py (lines 187-188), ensuring uniform behavior across platforms.
Configuring Sleep Intervals
Modifying the Global Configuration
The simplest method to adjust rate limiting involves editing the source configuration file directly. Open config/base_config.py and modify line 137:
# config/base_config.py
CRAWLER_MAX_SLEEP_SEC = 5 # Pauses 5 seconds between requests
This change applies universally to all crawler instances upon the next execution. Values typically range from 1 (aggressive) to 10+ seconds (conservative), depending on the target platform's tolerance.
Runtime Override via CLI Arguments
For dynamic adjustment without code modification, override the configuration at startup through the main entry point. The following pattern allows command-line specification of sleep intervals:
# main.py (entry point implementation)
import argparse
import config
parser = argparse.ArgumentParser()
parser.add_argument("--sleep-sec", type=int,
help="Override the default crawler sleep interval")
args = parser.parse_args()
if args.sleep_sec:
config.CRAWLER_MAX_SLEEP_SEC = args.sleep_sec
Execute the crawler with custom timing:
python main.py --sleep-sec 3
This approach assigns 3 to config.CRAWLER_MAX_SLEEP_SEC, overriding the default value before any crawler tasks initialize.
Platform-Specific Implementation Examples
Each media platform follows the same throttling pattern while maintaining separate logging contexts. The following examples demonstrate the consistency:
Zhihu Implementation (media_platform/zhihu/core.py):
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
Tieba Implementation (media_platform/tieba/core.py):
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
Weibo Implementation (media_platform/weibo/core.py):
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
All platforms reference the identical config object, ensuring that changing CRAWLER_MAX_SLEEP_SEC in one location affects the entire application uniformly.
Balancing Throughput and API Safety
The sleep interval interacts with other configuration limits to shape overall system throughput. MediaCrawler respects maximum content boundaries defined by CRAWLER_MAX_NOTES_COUNT and CRAWLER_MAX_COMMENTS_COUNT_SINGLENOTES.
When combined with higher sleep values, these limits create natural rate caps that prevent excessive API consumption. For instance, crawling 100 notes with a 5-second sleep interval generates a minimum total execution time of 500 seconds (approximately 8.3 minutes), keeping request volumes well within typical platform rate limits.
The logging system records every sleep event with the exact duration, enabling audit trails for debugging throttling behavior or verifying compliance with target service terms.
Summary
CRAWLER_MAX_SLEEP_SECinconfig/base_config.pycontrols the global delay between HTTP requests- The crawler uses
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)after each network call across all platforms - Modify the constant directly in
base_config.pyor override it programmatically at runtime - Sleep intervals work in conjunction with content limits (
CRAWLER_MAX_NOTES_COUNT) to manage total throughput - All sleep events are logged with duration details for operational transparency
Frequently Asked Questions
What is the default value for CRAWLER_MAX_SLEEP_SEC in MediaCrawler?
The default value varies by version, but the constant is defined at line 137 of config/base_config.py. Users should inspect this file directly, as the repository maintainers may adjust the baseline to reflect current platform tolerances. Typical defaults range between 1 and 3 seconds.
Can I set different sleep intervals for different platforms?
While CRAWLER_MAX_SLEEP_SEC is global, you can implement platform-specific overrides by conditionally checking the crawler type before the sleep call. However, the stock MediaCrawler implementation uses a unified constant to ensure consistent behavior across all media_platform/ modules. Forking the repository allows customization of individual platform files like zhihu/core.py or tieba/core.py to use distinct timing variables.
How does CRAWLER_MAX_SLEEP_SEC interact with randomization or jitter?
The base implementation uses a fixed interval. To add randomization (recommended for avoiding detection patterns), wrap the sleep call with random.uniform():
import random
await asyncio.sleep(random.uniform(1, config.CRAWLER_MAX_SLEEP_SEC))
This modification requires editing the specific platform core files in media_platform/, as the stock configuration does not include built-in jitter support.
Will reducing CRAWLER_MAX_SLEEP_SEC to 0 improve performance?
Setting the value to 0 eliminates intentional delays, significantly increasing request rates and substantially raising the risk of IP bans or account suspension. Most platforms implement automatic rate limiting that returns HTTP 429 (Too Many Requests) or temporary blocks when detecting rapid sequential requests. Maintain intervals of at least 1-2 seconds for production scraping tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →