Tuning MediaCrawler MAX_CONCURRENCY_NUM and CRAWLER_MAX_SLEEP_SEC for Performance

Adjust MAX_CONCURRENCY_NUM and CRAWLER_MAX_SLEEP_SEC in config/base_config.py or via CLI flags to balance crawling speed against rate-limit compliance, where higher concurrency increases parallel HTTP requests and shorter sleep intervals reduce delays between calls.

MediaCrawler is an open-source asynchronous scraping framework supporting platforms like Zhihu, XiaoHongShu, Weibo, and Douyin. Tuning MediaCrawler MAX_CONCURRENCY_NUM and CRAWLER_MAX_SLEEP_SEC for performance allows you to optimize throughput while minimizing the risk of IP bans or CAPTCHAs. These two parameters control the asyncio semaphore size and the fixed delay between requests across all platform implementations.

Understanding the Core Configuration Parameters

MAX_CONCURRENCY_NUM

Located in config/base_config.py at lines 105-138, this global constant defaults to 1 and determines the size of the asyncio.Semaphore instantiated in every crawler's core.py module. When you increase this value, more concurrent tasks can acquire the semaphore simultaneously, allowing parallel HTTP requests to the target platform.

CRAWLER_MAX_SLEEP_SEC

Also defined in config/base_config.py, this parameter defaults to 2 seconds. It controls the fixed pause inserted after each network request via await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC). This throttle protects against aggressive rate limiting but adds latency to the overall crawl duration.

Architectural Implementation Details

Semaphore-Based Concurrency Control

Every media platform crawler follows the same pattern in its respective core.py file. For example, in media_platform/zhihu/core.py at line 16, the code initializes:

semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)

This semaphore is passed to asynchronous tasks such as self.get_comments(content_item, semaphore). The same implementation appears in:

Fixed Sleep Interval Implementation

After executing HTTP requests, crawlers invoke a standardized sleep call. In media_platform/zhihu/core.py at line 42, and similarly in media_platform/xhs/core.py at line 182 and media_platform/weibo/core.py at line 27, the code executes:

await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)

This ensures consistent throttling regardless of the specific platform being crawled.

CLI Overrides and Configuration Methods

Command-Line Interface

The cmd_arg/arg.py module exposes these settings through Typer options. Lines 86-94 define the --max_concurrency_num flag:

max_concurrency_num: Annotated[
    int,
    typer.Option(
        "--max_concurrency_num",
        help="Maximum number of concurrent crawlers",
        rich_help_panel="Performance Configuration",
    ),
] = config.MAX_CONCURRENCY_NUM

At runtime (line 62), the CLI writes parsed values back to the config module:

config.MAX_CONCURRENCY_NUM = max_concurrency_num

Practical Tuning Strategies

High-Throughput Configuration

To maximize scraping speed for large datasets, increase both concurrency and reduce sleep:

python -m MediaCrawler.main --max_concurrency_num 8 --crawler_max_sleep_sec 0.5

This creates an asyncio.Semaphore(8) and reduces the inter-request delay to 0.5 seconds, significantly decreasing total runtime for platforms like XiaoHongShu or Douyin.

Conservative Stability Settings

For platforms with strict anti-scraping measures like Zhihu, use moderate values:

python -m MediaCrawler.main --max_concurrency_num 4 --crawler_max_sleep_sec 1

This configuration maintains four parallel connections while respecting rate limits through one-second pauses between requests.

Programmatic Configuration

When embedding MediaCrawler in Python applications, modify the config module directly before initialization:

from MediaCrawler.config import base_config as config

config.MAX_CONCURRENCY_NUM = 6
config.CRAWLER_MAX_SLEEP_SEC = 1.5

from MediaCrawler.main import run_crawler
run_crawler()

This approach bypasses CLI parsing and allows dynamic configuration based on runtime conditions.

Summary

  • MAX_CONCURRENCY_NUM controls the asyncio.Semaphore size in all platform core.py files, defaulting to 1 in config/base_config.py
  • CRAWLER_MAX_SLEEP_SEC sets the fixed delay between HTTP requests via asyncio.sleep(), defaulting to 2 seconds
  • Adjust these via CLI flags --max_concurrency_num and --crawler_max_sleep_sec defined in cmd_arg/arg.py
  • Higher concurrency increases parallel requests but risks rate limiting; shorter sleep intervals speed up crawls but may trigger anti-bot detection
  • All supported platforms (Zhihu, XHS, Weibo, Tieba, Kuaishou, Douyin, Bilibili) respect these global settings

Frequently Asked Questions

What is the default MAX_CONCURRENCY_NUM in MediaCrawler?

The default value is 1, defined in config/base_config.py. This conservative setting ensures only one concurrent request per platform runs at a time, minimizing the risk of IP bans on strict platforms like Zhihu or Weibo.

How does CRAWLER_MAX_SLEEP_SEC prevent IP bans?

This parameter inserts a fixed pause (default 2 seconds) between HTTP requests via await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) in each platform's core.py. The delay mimics human browsing patterns and avoids triggering rate-limiting algorithms that flag rapid sequential requests.

Can I override these settings via command line?

Yes. The CLI in cmd_arg/arg.py exposes --max_concurrency_num to adjust concurrency and --crawler_max_sleep_sec to modify the sleep interval. These flags override the defaults from config/base_config.py at runtime by writing the parsed values back to the config module.

Do all platforms support these concurrency settings?

Yes. Every platform implementation—including Zhihu, XiaoHongShu, Weibo, Tieba, Kuaishou, Douyin, and Bilibili—initializes asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) and calls asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) in their respective media_platform/{platform}/core.py files, ensuring consistent behavior across the framework.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →