How to Control Concurrency in MediaCrawler Using MAX_CONCURRENCY_NUM

Set MAX_CONCURRENCY_NUM in config/base_config.py or via the --max_concurrency_num CLI flag to throttle parallel requests using an asyncio.Semaphore, preventing rate limits while optimizing throughput.

The MAX_CONCURRENCY_NUM parameter is the central throttle governing asynchronous request parallelism in the NanmiCoder/MediaCrawler repository. Defined in the base configuration and enforced through semaphore-based synchronization, this setting balances crawling speed against server-side rate limits and local resource consumption. Understanding how to tune this value is essential for stable, high-performance data extraction across platforms like Zhihu, XiaoHongShu (XHS), Weibo, and Tieba.

Understanding MAX_CONCURRENCY_NUM Implementation

MediaCrawler implements concurrency control through a configurable semaphore pattern that spans configuration, CLI arguments, and platform-specific core modules.

Configuration and Defaults

The default concurrency limit is defined in config/base_config.py as MAX_CONCURRENCY_NUM = 1. This conservative default ensures safe operation out-of-the-box, preventing accidental denial-of-service patterns against target platforms. The constant is imported throughout the codebase as the single source of truth for parallel request limits.

Command-Line Override

Runtime adjustment is available via the --max_concurrency_num argument parsed in cmd_arg/arg.py (line 362). This CLI flag updates the global configuration without requiring source file modifications, enabling per-run customization based on target platform tolerance or network conditions.

Semaphore Enforcement

Each platform core instantiates an asyncio.Semaphore using the configured value. In media_platform/zhihu/core.py (lines 216 and 377), the semaphore is created as asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) and injected into coroutines performing network I/O. Identical patterns appear in media_platform/xhs/core.py, media_platform/weibo/core.py, and other platform modules. When the limit is reached, subsequent coroutines await semaphore release, ensuring simultaneous outbound requests never exceed the threshold.

Why Concurrency Control Matters

Uncontrolled parallelism creates cascading failures across multiple system layers:

  • Remote service rate limits trigger HTTP 429 responses or temporary IP bans when request volumes exceed platform tolerance.
  • Local resource exhaustion occurs when excessive concurrency saturates the asyncio event loop, increases context-switch overhead, and inflates memory footprints through unbounded connection pools.
  • Network congestion emerges on constrained connections, causing packet loss and timeout cascades that reduce effective throughput.
  • Compliance risks violate terms of service and fair-use policies, potentially exposing crawling operations to legal or access restrictions.

Best Practices for Configuring MAX_CONCURRENCY_NUM

Start with the Safe Default

Retain MAX_CONCURRENCY_NUM = 1 in config/base_config.py until you establish baseline metrics for your target platform. This single-threaded default prevents accidental rate limit violations during initial testing and development.

Increase Based on Empirical Testing

Scale concurrency deliberately using measured data rather than guesswork:

  1. Run a 10-second benchmark crawl with an elevated concurrency value.
  2. Monitor error rates (particularly HTTP 429 and timeouts) and system metrics (CPU, RAM).
  3. Select the highest value maintaining sub-5% error rates and acceptable resource utilization.

Align with API Rate Limits

When platform documentation specifies requests-per-second (RPS) limits, calculate a safe concurrency count using the formula: MAX_CONCURRENCY_NUM ≈ allowed_RPS / average_response_time. For example, if an API permits 10 RPS and average response latency is 0.8 seconds, a safe limit is approximately 12 concurrent requests.

Use CLI Flags for Runtime Adjustment

Avoid editing configuration files for temporary changes. Use the command-line interface for per-run granularity:


# Run with up to 8 concurrent tasks

python -m MediaCrawler.main --max_concurrency_num 8

The argument handler in cmd_arg/arg.py maps this flag directly to the configuration object, overriding the base default without code modification.

Avoid Hard-Coding Values

Never embed literal concurrency integers in platform core logic. Always reference config.MAX_CONCURRENCY_NUM when instantiating semaphores. This indirection ensures CLI overrides propagate correctly and centralizes configuration management in base_config.py.

Combine with Exponential Back-Off

Concurrency limits prevent rate limits but do not eliminate them. Implement exponential back-off retries when encountering HTTP 429 responses rather than increasing semaphore counts reactively. This resilience pattern complements static concurrency controls by handling transient throttling events gracefully.

Monitor Concurrency Usage

Insert debug logging around semaphore acquisition and release points in platform cores to verify limits are respected during long-running operations. Add instrumentation to track actual concurrency utilization versus configured limits, identifying opportunities for safe increases.

Consider Per-Endpoint Limits

When platforms expose sub-endpoints with varying rate limits (e.g., search versus detail APIs), instantiate separate semaphores per endpoint rather than using a single global limit. This granular approach maximizes throughput for permissive endpoints while protecting restricted ones.

Code Examples and Implementation Details

Configuring via Command Line

The CLI flag provides the most flexible adjustment mechanism:


# Conservative setting for sensitive platforms

python -m MediaCrawler.main --max_concurrency_num 2

# Aggressive setting for high-throughput scenarios

python -m MediaCrawler.main --max_concurrency_num 16

As implemented in cmd_arg/arg.py, these values populate the configuration object before platform core initialization.

Semaphore Usage in Platform Cores

The following pattern from media_platform/zhihu/core.py (lines 216, 377) demonstrates proper semaphore integration:

import asyncio
import httpx
from config import base_config

# Semaphore instantiation using configured limit

semaphore = asyncio.Semaphore(base_config.MAX_CONCURRENCY_NUM)

async def fetch_page(url: str):
    async with semaphore:  # Respects concurrency limit

        async with httpx.AsyncClient() as client:
            response = await client.get(url)
            response.raise_for_status()
            return response.json()

All platform implementations follow this identical pattern, ensuring consistent throttling across Zhihu, XHS, Weibo, Tieba, and other supported services.

Dynamic Adjustment Based on Latency

Calculate concurrency limits dynamically using observed performance metrics:

def compute_max_concurrency(avg_response_sec: float, target_rps: int) -> int:
    """
    Calculate safe concurrency based on API limits and measured latency.
    
    Args:
        avg_response_sec: Average HTTP response time in seconds
        target_rps: Maximum requests per second allowed by API
        
    Returns:
        Recommended MAX_CONCURRENCY_NUM value
    """
    return max(1, int(target_rps * avg_response_sec))

# Example: API allows 10 RPS, average response is 0.5 seconds

safe_concurrency = compute_max_concurrency(0.5, 10)
base_config.MAX_CONCURRENCY_NUM = safe_concurrency  # Results in 5

This data-driven approach prevents trial-and-error configuration while respecting platform-specific constraints.

Key Files in the MediaCrawler Repository

Summary

  • Treat MAX_CONCURRENCY_NUM as the single source of truth for request parallelism across all MediaCrawler platforms.
  • Begin with the default value of 1 and increase only after empirical testing confirms platform tolerance.
  • Use the --max_concurrency_num CLI flag for runtime adjustments without modifying source files.
  • Implement semaphore patterns via asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) in all network I/O coroutines.
  • Calculate safe limits using the formula allowed_RPS × average_response_time when API documentation provides rate limiting details.
  • Combine with retry logic and monitoring to handle transient rate limits and verify actual concurrency utilization.

Frequently Asked Questions

What is the default MAX_CONCURRENCY_NUM in MediaCrawler?

The default value is 1, defined in config/base_config.py. This conservative setting ensures single-threaded operation by default, preventing accidental rate limit violations or IP bans when users first run the crawler against target platforms.

How do I override MAX_CONCURRENCY_NUM without modifying config files?

Use the command-line argument --max_concurrency_num when launching the crawler. For example, python -m MediaCrawler.main --max_concurrency_num 5 sets the limit to 5 concurrent requests for that specific execution. The cmd_arg/arg.py module handles parsing at line 362 and updates the global configuration before platform cores initialize their semaphores.

Why does MediaCrawler use asyncio.Semaphore instead of other concurrency controls?

asyncio.Semaphore provides precise control over coroutine-level concurrency without blocking the event loop or creating OS-level threads. This approach is memory-efficient and aligns with MediaCrawler's async/await architecture (using httpx), allowing thousands of pending tasks with only MAX_CONCURRENCY_NUM active connections at any moment.

Can I set different concurrency limits for different platforms?

While the global configuration uses a single MAX_CONCURRENCY_NUM, you can implement per-platform limits by instantiating separate semaphores in individual platform core files (e.g., media_platform/zhihu/core.py vs. media_platform/xhs/core.py). Override the global value locally or add platform-specific configuration parameters to achieve granular control over high-traffic versus sensitive endpoints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →