# How to Control Concurrency in MediaCrawler Using MAX_CONCURRENCY_NUM

> Optimize MediaCrawler throughput by controlling concurrency. Learn to set MAX_CONCURRENCY_NUM via config or CLI to manage parallel requests and avoid rate limits effectively.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: best-practices
- Published: 2026-06-29

---

**Set `MAX_CONCURRENCY_NUM` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) or via the `--max_concurrency_num` CLI flag to throttle parallel requests using an `asyncio.Semaphore`, preventing rate limits while optimizing throughput.**

The `MAX_CONCURRENCY_NUM` parameter is the central throttle governing asynchronous request parallelism in the NanmiCoder/MediaCrawler repository. Defined in the base configuration and enforced through semaphore-based synchronization, this setting balances crawling speed against server-side rate limits and local resource consumption. Understanding how to tune this value is essential for stable, high-performance data extraction across platforms like Zhihu, XiaoHongShu (XHS), Weibo, and Tieba.

## Understanding MAX_CONCURRENCY_NUM Implementation

MediaCrawler implements concurrency control through a configurable semaphore pattern that spans configuration, CLI arguments, and platform-specific core modules.

### Configuration and Defaults

The default concurrency limit is defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) as `MAX_CONCURRENCY_NUM = 1`. This conservative default ensures safe operation out-of-the-box, preventing accidental denial-of-service patterns against target platforms. The constant is imported throughout the codebase as the single source of truth for parallel request limits.

### Command-Line Override

Runtime adjustment is available via the `--max_concurrency_num` argument parsed in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) (line 362). This CLI flag updates the global configuration without requiring source file modifications, enabling per-run customization based on target platform tolerance or network conditions.

### Semaphore Enforcement

Each platform core instantiates an `asyncio.Semaphore` using the configured value. In [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) (lines 216 and 377), the semaphore is created as `asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)` and injected into coroutines performing network I/O. Identical patterns appear in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py), [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py), and other platform modules. When the limit is reached, subsequent coroutines await semaphore release, ensuring simultaneous outbound requests never exceed the threshold.

## Why Concurrency Control Matters

Uncontrolled parallelism creates cascading failures across multiple system layers:

- **Remote service rate limits** trigger HTTP 429 responses or temporary IP bans when request volumes exceed platform tolerance.
- **Local resource exhaustion** occurs when excessive concurrency saturates the asyncio event loop, increases context-switch overhead, and inflates memory footprints through unbounded connection pools.
- **Network congestion** emerges on constrained connections, causing packet loss and timeout cascades that reduce effective throughput.
- **Compliance risks** violate terms of service and fair-use policies, potentially exposing crawling operations to legal or access restrictions.

## Best Practices for Configuring MAX_CONCURRENCY_NUM

### Start with the Safe Default

Retain `MAX_CONCURRENCY_NUM = 1` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) until you establish baseline metrics for your target platform. This single-threaded default prevents accidental rate limit violations during initial testing and development.

### Increase Based on Empirical Testing

Scale concurrency deliberately using measured data rather than guesswork:

1. Run a 10-second benchmark crawl with an elevated concurrency value.
2. Monitor error rates (particularly HTTP 429 and timeouts) and system metrics (CPU, RAM).
3. Select the highest value maintaining sub-5% error rates and acceptable resource utilization.

### Align with API Rate Limits

When platform documentation specifies requests-per-second (RPS) limits, calculate a safe concurrency count using the formula: `MAX_CONCURRENCY_NUM ≈ allowed_RPS / average_response_time`. For example, if an API permits 10 RPS and average response latency is 0.8 seconds, a safe limit is approximately 12 concurrent requests.

### Use CLI Flags for Runtime Adjustment

Avoid editing configuration files for temporary changes. Use the command-line interface for per-run granularity:

```bash

# Run with up to 8 concurrent tasks

python -m MediaCrawler.main --max_concurrency_num 8

```

The argument handler in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) maps this flag directly to the configuration object, overriding the base default without code modification.

### Avoid Hard-Coding Values

Never embed literal concurrency integers in platform core logic. Always reference `config.MAX_CONCURRENCY_NUM` when instantiating semaphores. This indirection ensures CLI overrides propagate correctly and centralizes configuration management in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py).

### Combine with Exponential Back-Off

Concurrency limits prevent rate limits but do not eliminate them. Implement exponential back-off retries when encountering HTTP 429 responses rather than increasing semaphore counts reactively. This resilience pattern complements static concurrency controls by handling transient throttling events gracefully.

### Monitor Concurrency Usage

Insert debug logging around semaphore acquisition and release points in platform cores to verify limits are respected during long-running operations. Add instrumentation to track actual concurrency utilization versus configured limits, identifying opportunities for safe increases.

### Consider Per-Endpoint Limits

When platforms expose sub-endpoints with varying rate limits (e.g., search versus detail APIs), instantiate separate semaphores per endpoint rather than using a single global limit. This granular approach maximizes throughput for permissive endpoints while protecting restricted ones.

## Code Examples and Implementation Details

### Configuring via Command Line

The CLI flag provides the most flexible adjustment mechanism:

```bash

# Conservative setting for sensitive platforms

python -m MediaCrawler.main --max_concurrency_num 2

# Aggressive setting for high-throughput scenarios

python -m MediaCrawler.main --max_concurrency_num 16

```

As implemented in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py), these values populate the configuration object before platform core initialization.

### Semaphore Usage in Platform Cores

The following pattern from [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) (lines 216, 377) demonstrates proper semaphore integration:

```python
import asyncio
import httpx
from config import base_config

# Semaphore instantiation using configured limit

semaphore = asyncio.Semaphore(base_config.MAX_CONCURRENCY_NUM)

async def fetch_page(url: str):
    async with semaphore:  # Respects concurrency limit

        async with httpx.AsyncClient() as client:
            response = await client.get(url)
            response.raise_for_status()
            return response.json()

```

All platform implementations follow this identical pattern, ensuring consistent throttling across Zhihu, XHS, Weibo, Tieba, and other supported services.

### Dynamic Adjustment Based on Latency

Calculate concurrency limits dynamically using observed performance metrics:

```python
def compute_max_concurrency(avg_response_sec: float, target_rps: int) -> int:
    """
    Calculate safe concurrency based on API limits and measured latency.
    
    Args:
        avg_response_sec: Average HTTP response time in seconds
        target_rps: Maximum requests per second allowed by API
        
    Returns:
        Recommended MAX_CONCURRENCY_NUM value
    """
    return max(1, int(target_rps * avg_response_sec))

# Example: API allows 10 RPS, average response is 0.5 seconds

safe_concurrency = compute_max_concurrency(0.5, 10)
base_config.MAX_CONCURRENCY_NUM = safe_concurrency  # Results in 5

```

This data-driven approach prevents trial-and-error configuration while respecting platform-specific constraints.

## Key Files in the MediaCrawler Repository

- **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)**: Defines the default `MAX_CONCURRENCY_NUM` constant and configuration class.
- **[`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py)**: Implements `--max_concurrency_num` CLI argument parsing (line 362).
- **[`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py)**: Reference implementation of semaphore creation and usage (lines 216, 377).
- **[`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py)**: Additional semaphore usage example for XiaoHongShu platform.
- **[`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py)**: Demonstrates concurrency control in Weibo crawling logic.
- **`docs/项目架构文档.md`**: Architecture documentation referencing concurrency settings (line 679).

## Summary

- **Treat `MAX_CONCURRENCY_NUM` as the single source of truth** for request parallelism across all MediaCrawler platforms.
- **Begin with the default value of 1** and increase only after empirical testing confirms platform tolerance.
- **Use the `--max_concurrency_num` CLI flag** for runtime adjustments without modifying source files.
- **Implement semaphore patterns** via `asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)` in all network I/O coroutines.
- **Calculate safe limits** using the formula `allowed_RPS × average_response_time` when API documentation provides rate limiting details.
- **Combine with retry logic** and monitoring to handle transient rate limits and verify actual concurrency utilization.

## Frequently Asked Questions

### What is the default MAX_CONCURRENCY_NUM in MediaCrawler?

The default value is `1`, defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). This conservative setting ensures single-threaded operation by default, preventing accidental rate limit violations or IP bans when users first run the crawler against target platforms.

### How do I override MAX_CONCURRENCY_NUM without modifying config files?

Use the command-line argument `--max_concurrency_num` when launching the crawler. For example, `python -m MediaCrawler.main --max_concurrency_num 5` sets the limit to 5 concurrent requests for that specific execution. The [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) module handles parsing at line 362 and updates the global configuration before platform cores initialize their semaphores.

### Why does MediaCrawler use asyncio.Semaphore instead of other concurrency controls?

`asyncio.Semaphore` provides precise control over coroutine-level concurrency without blocking the event loop or creating OS-level threads. This approach is memory-efficient and aligns with MediaCrawler's async/await architecture (using `httpx`), allowing thousands of pending tasks with only `MAX_CONCURRENCY_NUM` active connections at any moment.

### Can I set different concurrency limits for different platforms?

While the global configuration uses a single `MAX_CONCURRENCY_NUM`, you can implement per-platform limits by instantiating separate semaphores in individual platform core files (e.g., [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) vs. [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py)). Override the global value locally or add platform-specific configuration parameters to achieve granular control over high-traffic versus sensitive endpoints.