# Async Architecture and Concurrency Control (`MAX_CONCURRENCY_NUM`) in MediaCrawler

> Explore MediaCrawler's async architecture and MAX_CONCURRENCY_NUM for safe, serialized crawling. Learn how Python's asyncio and semaphores manage simultaneous operations effectively.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: internals
- Published: 2026-08-14

---

**MediaCrawler uses Python's `asyncio` with a semaphore-based concurrency limiter (`MAX_CONCURRENCY_NUM`) to control how many simultaneous network or browser operations execute at once, defaulting to 1 for safe, serialized crawling.**

MediaCrawler is an open-source multi-platform media crawler that asynchronously fetches data from Zhihu, XiaoHongShu (XHS), Weibo, Tieba, Kuaishou, Douyin, Bilibili, and more. Its async architecture and concurrency control through `MAX_CONCURRENCY_NUM` ensure efficient resource usage while respecting API rate limits. This article explains how the semaphore-based system works, where it's configured, and how to tune it for your crawling workloads.

## What Is `MAX_CONCURRENCY_NUM`?

`MAX_CONCURRENCY_NUM` is a global configuration constant that defines the maximum number of **concurrent coroutines** allowed to perform I/O-bound operations simultaneously. These operations include HTTP requests, browser automation with Playwright, and other network-dependent tasks.

The default value of `1` means all network calls are serialized by default. This conservative setting prevents overwhelming target servers and avoids triggering rate limits or IP bans.

## Where `MAX_CONCURRENCY_NUM` Is Defined and Configured

### Base Configuration File

The constant originates in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) at **line 105**:

```python

# config/base_config.py (excerpt)

MAX_CONCURRENCY_NUM = 1  # Controls simultaneous async operations

```

This file serves as the single source of truth for the default concurrency limit across all platform crawlers.

### CLI Argument Override

Users can override the default at runtime via the command-line interface. The argument parser in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) (lines **293-362**) exposes `--max_concurrency_num`:

```bash

# Run with 5 concurrent operations

python main.py --max_concurrency_num 5 --platform xhs

```

The CLI parser updates `config.MAX_CONCURRENCY_NUM` before any crawler modules initialize, ensuring the custom limit propagates throughout the system.

## How the Semaphore Pattern Works in Practice

Each platform's core module creates an `asyncio.Semaphore` from `MAX_CONCURRENCY_NUM` and passes it to async workers. The semaphore is acquired using `async with semaphore:` before any network or browser operation begins.

### Zhihu Implementation

In [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) (lines **216-226**):

```python

# media_platform/zhihu/core.py

import asyncio
import config

class ZhihuCrawler:
    def __init__(self):
        self.semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
    
    async def fetch_comments(self, answer_id: str):
        async with self.semaphore:  # Concurrency guard

            # HTTP request or browser automation happens here

            async with httpx.AsyncClient() as client:
                response = await client.get(f"/answers/{answer_id}/comments")
                return response.json()

```

### XiaoHongShu (XHS) Implementation

The same pattern appears in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) (lines **166-172**):

```python

# media_platform/xhs/core.py

async def fetch_note_detail(self, note_id: str):
    async with self.semaphore:
        # Render page or call XHS API

        page = await self.browser_context.new_page()
        # ... note extraction logic

```

### Consistent Pattern Across All Platforms

The identical semaphore usage exists in:

- [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py)
- [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py)
- [`media_platform/kuaishou/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/core.py)
- [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py)
- [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py)

This design ensures **unified concurrency control** regardless of which platform you're crawling.

## Practical Usage and Tuning

### Default Behavior (MAX_CONCURRENCY_NUM = 1)

With the default value, MediaCrawler processes one operation at a time. This is ideal for:

- Development and debugging
- Strictly rate-limited APIs
- Low-resource environments (small VPS, residential connections)
- Avoiding IP bans on sensitive platforms

### Increasing Concurrency

Raise the limit when targeting platforms with generous rate limits or when running on capable infrastructure:

```bash

# Example: Crawl XHS with 10 concurrent operations

python main.py --platform xhs --max_concurrency_num 10

# Example: Crawl multiple platforms with different limits

python main.py --platform zhihu --max_concurrency_num 5
python main.py --platform douyin --max_concurrency_num 20

```

### Custom Script Integration

For programmatic usage, reference the config directly:

```python
import asyncio
import config
from media_platform.xhs import XiaoHongShuCrawler

async def main():
    # Override config before instantiation

    config.MAX_CONCURRENCY_NUM = 15
    
    crawler = XiaoHongShuCrawler()
    await crawler.start()

if __name__ == "__main__":
    asyncio.run(main())

```

## Performance and Resource Considerations

| Setting | Memory Usage | Request Rate | Risk Level | Best For |
|---------|-----------|------------|-----------|----------|
| `1` | Minimal | ~1 req/sec | Very Low | All platforms, testing, strict APIs |
| `5-10` | Moderate | 5-10 req/sec | Low | Zhihu, Bilibili, Tieba |
| `20-50` | Higher | 20-50 req/sec | Medium | Douyin, Kuaishou (with proxy rotation) |
| `>50` | Significant | Burst rates | High | Only with residential proxies, rate limit handling |

Memory scales with concurrency because each async operation may hold:
- An `httpx.AsyncClient` connection
- A Playwright browser page or context
- Response data and parsing buffers

## How Rate Limiting Interacts with Concurrency

MediaCrawler's async architecture does not include built-in exponential backoff. The `MAX_CONCURRENCY_NUM` semaphore is your primary defense against rate limits. When platforms return `429 Too Many Requests` or temporary bans, reducing this value is the first remediation step.

Some crawlers combine the semaphore with additional safeguards:

```python

# Hypothetical enhancement: per-platform rate limiting

async def fetch_with_backoff(self, url: str, max_retries: int = 3):
    async with self.semaphore:
        for attempt in range(max_retries):
            try:
                response = await self.client.get(url)
                response.raise_for_status()
                return response
            except httpx.HTTPStatusError as e:
                if e.response.status_code == 429:
                    wait = 2 ** attempt  # Exponential backoff

                    await asyncio.sleep(wait)
                else:
                    raise
        raise RateLimitExceeded(url)

```

## Key Source Files Reference

| File | Purpose | Relevant Lines |
|------|---------|--------------|
| [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) | Default `MAX_CONCURRENCY_NUM` constant | Line 105 |
| [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) | CLI argument parsing and config override | Lines 293-362 |
| [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) | Zhihu semaphore usage | Lines 216-226 |
| [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) | XHS semaphore usage | Lines 166-172 |
| [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py) | Weibo semaphore usage | Core worker methods |
| [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py) | Tieba semaphore usage | Core worker methods |
| [`media_platform/kuaishou/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/core.py) | Kuaishou semaphore usage | Core worker methods |
| [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) | Douyin semaphore usage | Core worker methods |
| [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) | Bilibili semaphore usage | Core worker methods |

## Summary

- **MediaCrawler's async architecture** uses Python `asyncio` with platform-specific crawlers for simultaneous multi-source data collection.
- **`MAX_CONCURRENCY_NUM`** in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) controls maximum simultaneous I/O operations via `asyncio.Semaphore`, defaulting to `1`.
- **CLI override** via `--max_concurrency_num` allows runtime adjustment without code changes.
- **Every platform crawler** (Zhihu, XHS, Weibo, Tieba, Kuaishou, Douyin, Bilibili) implements the identical semaphore pattern for consistent concurrency control.
- **Tuning guidance**: Start at `1`, increase gradually based on target API limits, infrastructure capacity, and observed rate limit responses.

## Frequently Asked Questions

### What happens if I set MAX_CONCURRENCY_NUM too high?

Excessive concurrency causes rate limit errors (HTTP 429), IP bans, memory exhaustion, or connection pool depletion. Start conservative and increase based on platform tolerance and infrastructure monitoring. MediaCrawler does not implement automatic backoff, so manual tuning is required.

### Can I use different concurrency limits for different platforms?

Yes. Since `MAX_CONCURRENCY_NUM` is read from `config` at crawler initialization, you can run separate processes with different CLI values: `python main.py --platform zhihu --max_concurrency_num 3` and `python main.py --platform douyin --max_concurrency_num 10` execute independently with their own limits.

### Does MAX_CONCURRENCY_NUM affect browser automation crawling?

Absolutely. The same semaphore guards Playwright browser operations in `async with semaphore:` blocks. Browser contexts and pages consume more memory than HTTP requests, so keep limits lower when using headless browser modes (typically `1-5` versus `10-50` for API-only crawling).

### Why is the default only 1 instead of a higher number?

The default of `1` prioritizes safety and portability. Different platforms have dramatically different rate limits—Zhihu is stricter than Kuaishou. A conservative default ensures first-time users don't immediately trigger bans. Production deployments should benchmark and adjust based on their specific target platforms and infrastructure.