# How MediaCrawler Handles Async Concurrency for Parallel Crawling

> Discover how MediaCrawler manages async concurrency for parallel crawling using asyncio Semaphore and gather. Optimize your crawls and avoid rate limiting.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: internals
- Published: 2026-07-02

---

**MediaCrawler uses Python's `asyncio.Semaphore` combined with `asyncio.gather` to execute multiple crawling tasks concurrently while respecting a configurable global limit, preventing rate limiting and blocking.**

MediaCrawler is a multi-platform social media scraping framework that leverages Python's `asyncio` library to maximize I/O efficiency. The repository, hosted at `NanmiCoder/MediaCrawler`, implements a uniform async concurrency pattern across all supported platforms including Zhihu, Xiaohongshu, Weibo, and Douyin. This architecture allows the crawler to process hundreds of posts and comments simultaneously without overwhelming target servers or triggering anti-bot measures.

## The Core Concurrency Pattern

The implementation relies on five coordinated mechanisms that work together to manage parallel execution. According to the source code in [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py), the pattern creates a controlled execution environment where network I/O happens concurrently but never exceeds safe operational limits.

### Semaphore-Based Limiting

At the entry point of each crawling operation, the code instantiates an `asyncio.Semaphore` using a configurable value from `config.MAX_CONCURRENCY_NUM`. This semaphore acts as a global gatekeeper that limits the number of simultaneous HTTP requests across the entire crawling session.

```python
semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)   # media_platform/zhihu/core.py#L16

```

### Task Creation and Execution

Once the semaphore is established, the crawler generates individual tasks for each content item using `asyncio.create_task`. These tasks are collected into a list and executed simultaneously using `asyncio.gather`, which runs all coroutines concurrently while the event loop remains unblocked.

```python
task_list = [
    asyncio.create_task(self.get_comments(item, semaphore), name=item.content_id)
    for item in content_list
]

await asyncio.gather(*task_list)   # media_platform/zhihu/core.py#L23

```

### Worker Protection with Context Managers

Each worker coroutine (`get_comments`, `get_note_detail`, etc.) wraps its network operations inside an `async with semaphore:` block. This ensures that the global concurrency limit is respected even when individual coroutines are awaiting network I/O, preventing the crawler from spawning unlimited simultaneous connections.

```python
async with semaphore:   # media_platform/zhihu/core.py#L37

    await self.zhihu_client.get_note_all_comments(
        content=content_item,
        crawl_interval=config.CRAWLER_MAX_SLEEP_SEC,
        callback=zhihu_store.batch_update_zhihu_note_comments,
    )

```

## Step-by-Step Implementation Flow

The async concurrency model follows a strict execution sequence that appears consistently across platform implementations:

1. **Initialize the semaphore** – Creates the concurrency cap based on `config.MAX_CONCURRENCY_NUM` (typically set to 8 by default in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)).

2. **Spawn async tasks** – For every content item (post, comment, video), the crawler creates a task using `asyncio.create_task` with the semaphore passed as an argument.

3. **Execute concurrently** – `await asyncio.gather(*tasks)` schedules all tasks on the event loop simultaneously, allowing Python to switch between them during I/O wait times.

4. **Throttle requests** – After each network operation, `await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` inserts a configurable pause (typically 1-2 seconds) to avoid overwhelming target servers.

5. **Enforce limits per worker** – Each task uses `async with semaphore:` to acquire permission before making HTTP requests, ensuring the global limit is never exceeded regardless of how many tasks are in the gathering phase.

## Configuration and Rate Limiting

The concurrency behavior is controlled through [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), which defines two critical parameters:

```python

# Maximum number of concurrent HTTP requests across all platforms

MAX_CONCURRENCY_NUM = 8

# Fixed sleep between requests to avoid triggering anti-scraping measures

CRAWLER_MAX_SLEEP_SEC = 2

```

These settings propagate to all platform-specific crawlers, ensuring uniform behavior whether scraping Zhihu, Bilibili, or Kuaishou. The sleep interval works in conjunction with the semaphore—while the semaphore controls parallel execution, the sleep timer ensures sequential politeness between requests.

## Cross-Platform Consistency

The same async concurrency pattern appears identically across all platform modules. In [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py), the implementation mirrors the Zhihu approach:

```python
semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
tasks = [
    asyncio.create_task(self.get_note_detail(url, semaphore))
    for url in note_urls
]
await asyncio.gather(*tasks)

```

Similarly, [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) and [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py) implement identical semaphore-gather patterns, making the codebase predictable and maintainable. The abstract base class defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) establishes the interface contracts that enforce this consistency.

## Practical Implementation Example

Here is a complete example showing how the Zhihu crawler implements parallel comment fetching:

```python
async def fetch_comments_parallel(self, content_list):
    # Limit parallelism

    semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
    
    # Create one task per content item

    task_list = [
        asyncio.create_task(
            self.get_comments(item, semaphore), 
            name=item.content_id
        )
        for item in content_list
    ]
    
    # Run all tasks concurrently

    await asyncio.gather(*task_list)

async def get_comments(self, content_item, semaphore):
    async with semaphore:  # Respects global limit

        await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
        await self.zhihu_client.get_note_all_comments(
            content=content_item,
            crawl_interval=config.CRAWLER_MAX_SLEEP_SEC,
            callback=zhihu_store.batch_update_zhihu_note_comments,
        )

```

This pattern ensures that if `MAX_CONCURRENCY_NUM` is set to 8, only 8 HTTP requests will be in flight at any moment, even if the task list contains 1000 items.

## Summary

- **Semaphore-based limiting**: `asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)` creates a global concurrency cap that prevents overwhelming target servers.
- **Task gathering**: `asyncio.gather(*tasks)` executes all crawling operations concurrently while respecting the semaphore constraints.
- **Worker protection**: Each coroutine uses `async with semaphore:` to acquire permission before network I/O, ensuring limits are enforced at the operation level.
- **Configurable throttling**: `CRAWLER_MAX_SLEEP_SEC` adds polite delays between requests to avoid rate limiting.
- **Unified architecture**: The same pattern appears in [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py), [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py), [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py), and other platform modules.

## Frequently Asked Questions

### What is the maximum number of concurrent requests in MediaCrawler?

The default maximum is **8 concurrent requests**, defined by `MAX_CONCURRENCY_NUM` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). You can adjust this value based on your network capacity and the target platform's rate limits, though higher values increase the risk of IP bans or CAPTCHA challenges.

### How does MediaCrawler prevent rate limiting?

MediaCrawler implements **two-layer protection**: first, the `asyncio.Semaphore` caps simultaneous connections; second, `asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` inserts a configurable delay (typically 1-2 seconds) after each request. This combination ensures the crawler behaves politely and avoids triggering anti-bot mechanisms.

### Why use `asyncio.Semaphore` instead of `asyncio.Queue`?

The semaphore pattern is more efficient for I/O-bound crawling because it allows the event loop to manage context switching naturally during network waits. While `Queue` could limit concurrency, the semaphore approach in [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) provides cleaner syntax with `async with` context managers and integrates seamlessly with `asyncio.gather` for batch task execution.

### Is the concurrency pattern identical across all platforms?

Yes. Whether crawling Zhihu, Xiaohongshu, Weibo, or Bilibili, the implementation uses the same five-step pattern: semaphore creation, task spawning with `create_task`, concurrent execution via `gather`, semaphore-protected workers, and configurable sleep intervals. This uniformity is enforced by the abstract interfaces in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py).