How MediaCrawler Handles Async Concurrency for Parallel Crawling

MediaCrawler uses Python's asyncio.Semaphore combined with asyncio.gather to execute multiple crawling tasks concurrently while respecting a configurable global limit, preventing rate limiting and blocking.

MediaCrawler is a multi-platform social media scraping framework that leverages Python's asyncio library to maximize I/O efficiency. The repository, hosted at NanmiCoder/MediaCrawler, implements a uniform async concurrency pattern across all supported platforms including Zhihu, Xiaohongshu, Weibo, and Douyin. This architecture allows the crawler to process hundreds of posts and comments simultaneously without overwhelming target servers or triggering anti-bot measures.

The Core Concurrency Pattern

The implementation relies on five coordinated mechanisms that work together to manage parallel execution. According to the source code in media_platform/zhihu/core.py, the pattern creates a controlled execution environment where network I/O happens concurrently but never exceeds safe operational limits.

Semaphore-Based Limiting

At the entry point of each crawling operation, the code instantiates an asyncio.Semaphore using a configurable value from config.MAX_CONCURRENCY_NUM. This semaphore acts as a global gatekeeper that limits the number of simultaneous HTTP requests across the entire crawling session.

semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)   # media_platform/zhihu/core.py#L16

Task Creation and Execution

Once the semaphore is established, the crawler generates individual tasks for each content item using asyncio.create_task. These tasks are collected into a list and executed simultaneously using asyncio.gather, which runs all coroutines concurrently while the event loop remains unblocked.

task_list = [
    asyncio.create_task(self.get_comments(item, semaphore), name=item.content_id)
    for item in content_list
]

await asyncio.gather(*task_list)   # media_platform/zhihu/core.py#L23

Worker Protection with Context Managers

Each worker coroutine (get_comments, get_note_detail, etc.) wraps its network operations inside an async with semaphore: block. This ensures that the global concurrency limit is respected even when individual coroutines are awaiting network I/O, preventing the crawler from spawning unlimited simultaneous connections.

async with semaphore:   # media_platform/zhihu/core.py#L37

    await self.zhihu_client.get_note_all_comments(
        content=content_item,
        crawl_interval=config.CRAWLER_MAX_SLEEP_SEC,
        callback=zhihu_store.batch_update_zhihu_note_comments,
    )

Step-by-Step Implementation Flow

The async concurrency model follows a strict execution sequence that appears consistently across platform implementations:

  1. Initialize the semaphore – Creates the concurrency cap based on config.MAX_CONCURRENCY_NUM (typically set to 8 by default in config/base_config.py).

  2. Spawn async tasks – For every content item (post, comment, video), the crawler creates a task using asyncio.create_task with the semaphore passed as an argument.

  3. Execute concurrently – await asyncio.gather(*tasks) schedules all tasks on the event loop simultaneously, allowing Python to switch between them during I/O wait times.

  4. Throttle requests – After each network operation, await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) inserts a configurable pause (typically 1-2 seconds) to avoid overwhelming target servers.

  5. Enforce limits per worker – Each task uses async with semaphore: to acquire permission before making HTTP requests, ensuring the global limit is never exceeded regardless of how many tasks are in the gathering phase.

Configuration and Rate Limiting

The concurrency behavior is controlled through config/base_config.py, which defines two critical parameters:


# Maximum number of concurrent HTTP requests across all platforms

MAX_CONCURRENCY_NUM = 8

# Fixed sleep between requests to avoid triggering anti-scraping measures

CRAWLER_MAX_SLEEP_SEC = 2

These settings propagate to all platform-specific crawlers, ensuring uniform behavior whether scraping Zhihu, Bilibili, or Kuaishou. The sleep interval works in conjunction with the semaphore—while the semaphore controls parallel execution, the sleep timer ensures sequential politeness between requests.

Cross-Platform Consistency

The same async concurrency pattern appears identically across all platform modules. In media_platform/tieba/core.py, the implementation mirrors the Zhihu approach:

semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
tasks = [
    asyncio.create_task(self.get_note_detail(url, semaphore))
    for url in note_urls
]
await asyncio.gather(*tasks)

Similarly, media_platform/xhs/core.py and media_platform/weibo/core.py implement identical semaphore-gather patterns, making the codebase predictable and maintainable. The abstract base class defined in base/base_crawler.py establishes the interface contracts that enforce this consistency.

Practical Implementation Example

Here is a complete example showing how the Zhihu crawler implements parallel comment fetching:

async def fetch_comments_parallel(self, content_list):
    # Limit parallelism

    semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
    
    # Create one task per content item

    task_list = [
        asyncio.create_task(
            self.get_comments(item, semaphore), 
            name=item.content_id
        )
        for item in content_list
    ]
    
    # Run all tasks concurrently

    await asyncio.gather(*task_list)

async def get_comments(self, content_item, semaphore):
    async with semaphore:  # Respects global limit

        await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
        await self.zhihu_client.get_note_all_comments(
            content=content_item,
            crawl_interval=config.CRAWLER_MAX_SLEEP_SEC,
            callback=zhihu_store.batch_update_zhihu_note_comments,
        )

This pattern ensures that if MAX_CONCURRENCY_NUM is set to 8, only 8 HTTP requests will be in flight at any moment, even if the task list contains 1000 items.

Summary

  • Semaphore-based limiting: asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) creates a global concurrency cap that prevents overwhelming target servers.
  • Task gathering: asyncio.gather(*tasks) executes all crawling operations concurrently while respecting the semaphore constraints.
  • Worker protection: Each coroutine uses async with semaphore: to acquire permission before network I/O, ensuring limits are enforced at the operation level.
  • Configurable throttling: CRAWLER_MAX_SLEEP_SEC adds polite delays between requests to avoid rate limiting.
  • Unified architecture: The same pattern appears in media_platform/zhihu/core.py, media_platform/xhs/core.py, media_platform/tieba/core.py, and other platform modules.

Frequently Asked Questions

What is the maximum number of concurrent requests in MediaCrawler?

The default maximum is 8 concurrent requests, defined by MAX_CONCURRENCY_NUM in config/base_config.py. You can adjust this value based on your network capacity and the target platform's rate limits, though higher values increase the risk of IP bans or CAPTCHA challenges.

How does MediaCrawler prevent rate limiting?

MediaCrawler implements two-layer protection: first, the asyncio.Semaphore caps simultaneous connections; second, asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) inserts a configurable delay (typically 1-2 seconds) after each request. This combination ensures the crawler behaves politely and avoids triggering anti-bot mechanisms.

Why use asyncio.Semaphore instead of asyncio.Queue?

The semaphore pattern is more efficient for I/O-bound crawling because it allows the event loop to manage context switching naturally during network waits. While Queue could limit concurrency, the semaphore approach in media_platform/zhihu/core.py provides cleaner syntax with async with context managers and integrates seamlessly with asyncio.gather for batch task execution.

Is the concurrency pattern identical across all platforms?

Yes. Whether crawling Zhihu, Xiaohongshu, Weibo, or Bilibili, the implementation uses the same five-step pattern: semaphore creation, task spawning with create_task, concurrent execution via gather, semaphore-protected workers, and configurable sleep intervals. This uniformity is enforced by the abstract interfaces in base/base_crawler.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →