How MediaCrawler Handles Async Concurrency for Parallel Crawling
MediaCrawler uses Python's asyncio.Semaphore combined with asyncio.gather to execute multiple crawling tasks concurrently while respecting a configurable global limit, preventing rate limiting and blocking.
MediaCrawler is a multi-platform social media scraping framework that leverages Python's asyncio library to maximize I/O efficiency. The repository, hosted at NanmiCoder/MediaCrawler, implements a uniform async concurrency pattern across all supported platforms including Zhihu, Xiaohongshu, Weibo, and Douyin. This architecture allows the crawler to process hundreds of posts and comments simultaneously without overwhelming target servers or triggering anti-bot measures.
The Core Concurrency Pattern
The implementation relies on five coordinated mechanisms that work together to manage parallel execution. According to the source code in media_platform/zhihu/core.py, the pattern creates a controlled execution environment where network I/O happens concurrently but never exceeds safe operational limits.
Semaphore-Based Limiting
At the entry point of each crawling operation, the code instantiates an asyncio.Semaphore using a configurable value from config.MAX_CONCURRENCY_NUM. This semaphore acts as a global gatekeeper that limits the number of simultaneous HTTP requests across the entire crawling session.
semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) # media_platform/zhihu/core.py#L16
Task Creation and Execution
Once the semaphore is established, the crawler generates individual tasks for each content item using asyncio.create_task. These tasks are collected into a list and executed simultaneously using asyncio.gather, which runs all coroutines concurrently while the event loop remains unblocked.
task_list = [
asyncio.create_task(self.get_comments(item, semaphore), name=item.content_id)
for item in content_list
]
await asyncio.gather(*task_list) # media_platform/zhihu/core.py#L23
Worker Protection with Context Managers
Each worker coroutine (get_comments, get_note_detail, etc.) wraps its network operations inside an async with semaphore: block. This ensures that the global concurrency limit is respected even when individual coroutines are awaiting network I/O, preventing the crawler from spawning unlimited simultaneous connections.
async with semaphore: # media_platform/zhihu/core.py#L37
await self.zhihu_client.get_note_all_comments(
content=content_item,
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC,
callback=zhihu_store.batch_update_zhihu_note_comments,
)
Step-by-Step Implementation Flow
The async concurrency model follows a strict execution sequence that appears consistently across platform implementations:
-
Initialize the semaphore – Creates the concurrency cap based on
config.MAX_CONCURRENCY_NUM(typically set to 8 by default inconfig/base_config.py). -
Spawn async tasks – For every content item (post, comment, video), the crawler creates a task using
asyncio.create_taskwith the semaphore passed as an argument. -
Execute concurrently –
await asyncio.gather(*tasks)schedules all tasks on the event loop simultaneously, allowing Python to switch between them during I/O wait times. -
Throttle requests – After each network operation,
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)inserts a configurable pause (typically 1-2 seconds) to avoid overwhelming target servers. -
Enforce limits per worker – Each task uses
async with semaphore:to acquire permission before making HTTP requests, ensuring the global limit is never exceeded regardless of how many tasks are in the gathering phase.
Configuration and Rate Limiting
The concurrency behavior is controlled through config/base_config.py, which defines two critical parameters:
# Maximum number of concurrent HTTP requests across all platforms
MAX_CONCURRENCY_NUM = 8
# Fixed sleep between requests to avoid triggering anti-scraping measures
CRAWLER_MAX_SLEEP_SEC = 2
These settings propagate to all platform-specific crawlers, ensuring uniform behavior whether scraping Zhihu, Bilibili, or Kuaishou. The sleep interval works in conjunction with the semaphore—while the semaphore controls parallel execution, the sleep timer ensures sequential politeness between requests.
Cross-Platform Consistency
The same async concurrency pattern appears identically across all platform modules. In media_platform/tieba/core.py, the implementation mirrors the Zhihu approach:
semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
tasks = [
asyncio.create_task(self.get_note_detail(url, semaphore))
for url in note_urls
]
await asyncio.gather(*tasks)
Similarly, media_platform/xhs/core.py and media_platform/weibo/core.py implement identical semaphore-gather patterns, making the codebase predictable and maintainable. The abstract base class defined in base/base_crawler.py establishes the interface contracts that enforce this consistency.
Practical Implementation Example
Here is a complete example showing how the Zhihu crawler implements parallel comment fetching:
async def fetch_comments_parallel(self, content_list):
# Limit parallelism
semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
# Create one task per content item
task_list = [
asyncio.create_task(
self.get_comments(item, semaphore),
name=item.content_id
)
for item in content_list
]
# Run all tasks concurrently
await asyncio.gather(*task_list)
async def get_comments(self, content_item, semaphore):
async with semaphore: # Respects global limit
await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)
await self.zhihu_client.get_note_all_comments(
content=content_item,
crawl_interval=config.CRAWLER_MAX_SLEEP_SEC,
callback=zhihu_store.batch_update_zhihu_note_comments,
)
This pattern ensures that if MAX_CONCURRENCY_NUM is set to 8, only 8 HTTP requests will be in flight at any moment, even if the task list contains 1000 items.
Summary
- Semaphore-based limiting:
asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)creates a global concurrency cap that prevents overwhelming target servers. - Task gathering:
asyncio.gather(*tasks)executes all crawling operations concurrently while respecting the semaphore constraints. - Worker protection: Each coroutine uses
async with semaphore:to acquire permission before network I/O, ensuring limits are enforced at the operation level. - Configurable throttling:
CRAWLER_MAX_SLEEP_SECadds polite delays between requests to avoid rate limiting. - Unified architecture: The same pattern appears in
media_platform/zhihu/core.py,media_platform/xhs/core.py,media_platform/tieba/core.py, and other platform modules.
Frequently Asked Questions
What is the maximum number of concurrent requests in MediaCrawler?
The default maximum is 8 concurrent requests, defined by MAX_CONCURRENCY_NUM in config/base_config.py. You can adjust this value based on your network capacity and the target platform's rate limits, though higher values increase the risk of IP bans or CAPTCHA challenges.
How does MediaCrawler prevent rate limiting?
MediaCrawler implements two-layer protection: first, the asyncio.Semaphore caps simultaneous connections; second, asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) inserts a configurable delay (typically 1-2 seconds) after each request. This combination ensures the crawler behaves politely and avoids triggering anti-bot mechanisms.
Why use asyncio.Semaphore instead of asyncio.Queue?
The semaphore pattern is more efficient for I/O-bound crawling because it allows the event loop to manage context switching naturally during network waits. While Queue could limit concurrency, the semaphore approach in media_platform/zhihu/core.py provides cleaner syntax with async with context managers and integrates seamlessly with asyncio.gather for batch task execution.
Is the concurrency pattern identical across all platforms?
Yes. Whether crawling Zhihu, Xiaohongshu, Weibo, or Bilibili, the implementation uses the same five-step pattern: semaphore creation, task spawning with create_task, concurrent execution via gather, semaphore-protected workers, and configurable sleep intervals. This uniformity is enforced by the abstract interfaces in base/base_crawler.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →