Async Architecture and Concurrency Control (`MAX_CONCURRENCY_NUM`) in MediaCrawler
MediaCrawler uses Python's asyncio with a semaphore-based concurrency limiter (MAX_CONCURRENCY_NUM) to control how many simultaneous network or browser operations execute at once, defaulting to 1 for safe, serialized crawling.
MediaCrawler is an open-source multi-platform media crawler that asynchronously fetches data from Zhihu, XiaoHongShu (XHS), Weibo, Tieba, Kuaishou, Douyin, Bilibili, and more. Its async architecture and concurrency control through MAX_CONCURRENCY_NUM ensure efficient resource usage while respecting API rate limits. This article explains how the semaphore-based system works, where it's configured, and how to tune it for your crawling workloads.
What Is MAX_CONCURRENCY_NUM?
MAX_CONCURRENCY_NUM is a global configuration constant that defines the maximum number of concurrent coroutines allowed to perform I/O-bound operations simultaneously. These operations include HTTP requests, browser automation with Playwright, and other network-dependent tasks.
The default value of 1 means all network calls are serialized by default. This conservative setting prevents overwhelming target servers and avoids triggering rate limits or IP bans.
Where MAX_CONCURRENCY_NUM Is Defined and Configured
Base Configuration File
The constant originates in config/base_config.py at line 105:
# config/base_config.py (excerpt)
MAX_CONCURRENCY_NUM = 1 # Controls simultaneous async operations
This file serves as the single source of truth for the default concurrency limit across all platform crawlers.
CLI Argument Override
Users can override the default at runtime via the command-line interface. The argument parser in cmd_arg/arg.py (lines 293-362) exposes --max_concurrency_num:
# Run with 5 concurrent operations
python main.py --max_concurrency_num 5 --platform xhs
The CLI parser updates config.MAX_CONCURRENCY_NUM before any crawler modules initialize, ensuring the custom limit propagates throughout the system.
How the Semaphore Pattern Works in Practice
Each platform's core module creates an asyncio.Semaphore from MAX_CONCURRENCY_NUM and passes it to async workers. The semaphore is acquired using async with semaphore: before any network or browser operation begins.
Zhihu Implementation
In media_platform/zhihu/core.py (lines 216-226):
# media_platform/zhihu/core.py
import asyncio
import config
class ZhihuCrawler:
def __init__(self):
self.semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
async def fetch_comments(self, answer_id: str):
async with self.semaphore: # Concurrency guard
# HTTP request or browser automation happens here
async with httpx.AsyncClient() as client:
response = await client.get(f"/answers/{answer_id}/comments")
return response.json()
XiaoHongShu (XHS) Implementation
The same pattern appears in media_platform/xhs/core.py (lines 166-172):
# media_platform/xhs/core.py
async def fetch_note_detail(self, note_id: str):
async with self.semaphore:
# Render page or call XHS API
page = await self.browser_context.new_page()
# ... note extraction logic
Consistent Pattern Across All Platforms
The identical semaphore usage exists in:
media_platform/weibo/core.pymedia_platform/tieba/core.pymedia_platform/kuaishou/core.pymedia_platform/douyin/core.pymedia_platform/bilibili/core.py
This design ensures unified concurrency control regardless of which platform you're crawling.
Practical Usage and Tuning
Default Behavior (MAX_CONCURRENCY_NUM = 1)
With the default value, MediaCrawler processes one operation at a time. This is ideal for:
- Development and debugging
- Strictly rate-limited APIs
- Low-resource environments (small VPS, residential connections)
- Avoiding IP bans on sensitive platforms
Increasing Concurrency
Raise the limit when targeting platforms with generous rate limits or when running on capable infrastructure:
# Example: Crawl XHS with 10 concurrent operations
python main.py --platform xhs --max_concurrency_num 10
# Example: Crawl multiple platforms with different limits
python main.py --platform zhihu --max_concurrency_num 5
python main.py --platform douyin --max_concurrency_num 20
Custom Script Integration
For programmatic usage, reference the config directly:
import asyncio
import config
from media_platform.xhs import XiaoHongShuCrawler
async def main():
# Override config before instantiation
config.MAX_CONCURRENCY_NUM = 15
crawler = XiaoHongShuCrawler()
await crawler.start()
if __name__ == "__main__":
asyncio.run(main())
Performance and Resource Considerations
| Setting | Memory Usage | Request Rate | Risk Level | Best For |
|---|---|---|---|---|
1 |
Minimal | ~1 req/sec | Very Low | All platforms, testing, strict APIs |
5-10 |
Moderate | 5-10 req/sec | Low | Zhihu, Bilibili, Tieba |
20-50 |
Higher | 20-50 req/sec | Medium | Douyin, Kuaishou (with proxy rotation) |
>50 |
Significant | Burst rates | High | Only with residential proxies, rate limit handling |
Memory scales with concurrency because each async operation may hold:
- An
httpx.AsyncClientconnection - A Playwright browser page or context
- Response data and parsing buffers
How Rate Limiting Interacts with Concurrency
MediaCrawler's async architecture does not include built-in exponential backoff. The MAX_CONCURRENCY_NUM semaphore is your primary defense against rate limits. When platforms return 429 Too Many Requests or temporary bans, reducing this value is the first remediation step.
Some crawlers combine the semaphore with additional safeguards:
# Hypothetical enhancement: per-platform rate limiting
async def fetch_with_backoff(self, url: str, max_retries: int = 3):
async with self.semaphore:
for attempt in range(max_retries):
try:
response = await self.client.get(url)
response.raise_for_status()
return response
except httpx.HTTPStatusError as e:
if e.response.status_code == 429:
wait = 2 ** attempt # Exponential backoff
await asyncio.sleep(wait)
else:
raise
raise RateLimitExceeded(url)
Key Source Files Reference
| File | Purpose | Relevant Lines |
|---|---|---|
config/base_config.py |
Default MAX_CONCURRENCY_NUM constant |
Line 105 |
cmd_arg/arg.py |
CLI argument parsing and config override | Lines 293-362 |
media_platform/zhihu/core.py |
Zhihu semaphore usage | Lines 216-226 |
media_platform/xhs/core.py |
XHS semaphore usage | Lines 166-172 |
media_platform/weibo/core.py |
Weibo semaphore usage | Core worker methods |
media_platform/tieba/core.py |
Tieba semaphore usage | Core worker methods |
media_platform/kuaishou/core.py |
Kuaishou semaphore usage | Core worker methods |
media_platform/douyin/core.py |
Douyin semaphore usage | Core worker methods |
media_platform/bilibili/core.py |
Bilibili semaphore usage | Core worker methods |
Summary
- MediaCrawler's async architecture uses Python
asynciowith platform-specific crawlers for simultaneous multi-source data collection. MAX_CONCURRENCY_NUMinconfig/base_config.pycontrols maximum simultaneous I/O operations viaasyncio.Semaphore, defaulting to1.- CLI override via
--max_concurrency_numallows runtime adjustment without code changes. - Every platform crawler (Zhihu, XHS, Weibo, Tieba, Kuaishou, Douyin, Bilibili) implements the identical semaphore pattern for consistent concurrency control.
- Tuning guidance: Start at
1, increase gradually based on target API limits, infrastructure capacity, and observed rate limit responses.
Frequently Asked Questions
What happens if I set MAX_CONCURRENCY_NUM too high?
Excessive concurrency causes rate limit errors (HTTP 429), IP bans, memory exhaustion, or connection pool depletion. Start conservative and increase based on platform tolerance and infrastructure monitoring. MediaCrawler does not implement automatic backoff, so manual tuning is required.
Can I use different concurrency limits for different platforms?
Yes. Since MAX_CONCURRENCY_NUM is read from config at crawler initialization, you can run separate processes with different CLI values: python main.py --platform zhihu --max_concurrency_num 3 and python main.py --platform douyin --max_concurrency_num 10 execute independently with their own limits.
Does MAX_CONCURRENCY_NUM affect browser automation crawling?
Absolutely. The same semaphore guards Playwright browser operations in async with semaphore: blocks. Browser contexts and pages consume more memory than HTTP requests, so keep limits lower when using headless browser modes (typically 1-5 versus 10-50 for API-only crawling).
Why is the default only 1 instead of a higher number?
The default of 1 prioritizes safety and portability. Different platforms have dramatically different rate limits—Zhihu is stricter than Kuaishou. A conservative default ensures first-time users don't immediately trigger bans. Production deployments should benchmark and adjust based on their specific target platforms and infrastructure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →