Async Architecture and Concurrency Control (`MAX_CONCURRENCY_NUM`) in MediaCrawler

MediaCrawler uses Python's asyncio with a semaphore-based concurrency limiter (MAX_CONCURRENCY_NUM) to control how many simultaneous network or browser operations execute at once, defaulting to 1 for safe, serialized crawling.

MediaCrawler is an open-source multi-platform media crawler that asynchronously fetches data from Zhihu, XiaoHongShu (XHS), Weibo, Tieba, Kuaishou, Douyin, Bilibili, and more. Its async architecture and concurrency control through MAX_CONCURRENCY_NUM ensure efficient resource usage while respecting API rate limits. This article explains how the semaphore-based system works, where it's configured, and how to tune it for your crawling workloads.

What Is MAX_CONCURRENCY_NUM?

MAX_CONCURRENCY_NUM is a global configuration constant that defines the maximum number of concurrent coroutines allowed to perform I/O-bound operations simultaneously. These operations include HTTP requests, browser automation with Playwright, and other network-dependent tasks.

The default value of 1 means all network calls are serialized by default. This conservative setting prevents overwhelming target servers and avoids triggering rate limits or IP bans.

Where MAX_CONCURRENCY_NUM Is Defined and Configured

Base Configuration File

The constant originates in config/base_config.py at line 105:


# config/base_config.py (excerpt)

MAX_CONCURRENCY_NUM = 1  # Controls simultaneous async operations

This file serves as the single source of truth for the default concurrency limit across all platform crawlers.

CLI Argument Override

Users can override the default at runtime via the command-line interface. The argument parser in cmd_arg/arg.py (lines 293-362) exposes --max_concurrency_num:


# Run with 5 concurrent operations

python main.py --max_concurrency_num 5 --platform xhs

The CLI parser updates config.MAX_CONCURRENCY_NUM before any crawler modules initialize, ensuring the custom limit propagates throughout the system.

How the Semaphore Pattern Works in Practice

Each platform's core module creates an asyncio.Semaphore from MAX_CONCURRENCY_NUM and passes it to async workers. The semaphore is acquired using async with semaphore: before any network or browser operation begins.

Zhihu Implementation

In media_platform/zhihu/core.py (lines 216-226):


# media_platform/zhihu/core.py

import asyncio
import config

class ZhihuCrawler:
    def __init__(self):
        self.semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
    
    async def fetch_comments(self, answer_id: str):
        async with self.semaphore:  # Concurrency guard

            # HTTP request or browser automation happens here

            async with httpx.AsyncClient() as client:
                response = await client.get(f"/answers/{answer_id}/comments")
                return response.json()

XiaoHongShu (XHS) Implementation

The same pattern appears in media_platform/xhs/core.py (lines 166-172):


# media_platform/xhs/core.py

async def fetch_note_detail(self, note_id: str):
    async with self.semaphore:
        # Render page or call XHS API

        page = await self.browser_context.new_page()
        # ... note extraction logic

Consistent Pattern Across All Platforms

The identical semaphore usage exists in:

This design ensures unified concurrency control regardless of which platform you're crawling.

Practical Usage and Tuning

Default Behavior (MAX_CONCURRENCY_NUM = 1)

With the default value, MediaCrawler processes one operation at a time. This is ideal for:

  • Development and debugging
  • Strictly rate-limited APIs
  • Low-resource environments (small VPS, residential connections)
  • Avoiding IP bans on sensitive platforms

Increasing Concurrency

Raise the limit when targeting platforms with generous rate limits or when running on capable infrastructure:


# Example: Crawl XHS with 10 concurrent operations

python main.py --platform xhs --max_concurrency_num 10

# Example: Crawl multiple platforms with different limits

python main.py --platform zhihu --max_concurrency_num 5
python main.py --platform douyin --max_concurrency_num 20

Custom Script Integration

For programmatic usage, reference the config directly:

import asyncio
import config
from media_platform.xhs import XiaoHongShuCrawler

async def main():
    # Override config before instantiation

    config.MAX_CONCURRENCY_NUM = 15
    
    crawler = XiaoHongShuCrawler()
    await crawler.start()

if __name__ == "__main__":
    asyncio.run(main())

Performance and Resource Considerations

Setting Memory Usage Request Rate Risk Level Best For
1 Minimal ~1 req/sec Very Low All platforms, testing, strict APIs
5-10 Moderate 5-10 req/sec Low Zhihu, Bilibili, Tieba
20-50 Higher 20-50 req/sec Medium Douyin, Kuaishou (with proxy rotation)
>50 Significant Burst rates High Only with residential proxies, rate limit handling

Memory scales with concurrency because each async operation may hold:

  • An httpx.AsyncClient connection
  • A Playwright browser page or context
  • Response data and parsing buffers

How Rate Limiting Interacts with Concurrency

MediaCrawler's async architecture does not include built-in exponential backoff. The MAX_CONCURRENCY_NUM semaphore is your primary defense against rate limits. When platforms return 429 Too Many Requests or temporary bans, reducing this value is the first remediation step.

Some crawlers combine the semaphore with additional safeguards:


# Hypothetical enhancement: per-platform rate limiting

async def fetch_with_backoff(self, url: str, max_retries: int = 3):
    async with self.semaphore:
        for attempt in range(max_retries):
            try:
                response = await self.client.get(url)
                response.raise_for_status()
                return response
            except httpx.HTTPStatusError as e:
                if e.response.status_code == 429:
                    wait = 2 ** attempt  # Exponential backoff

                    await asyncio.sleep(wait)
                else:
                    raise
        raise RateLimitExceeded(url)

Key Source Files Reference

File Purpose Relevant Lines
config/base_config.py Default MAX_CONCURRENCY_NUM constant Line 105
cmd_arg/arg.py CLI argument parsing and config override Lines 293-362
media_platform/zhihu/core.py Zhihu semaphore usage Lines 216-226
media_platform/xhs/core.py XHS semaphore usage Lines 166-172
media_platform/weibo/core.py Weibo semaphore usage Core worker methods
media_platform/tieba/core.py Tieba semaphore usage Core worker methods
media_platform/kuaishou/core.py Kuaishou semaphore usage Core worker methods
media_platform/douyin/core.py Douyin semaphore usage Core worker methods
media_platform/bilibili/core.py Bilibili semaphore usage Core worker methods

Summary

  • MediaCrawler's async architecture uses Python asyncio with platform-specific crawlers for simultaneous multi-source data collection.
  • MAX_CONCURRENCY_NUM in config/base_config.py controls maximum simultaneous I/O operations via asyncio.Semaphore, defaulting to 1.
  • CLI override via --max_concurrency_num allows runtime adjustment without code changes.
  • Every platform crawler (Zhihu, XHS, Weibo, Tieba, Kuaishou, Douyin, Bilibili) implements the identical semaphore pattern for consistent concurrency control.
  • Tuning guidance: Start at 1, increase gradually based on target API limits, infrastructure capacity, and observed rate limit responses.

Frequently Asked Questions

What happens if I set MAX_CONCURRENCY_NUM too high?

Excessive concurrency causes rate limit errors (HTTP 429), IP bans, memory exhaustion, or connection pool depletion. Start conservative and increase based on platform tolerance and infrastructure monitoring. MediaCrawler does not implement automatic backoff, so manual tuning is required.

Can I use different concurrency limits for different platforms?

Yes. Since MAX_CONCURRENCY_NUM is read from config at crawler initialization, you can run separate processes with different CLI values: python main.py --platform zhihu --max_concurrency_num 3 and python main.py --platform douyin --max_concurrency_num 10 execute independently with their own limits.

Does MAX_CONCURRENCY_NUM affect browser automation crawling?

Absolutely. The same semaphore guards Playwright browser operations in async with semaphore: blocks. Browser contexts and pages consume more memory than HTTP requests, so keep limits lower when using headless browser modes (typically 1-5 versus 10-50 for API-only crawling).

Why is the default only 1 instead of a higher number?

The default of 1 prioritizes safety and portability. Different platforms have dramatically different rate limits—Zhihu is stricter than Kuaishou. A conservative default ensures first-time users don't immediately trigger bans. Production deployments should benchmark and adjust based on their specific target platforms and infrastructure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →