MediaCrawler Performance Considerations: Optimizing Throughput and Resource Usage
MediaCrawler performance depends on balancing async concurrency limits, request politeness delays, browser launch modes, and storage backend I/O efficiency to maximize throughput while avoiding detection.
MediaCrawler is a high-concurrency asynchronous web scraper targeting short-video and social platforms including Xiaohongshu, Douyin, Bilibili, and Zhihu. Optimizing its performance requires tuning architectural knobs in config/base_config.py and understanding how semaphore-based rate limiting, browser automation overhead, and persistence strategies interact. This guide examines the specific implementation details in the NanmiCoder/MediaCrawler repository that determine crawling speed, latency, and resource consumption.
Async Concurrency and Semaphore-Based Rate Limiting
All platform crawlers inherit from AbstractCrawler in base/base_crawler.py and run inside an asyncio event loop using Playwright’s async API. To prevent overwhelming target sites, each implementation (such as media_platform/xhs/core.py) creates an asyncio.Semaphore(config.MAX_CONCURRENCY_NUM) that wraps network-heavy operations.
# From media_platform/xhs/core.py (line 161)
semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
async with semaphore:
await self.get_comments(note_id)
Raising MAX_CONCURRENCY_NUM increases raw throughput but consumes more CPU and memory. Conversely, lowering it reduces resource usage but creates bottlenecks during I/O wait times.
Request Politeness and Sleep Intervals
After each request, crawlers invoke await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC) to avoid detection. In media_platform/zhihu/core.py (line 189), this fixed delay defaults to 2 seconds. Decreasing this value speeds up crawling but increases the risk of IP bans or rate-limiting.
Task Batching Memory Impact
Crawlers batch tasks using asyncio.gather() as shown in media_platform/douyin/core.py (line 209):
task_list = [self.get_note_info_task(id) for id in note_ids]
await asyncio.gather(*task_list)
Large batches increase concurrency but also memory pressure, as each asyncio.Task holds references to its entire call stack. For crawls exceeding 10,000 notes, splitting batches keeps RAM usage manageable.
Tuning Concurrency Parameters
Adjust these values in config/base_config.py based on hardware capacity:
MAX_CONCURRENCY_NUM = 8 # Parallel network calls (6-8 for 8-core/16GB)
CRAWLER_MAX_SLEEP_SEC = 0.5 # Reduced delay (monitor for bans)
For an 8-core machine with 16GB RAM, MAX_CONCURRENCY_NUM of 6–8 provides optimal throughput without excessive context switching.
Browser Launch Modes: CDP vs Standard Playwright
MediaCrawler supports two browser initialization strategies that significantly impact startup latency and resource isolation.
CDP (Chrome DevTools Protocol) Mode (tools/cdp_browser.py) connects to an existing Chrome instance via remote debugging (config.CDP_CONNECT_EXISTING=True). This eliminates launch overhead and leverages cached sessions but requires a running Chrome browser with debugging enabled.
Standard Playwright Mode (base/base_crawler.py, line 54) launches a fresh Chromium process (config.ENABLE_CDP_MODE=False). While this adds seconds to startup and increases memory footprint, it provides complete isolation from user profiles and extensions, making it safer for CI/CD pipelines.
# config/base_config.py
ENABLE_CDP_MODE = False # Use fresh browser instance
HEADLESS = True # Disable UI for server environments
BROWSER_LAUNCH_TIMEOUT = 30
Storage Backend I/O Performance
The SAVE_DATA_OPTION in config/base_config.py determines how crawled data persists. The implementation resides in platform-specific store classes (e.g., store/xhs/_store_impl.py), where AsyncFileWriter handles writes.
| Backend | Throughput | Memory Profile | Best For |
|---|---|---|---|
| JSONL | ~10k records/sec | Low (streamed) | Large dumps, easy import |
| SQLite | ~5k inserts/sec | Moderate | Local relational queries |
| MySQL/Postgres | >20k inserts/sec | Higher | Production deduplication |
| MongoDB | ~15k docs/sec | Moderate-High | Schema-flexible pipelines |
| CSV | Slow | High (header overhead) | Small datasets only |
Specifically, XhsJsonStoreImplement uses AsyncFileWriter.write_single_item_to_json (from store/xhs/_store_impl.py, line 71) for batched JSON writes. Avoid CSV for crawls exceeding 1 million records, as csv.DictWriter retains the entire header structure in memory for each write operation.
Caching Strategy: Redis vs Local Expiring Cache
MediaCrawler provides two caching implementations to reduce redundant network calls and manage proxy tokens.
RedisCache (cache/redis_cache.py) offers ~0.5ms latency and supports shared state across distributed crawler instances. Use this when running multiple workers or when proxy tokens (IP_PROXY_PROVIDER_NAME) must persist between runs.
LocalExpiringCache (cache/local_cache.py) provides ~0.05ms access latency for single-process deduplication of URLs and temporary session storage without network overhead.
When ENABLE_IP_PROXY=True, the ProxyIPPool class (proxy/proxy_ip_pool.py) caches live proxies; storing this in Redis prevents re-fetching from the provider on every execution.
Proxy Pool Configuration
Concurrent proxy connections are controlled by IP_PROXY_POOL_COUNT (default: 2) in config/base_config.py. Each crawler thread acquires proxies from ProxyIPPool before requests:
IP_PROXY_POOL_COUNT = 4 # Parallel proxy sessions
Over-provisioning proxies reduces per-request latency but increases the risk of proxy-side throttling or IP bans from the proxy provider.
Memory and CPU Optimization
High concurrency creates significant memory pressure. Each asyncio.Task object retains references to its call stack; with MAX_CONCURRENCY_NUM=10 and 10,000 pending notes, resident memory can exceed 1GB.
The word-cloud generator (tools/words.py) adds CPU-intensive processing when ENABLE_GET_WORDCLOUD=True, scaling linearly with comment volume. For crawls exceeding 100,000 comments, disable this feature or offload processing to a separate worker process.
File writers use a per-platform asyncio.Lock (in AsyncFileWriter) that serializes writes but remains lightweight; I/O costs are dominated by underlying disk speed rather than the locking mechanism itself.
Practical Configuration Examples
High-Throughput Tuning
# config/base_config.py
MAX_CONCURRENCY_NUM = 12
CRAWLER_MAX_SLEEP_SEC = 0.2
SAVE_DATA_OPTION = "jsonl" # Fastest local storage
ENABLE_GET_WORDCLOUD = False # Reduce CPU load
# Run command
uv run main.py --platform xhs --lt qrcode --type search
Database Persistence for Deduplication
# config/base_config.py
SAVE_DATA_OPTION = "sqlite"
MAX_CONCURRENCY_NUM = 4 # Lower concurrency prevents DB lock contention
# Stores data via store/xhs/sqlite_store.py
Distributed Caching with Redis
# .env
REDIS_DB_HOST=127.0.0.1
REDIS_DB_PORT=6379
REDIS_DB_NUM=0
# config/base_config.py
ENABLE_IP_PROXY = True
IP_PROXY_PROVIDER_NAME = "kuaidaili"
Summary
- Concurrency tuning: Adjust
MAX_CONCURRENCY_NUMandCRAWLER_MAX_SLEEP_SECinconfig/base_config.pyto balance throughput against politeness; values of 6–8 work well for mid-range hardware. - Browser mode selection: Use CDP mode (
tools/cdp_browser.py) for low-latency development, and standard Playwright mode withHEADLESS=Truefor isolated server deployments. - Storage optimization: Prefer JSONL or database backends (MySQL/Postgres) over CSV for large-scale crawls to minimize memory overhead and maximize insert throughput.
- Caching architecture: Deploy
RedisCachefor multi-instance coordination and proxy token persistence; useLocalExpiringCachefor single-process URL deduplication. - Resource management: Monitor memory consumption when batching thousands of tasks simultaneously, and disable word-cloud generation (
ENABLE_GET_WORDCLOUD) for high-volume comment extraction.
Frequently Asked Questions
How does MAX_CONCURRENCY_NUM affect MediaCrawler performance?
This parameter controls the asyncio.Semaphore limit that governs how many simultaneous network requests (Playwright page loads or API calls) can execute. Increasing it improves throughput up to the point of CPU saturation or target site rate-limiting, while decreasing it reduces memory usage and detection risk but slows overall crawling speed.
What is the fastest storage backend for large-scale crawls?
JSONL provides the highest local throughput (~10,000 records/sec) with minimal memory footprint due to streaming writes. For production environments requiring complex queries or deduplication, MySQL or PostgreSQL exceeds 20,000 inserts per second when connection pools are tuned properly, significantly outperforming CSV and SQLite for concurrent writes.
Should I use CDP mode or standard Playwright for production?
Use CDP mode (ENABLE_CDP_MODE=True) when developing locally or debugging, as it connects to an existing Chrome instance and eliminates browser launch latency. Use standard Playwright mode (ENABLE_CDP_MODE=False with HEADLESS=True) for CI/CD pipelines and server deployments, as it provides process isolation and avoids dependencies on GUI availability.
Why does MediaCrawler consume excessive memory during large crawls?
High memory usage typically stems from large asyncio.gather() batches where thousands of Task objects retain call stack references simultaneously. Reduce MAX_CONCURRENCY_NUM or implement smaller batch sizes in platform core files (e.g., media_platform/douyin/core.py) to allow garbage collection between groups. Additionally, ensure ENABLE_GET_WORDCLOUD=False to prevent in-memory text processing of large comment datasets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →