# MediaCrawler Performance Considerations: Optimizing Throughput and Resource Usage

> Optimize MediaCrawler performance by balancing concurrency, politeness, browser modes, and I/O for maximum throughput and minimal resource usage. Learn how to enhance your crawling now.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: performance
- Published: 2026-07-29

---

**MediaCrawler performance depends on balancing async concurrency limits, request politeness delays, browser launch modes, and storage backend I/O efficiency to maximize throughput while avoiding detection.**

MediaCrawler is a high-concurrency asynchronous web scraper targeting short-video and social platforms including Xiaohongshu, Douyin, Bilibili, and Zhihu. Optimizing its performance requires tuning architectural knobs in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and understanding how semaphore-based rate limiting, browser automation overhead, and persistence strategies interact. This guide examines the specific implementation details in the NanmiCoder/MediaCrawler repository that determine crawling speed, latency, and resource consumption.

## Async Concurrency and Semaphore-Based Rate Limiting

All platform crawlers inherit from `AbstractCrawler` in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) and run inside an asyncio event loop using Playwright’s async API. To prevent overwhelming target sites, each implementation (such as [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py)) creates an `asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)` that wraps network-heavy operations.

```python

# From media_platform/xhs/core.py (line 161)

semaphore = asyncio.Semaphore(config.MAX_CONCURRENCY_NUM)
async with semaphore:
    await self.get_comments(note_id)

```

Raising `MAX_CONCURRENCY_NUM` increases raw throughput but consumes more CPU and memory. Conversely, lowering it reduces resource usage but creates bottlenecks during I/O wait times.

### Request Politeness and Sleep Intervals

After each request, crawlers invoke `await asyncio.sleep(config.CRAWLER_MAX_SLEEP_SEC)` to avoid detection. In [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) (line 189), this fixed delay defaults to **2 seconds**. Decreasing this value speeds up crawling but increases the risk of IP bans or rate-limiting.

### Task Batching Memory Impact

Crawlers batch tasks using `asyncio.gather()` as shown in [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) (line 209):

```python
task_list = [self.get_note_info_task(id) for id in note_ids]
await asyncio.gather(*task_list)

```

Large batches increase concurrency but also memory pressure, as each `asyncio.Task` holds references to its entire call stack. For crawls exceeding 10,000 notes, splitting batches keeps RAM usage manageable.

### Tuning Concurrency Parameters

Adjust these values in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) based on hardware capacity:

```python
MAX_CONCURRENCY_NUM = 8          # Parallel network calls (6-8 for 8-core/16GB)

CRAWLER_MAX_SLEEP_SEC = 0.5      # Reduced delay (monitor for bans)

```

For an 8-core machine with 16GB RAM, `MAX_CONCURRENCY_NUM` of 6–8 provides optimal throughput without excessive context switching.

## Browser Launch Modes: CDP vs Standard Playwright

MediaCrawler supports two browser initialization strategies that significantly impact startup latency and resource isolation.

**CDP (Chrome DevTools Protocol) Mode** ([`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py)) connects to an existing Chrome instance via remote debugging (`config.CDP_CONNECT_EXISTING=True`). This eliminates launch overhead and leverages cached sessions but requires a running Chrome browser with debugging enabled.

**Standard Playwright Mode** ([`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py), line 54) launches a fresh Chromium process (`config.ENABLE_CDP_MODE=False`). While this adds seconds to startup and increases memory footprint, it provides complete isolation from user profiles and extensions, making it safer for CI/CD pipelines.

```python

# config/base_config.py

ENABLE_CDP_MODE = False   # Use fresh browser instance

HEADLESS = True           # Disable UI for server environments

BROWSER_LAUNCH_TIMEOUT = 30

```

## Storage Backend I/O Performance

The `SAVE_DATA_OPTION` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) determines how crawled data persists. The implementation resides in platform-specific store classes (e.g., [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py)), where `AsyncFileWriter` handles writes.

| Backend | Throughput | Memory Profile | Best For |
|---------|-----------|---------------|----------|
| **JSONL** | ~10k records/sec | Low (streamed) | Large dumps, easy import |
| **SQLite** | ~5k inserts/sec | Moderate | Local relational queries |
| **MySQL/Postgres** | >20k inserts/sec | Higher | Production deduplication |
| **MongoDB** | ~15k docs/sec | Moderate-High | Schema-flexible pipelines |
| **CSV** | Slow | High (header overhead) | Small datasets only |

Specifically, `XhsJsonStoreImplement` uses `AsyncFileWriter.write_single_item_to_json` (from [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py), line 71) for batched JSON writes. **Avoid CSV for crawls exceeding 1 million records**, as `csv.DictWriter` retains the entire header structure in memory for each write operation.

## Caching Strategy: Redis vs Local Expiring Cache

MediaCrawler provides two caching implementations to reduce redundant network calls and manage proxy tokens.

**RedisCache** ([`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py)) offers ~0.5ms latency and supports shared state across distributed crawler instances. Use this when running multiple workers or when proxy tokens (`IP_PROXY_PROVIDER_NAME`) must persist between runs.

**LocalExpiringCache** ([`cache/local_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/local_cache.py)) provides ~0.05ms access latency for single-process deduplication of URLs and temporary session storage without network overhead.

When `ENABLE_IP_PROXY=True`, the `ProxyIPPool` class ([`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py)) caches live proxies; storing this in Redis prevents re-fetching from the provider on every execution.

## Proxy Pool Configuration

Concurrent proxy connections are controlled by `IP_PROXY_POOL_COUNT` (default: 2) in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). Each crawler thread acquires proxies from `ProxyIPPool` before requests:

```python
IP_PROXY_POOL_COUNT = 4        # Parallel proxy sessions

```

Over-provisioning proxies reduces per-request latency but increases the risk of proxy-side throttling or IP bans from the proxy provider.

## Memory and CPU Optimization

High concurrency creates significant memory pressure. Each `asyncio.Task` object retains references to its call stack; with `MAX_CONCURRENCY_NUM=10` and 10,000 pending notes, resident memory can exceed 1GB.

The word-cloud generator ([`tools/words.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/words.py)) adds CPU-intensive processing when `ENABLE_GET_WORDCLOUD=True`, scaling linearly with comment volume. For crawls exceeding 100,000 comments, disable this feature or offload processing to a separate worker process.

File writers use a per-platform `asyncio.Lock` (in `AsyncFileWriter`) that serializes writes but remains lightweight; I/O costs are dominated by underlying disk speed rather than the locking mechanism itself.

## Practical Configuration Examples

### High-Throughput Tuning

```python

# config/base_config.py

MAX_CONCURRENCY_NUM = 12
CRAWLER_MAX_SLEEP_SEC = 0.2
SAVE_DATA_OPTION = "jsonl"    # Fastest local storage

ENABLE_GET_WORDCLOUD = False   # Reduce CPU load

# Run command

uv run main.py --platform xhs --lt qrcode --type search

```

### Database Persistence for Deduplication

```python

# config/base_config.py

SAVE_DATA_OPTION = "sqlite"
MAX_CONCURRENCY_NUM = 4        # Lower concurrency prevents DB lock contention

# Stores data via store/xhs/sqlite_store.py

```

### Distributed Caching with Redis

```python

# .env

REDIS_DB_HOST=127.0.0.1
REDIS_DB_PORT=6379
REDIS_DB_NUM=0

# config/base_config.py

ENABLE_IP_PROXY = True
IP_PROXY_PROVIDER_NAME = "kuaidaili"

```

## Summary

- **Concurrency tuning**: Adjust `MAX_CONCURRENCY_NUM` and `CRAWLER_MAX_SLEEP_SEC` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to balance throughput against politeness; values of 6–8 work well for mid-range hardware.
- **Browser mode selection**: Use CDP mode ([`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py)) for low-latency development, and standard Playwright mode with `HEADLESS=True` for isolated server deployments.
- **Storage optimization**: Prefer JSONL or database backends (MySQL/Postgres) over CSV for large-scale crawls to minimize memory overhead and maximize insert throughput.
- **Caching architecture**: Deploy `RedisCache` for multi-instance coordination and proxy token persistence; use `LocalExpiringCache` for single-process URL deduplication.
- **Resource management**: Monitor memory consumption when batching thousands of tasks simultaneously, and disable word-cloud generation (`ENABLE_GET_WORDCLOUD`) for high-volume comment extraction.

## Frequently Asked Questions

### How does `MAX_CONCURRENCY_NUM` affect MediaCrawler performance?

This parameter controls the `asyncio.Semaphore` limit that governs how many simultaneous network requests (Playwright page loads or API calls) can execute. Increasing it improves throughput up to the point of CPU saturation or target site rate-limiting, while decreasing it reduces memory usage and detection risk but slows overall crawling speed.

### What is the fastest storage backend for large-scale crawls?

**JSONL** provides the highest local throughput (~10,000 records/sec) with minimal memory footprint due to streaming writes. For production environments requiring complex queries or deduplication, **MySQL or PostgreSQL** exceeds 20,000 inserts per second when connection pools are tuned properly, significantly outperforming CSV and SQLite for concurrent writes.

### Should I use CDP mode or standard Playwright for production?

Use **CDP mode** (`ENABLE_CDP_MODE=True`) when developing locally or debugging, as it connects to an existing Chrome instance and eliminates browser launch latency. Use **standard Playwright mode** (`ENABLE_CDP_MODE=False` with `HEADLESS=True`) for CI/CD pipelines and server deployments, as it provides process isolation and avoids dependencies on GUI availability.

### Why does MediaCrawler consume excessive memory during large crawls?

High memory usage typically stems from large `asyncio.gather()` batches where thousands of `Task` objects retain call stack references simultaneously. Reduce `MAX_CONCURRENCY_NUM` or implement smaller batch sizes in platform core files (e.g., [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py)) to allow garbage collection between groups. Additionally, ensure `ENABLE_GET_WORDCLOUD=False` to prevent in-memory text processing of large comment datasets.