# What is the Underlying Technology Stack of MediaCrawler? A Deep Dive into This Async Python Framework

> Explore the async Python framework MediaCrawler its tech stack including httpx Playwright SQLAlchemy and MongoDB Understand its architecture for scraping Chinese media platforms

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: deep-dive
- Published: 2026-08-12

---

**MediaCrawler is a Python 3.10+ asynchronous web scraping framework built on `httpx`, Playwright CDP, SQLAlchemy, and MongoDB, designed to crawl Chinese media platforms like Weibo, Zhihu, Bilibili, and DouYin.**

The **MediaCrawler technology stack** combines modern async I/O primitives with flexible data persistence and optional browser automation. This article breaks down every layer of its architecture, from core dependencies to platform-specific crawler implementations, based on the source code in the `NanmiCoder/MediaCrawler` repository.

## Core Runtime: Python 3.10+ and Asyncio

The entire codebase is written in **Python 3.10 or higher** and makes extensive use of **type hints** and the **`asyncio`** event loop. This choice enables non-blocking I/O throughout the crawling pipeline, from HTTP requests to database writes and file exports.

In [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), the entry point demonstrates this pattern:

```python
async def main():
    await crawler.start()
    await crawler.close()

```

The async-first design ensures high concurrency when scraping multiple pages or platforms simultaneously.

## HTTP and Browser Automation Layer

MediaCrawler uses two complementary tools for fetching content:

- **`httpx`** – An asynchronous HTTP client for API-style requests to endpoints that return JSON or static HTML.
- **Playwright CDP** – A Chrome DevTools Protocol wrapper in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) that drives headless Chromium for JavaScript-heavy pages requiring interaction or rendering.

This dual approach allows the framework to handle both lightweight API calls and complex SPAs (Single Page Applications) without hardcoding browser usage everywhere.

## Crawler Architecture: Abstract Base and Platform Implementations

The **MediaCrawler technology stack** enforces a clean abstraction through `AbstractCrawler` in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This base class defines the contract that every platform-specific crawler must implement:

```python
class AbstractCrawler(ABC):
    @abstractmethod
    async def start(self): ...
    
    @abstractmethod
    async def close(self): ...

```

Concrete implementations live in `media_platform/` and include:

- `WeiboCrawler` ([`weibo.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/weibo.py))
- `ZhihuCrawler` ([`zhihu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/zhihu.py))
- `BilibiliCrawler` ([`bilibili.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/bilibili.py))
- `DouYinCrawler` ([`dy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/dy.py))

The `CrawlerFactory.create_crawler()` method (invoked in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) lines 50-67) instantiates the appropriate class based on `config.PLATFORM`.

## Configuration Management

Settings are centralized in `config/*.py` files such as [`zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/zhihu_config.py) and [`dy_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/dy_config.py). These modules expose declarative constants for:

- Platform selection (`PLATFORM`)
- Proxy configuration
- Storage backend choice (`SAVE_DATA_OPTION`)

This design keeps environment-specific values out of the crawler logic itself.

## Data Persistence: Multi-Backend Support

The **MediaCrawler technology stack** supports four storage backends, selected via configuration:

| Backend | Implementation | Use Case |
|---------|---------------|----------|
| **MongoDB** | [`database/mongodb_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/mongodb_store_base.py) | Document-oriented storage for flexible schema |
| **SQLite / MySQL / PostgreSQL** | [`database/db.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db.py) via SQLAlchemy | Relational storage with ORM convenience |
| **Excel** | `pandas`/`openpyxl` via [`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) | Business-friendly tabular exports |
| **JSON/JSONL** | [`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) | Raw data dumps for downstream processing |

A separate **WordCloud** generator can visualize comment text frequencies after crawling completes.

## Proxy Management and Anti-Detection

IP rotation is handled by [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py) with provider-specific integrations in `proxy/providers/*.py`. This layer performs health checks and rotation to mitigate rate-limiting and bans from target platforms.

## Post-Processing and Utilities

Heavy I/O operations are offloaded to specialized tools:

- [`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) – Asynchronous JSON/Excel serialization
- `tools/wordcloud` – Text visualization from crawled comments
- [`tools/httpx_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/httpx_util.py) – Request throttling and retry logic
- [`tools/slider_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/slider_util.py) – CAPTCHA handling helpers

## CLI and Runtime Infrastructure

The command-line interface is defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py), while [`tools/app_runner.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/app_runner.py) manages graceful startup and shutdown. The entry script [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) orchestrates these pieces:

```python

# main.py lines 35-48: imports

# main.py lines 99-121: execution and cleanup

```

## Optional Web UI

A modern front-end built with **Vite**, **TypeScript**, and **Tailwind CSS** resides in `webui/`. This optional component provides result visualization and crawler management through a browser interface, compiled to static assets for deployment.

## Launching a Crawl: Practical Example

To start crawling Weibo using the public API:

```python

# example.py

import asyncio
from media_platform.weibo import WeiboCrawler
from base.base_crawler import AbstractCrawler

async def run():
    crawler: AbstractCrawler = WeiboCrawler()
    await crawler.start()
    await crawler.close()

if __name__ == "__main__":
    asyncio.run(run())

```

Export results and generate a word-cloud:

```python
import asyncio
from tools.async_file_writer import AsyncFileWriter
from var import crawler_type_var

async def export():
    writer = AsyncFileWriter(
        platform="dy",
        crawler_type=crawler_type_var.get(),
    )
    await writer.generate_wordcloud_from_comments()

asyncio.run(export())

```

## Summary

The **MediaCrawler technology stack** delivers a production-ready crawling solution through:

- **Async Python 3.10+** as the foundational runtime
- **`httpx`** and **Playwright CDP** for flexible HTTP and browser automation
- **AbstractCrawler** pattern with platform-specific subclasses in `media_platform/`
- **Multi-backend persistence** via MongoDB, SQLAlchemy-supported SQL databases, and file formats
- **Built-in proxy rotation** and CAPTCHA handling tools
- **Optional Vite/TypeScript frontend** for result visualization

All components wire together through [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), with configuration driving behavior and async utilities ensuring efficient resource use.

## Frequently Asked Questions

### What Python version does MediaCrawler require?

MediaCrawler requires **Python 3.10 or higher**. The codebase leverages modern type hint syntax and `asyncio` patterns that depend on this version.

### Can MediaCrawler store data in both SQL and NoSQL databases simultaneously?

Yes, the architecture supports multiple backends through separate modules. You can configure `SAVE_DATA_OPTION` to select one or combine outputs—[`database/db.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db.py) handles SQLAlchemy connections while [`database/mongodb_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/mongodb_store_base.py) manages MongoDB, and both can be used in the same pipeline.

### How does MediaCrawler handle JavaScript-rendered pages?

For pages requiring JavaScript execution, MediaCrawler uses **Playwright CDP** via [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py). This wrapper communicates with headless Chromium through the Chrome DevTools Protocol, enabling interaction with dynamic content while falling back to `httpx` for simpler requests.

### Is the web UI mandatory for running MediaCrawler?

No, the web UI in `webui/` is entirely optional. The core crawler runs via command line (`python main.py`) without any frontend dependencies. The Vite/TypeScript interface provides convenience for visualization but does not affect crawling functionality.