What is the Underlying Technology Stack of MediaCrawler? A Deep Dive into This Async Python Framework

MediaCrawler is a Python 3.10+ asynchronous web scraping framework built on httpx, Playwright CDP, SQLAlchemy, and MongoDB, designed to crawl Chinese media platforms like Weibo, Zhihu, Bilibili, and DouYin.

The MediaCrawler technology stack combines modern async I/O primitives with flexible data persistence and optional browser automation. This article breaks down every layer of its architecture, from core dependencies to platform-specific crawler implementations, based on the source code in the NanmiCoder/MediaCrawler repository.

Core Runtime: Python 3.10+ and Asyncio

The entire codebase is written in Python 3.10 or higher and makes extensive use of type hints and the asyncio event loop. This choice enables non-blocking I/O throughout the crawling pipeline, from HTTP requests to database writes and file exports.

In main.py, the entry point demonstrates this pattern:

async def main():
    await crawler.start()
    await crawler.close()

The async-first design ensures high concurrency when scraping multiple pages or platforms simultaneously.

HTTP and Browser Automation Layer

MediaCrawler uses two complementary tools for fetching content:

  • httpx – An asynchronous HTTP client for API-style requests to endpoints that return JSON or static HTML.
  • Playwright CDP – A Chrome DevTools Protocol wrapper in tools/cdp_browser.py that drives headless Chromium for JavaScript-heavy pages requiring interaction or rendering.

This dual approach allows the framework to handle both lightweight API calls and complex SPAs (Single Page Applications) without hardcoding browser usage everywhere.

Crawler Architecture: Abstract Base and Platform Implementations

The MediaCrawler technology stack enforces a clean abstraction through AbstractCrawler in base/base_crawler.py. This base class defines the contract that every platform-specific crawler must implement:

class AbstractCrawler(ABC):
    @abstractmethod
    async def start(self): ...
    
    @abstractmethod
    async def close(self): ...

Concrete implementations live in media_platform/ and include:

The CrawlerFactory.create_crawler() method (invoked in main.py lines 50-67) instantiates the appropriate class based on config.PLATFORM.

Configuration Management

Settings are centralized in config/*.py files such as zhihu_config.py and dy_config.py. These modules expose declarative constants for:

  • Platform selection (PLATFORM)
  • Proxy configuration
  • Storage backend choice (SAVE_DATA_OPTION)

This design keeps environment-specific values out of the crawler logic itself.

Data Persistence: Multi-Backend Support

The MediaCrawler technology stack supports four storage backends, selected via configuration:

Backend Implementation Use Case
MongoDB database/mongodb_store_base.py Document-oriented storage for flexible schema
SQLite / MySQL / PostgreSQL database/db.py via SQLAlchemy Relational storage with ORM convenience
Excel pandas/openpyxl via tools/async_file_writer.py Business-friendly tabular exports
JSON/JSONL tools/async_file_writer.py Raw data dumps for downstream processing

A separate WordCloud generator can visualize comment text frequencies after crawling completes.

Proxy Management and Anti-Detection

IP rotation is handled by proxy/proxy_ip_pool.py with provider-specific integrations in proxy/providers/*.py. This layer performs health checks and rotation to mitigate rate-limiting and bans from target platforms.

Post-Processing and Utilities

Heavy I/O operations are offloaded to specialized tools:

CLI and Runtime Infrastructure

The command-line interface is defined in cmd_arg/arg.py, while tools/app_runner.py manages graceful startup and shutdown. The entry script main.py orchestrates these pieces:


# main.py lines 35-48: imports

# main.py lines 99-121: execution and cleanup

Optional Web UI

A modern front-end built with Vite, TypeScript, and Tailwind CSS resides in webui/. This optional component provides result visualization and crawler management through a browser interface, compiled to static assets for deployment.

Launching a Crawl: Practical Example

To start crawling Weibo using the public API:


# example.py

import asyncio
from media_platform.weibo import WeiboCrawler
from base.base_crawler import AbstractCrawler

async def run():
    crawler: AbstractCrawler = WeiboCrawler()
    await crawler.start()
    await crawler.close()

if __name__ == "__main__":
    asyncio.run(run())

Export results and generate a word-cloud:

import asyncio
from tools.async_file_writer import AsyncFileWriter
from var import crawler_type_var

async def export():
    writer = AsyncFileWriter(
        platform="dy",
        crawler_type=crawler_type_var.get(),
    )
    await writer.generate_wordcloud_from_comments()

asyncio.run(export())

Summary

The MediaCrawler technology stack delivers a production-ready crawling solution through:

  • Async Python 3.10+ as the foundational runtime
  • httpx and Playwright CDP for flexible HTTP and browser automation
  • AbstractCrawler pattern with platform-specific subclasses in media_platform/
  • Multi-backend persistence via MongoDB, SQLAlchemy-supported SQL databases, and file formats
  • Built-in proxy rotation and CAPTCHA handling tools
  • Optional Vite/TypeScript frontend for result visualization

All components wire together through main.py, with configuration driving behavior and async utilities ensuring efficient resource use.

Frequently Asked Questions

What Python version does MediaCrawler require?

MediaCrawler requires Python 3.10 or higher. The codebase leverages modern type hint syntax and asyncio patterns that depend on this version.

Can MediaCrawler store data in both SQL and NoSQL databases simultaneously?

Yes, the architecture supports multiple backends through separate modules. You can configure SAVE_DATA_OPTION to select one or combine outputs—database/db.py handles SQLAlchemy connections while database/mongodb_store_base.py manages MongoDB, and both can be used in the same pipeline.

How does MediaCrawler handle JavaScript-rendered pages?

For pages requiring JavaScript execution, MediaCrawler uses Playwright CDP via tools/cdp_browser.py. This wrapper communicates with headless Chromium through the Chrome DevTools Protocol, enabling interaction with dynamic content while falling back to httpx for simpler requests.

Is the web UI mandatory for running MediaCrawler?

No, the web UI in webui/ is entirely optional. The core crawler runs via command line (python main.py) without any frontend dependencies. The Vite/TypeScript interface provides convenience for visualization but does not affect crawling functionality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →