What is the Underlying Technology Stack of MediaCrawler? A Deep Dive into This Async Python Framework
MediaCrawler is a Python 3.10+ asynchronous web scraping framework built on httpx, Playwright CDP, SQLAlchemy, and MongoDB, designed to crawl Chinese media platforms like Weibo, Zhihu, Bilibili, and DouYin.
The MediaCrawler technology stack combines modern async I/O primitives with flexible data persistence and optional browser automation. This article breaks down every layer of its architecture, from core dependencies to platform-specific crawler implementations, based on the source code in the NanmiCoder/MediaCrawler repository.
Core Runtime: Python 3.10+ and Asyncio
The entire codebase is written in Python 3.10 or higher and makes extensive use of type hints and the asyncio event loop. This choice enables non-blocking I/O throughout the crawling pipeline, from HTTP requests to database writes and file exports.
In main.py, the entry point demonstrates this pattern:
async def main():
await crawler.start()
await crawler.close()
The async-first design ensures high concurrency when scraping multiple pages or platforms simultaneously.
HTTP and Browser Automation Layer
MediaCrawler uses two complementary tools for fetching content:
httpx– An asynchronous HTTP client for API-style requests to endpoints that return JSON or static HTML.- Playwright CDP – A Chrome DevTools Protocol wrapper in
tools/cdp_browser.pythat drives headless Chromium for JavaScript-heavy pages requiring interaction or rendering.
This dual approach allows the framework to handle both lightweight API calls and complex SPAs (Single Page Applications) without hardcoding browser usage everywhere.
Crawler Architecture: Abstract Base and Platform Implementations
The MediaCrawler technology stack enforces a clean abstraction through AbstractCrawler in base/base_crawler.py. This base class defines the contract that every platform-specific crawler must implement:
class AbstractCrawler(ABC):
@abstractmethod
async def start(self): ...
@abstractmethod
async def close(self): ...
Concrete implementations live in media_platform/ and include:
WeiboCrawler(weibo.py)ZhihuCrawler(zhihu.py)BilibiliCrawler(bilibili.py)DouYinCrawler(dy.py)
The CrawlerFactory.create_crawler() method (invoked in main.py lines 50-67) instantiates the appropriate class based on config.PLATFORM.
Configuration Management
Settings are centralized in config/*.py files such as zhihu_config.py and dy_config.py. These modules expose declarative constants for:
- Platform selection (
PLATFORM) - Proxy configuration
- Storage backend choice (
SAVE_DATA_OPTION)
This design keeps environment-specific values out of the crawler logic itself.
Data Persistence: Multi-Backend Support
The MediaCrawler technology stack supports four storage backends, selected via configuration:
| Backend | Implementation | Use Case |
|---|---|---|
| MongoDB | database/mongodb_store_base.py |
Document-oriented storage for flexible schema |
| SQLite / MySQL / PostgreSQL | database/db.py via SQLAlchemy |
Relational storage with ORM convenience |
| Excel | pandas/openpyxl via tools/async_file_writer.py |
Business-friendly tabular exports |
| JSON/JSONL | tools/async_file_writer.py |
Raw data dumps for downstream processing |
A separate WordCloud generator can visualize comment text frequencies after crawling completes.
Proxy Management and Anti-Detection
IP rotation is handled by proxy/proxy_ip_pool.py with provider-specific integrations in proxy/providers/*.py. This layer performs health checks and rotation to mitigate rate-limiting and bans from target platforms.
Post-Processing and Utilities
Heavy I/O operations are offloaded to specialized tools:
tools/async_file_writer.py– Asynchronous JSON/Excel serializationtools/wordcloud– Text visualization from crawled commentstools/httpx_util.py– Request throttling and retry logictools/slider_util.py– CAPTCHA handling helpers
CLI and Runtime Infrastructure
The command-line interface is defined in cmd_arg/arg.py, while tools/app_runner.py manages graceful startup and shutdown. The entry script main.py orchestrates these pieces:
# main.py lines 35-48: imports
# main.py lines 99-121: execution and cleanup
Optional Web UI
A modern front-end built with Vite, TypeScript, and Tailwind CSS resides in webui/. This optional component provides result visualization and crawler management through a browser interface, compiled to static assets for deployment.
Launching a Crawl: Practical Example
To start crawling Weibo using the public API:
# example.py
import asyncio
from media_platform.weibo import WeiboCrawler
from base.base_crawler import AbstractCrawler
async def run():
crawler: AbstractCrawler = WeiboCrawler()
await crawler.start()
await crawler.close()
if __name__ == "__main__":
asyncio.run(run())
Export results and generate a word-cloud:
import asyncio
from tools.async_file_writer import AsyncFileWriter
from var import crawler_type_var
async def export():
writer = AsyncFileWriter(
platform="dy",
crawler_type=crawler_type_var.get(),
)
await writer.generate_wordcloud_from_comments()
asyncio.run(export())
Summary
The MediaCrawler technology stack delivers a production-ready crawling solution through:
- Async Python 3.10+ as the foundational runtime
httpxand Playwright CDP for flexible HTTP and browser automation- AbstractCrawler pattern with platform-specific subclasses in
media_platform/ - Multi-backend persistence via MongoDB, SQLAlchemy-supported SQL databases, and file formats
- Built-in proxy rotation and CAPTCHA handling tools
- Optional Vite/TypeScript frontend for result visualization
All components wire together through main.py, with configuration driving behavior and async utilities ensuring efficient resource use.
Frequently Asked Questions
What Python version does MediaCrawler require?
MediaCrawler requires Python 3.10 or higher. The codebase leverages modern type hint syntax and asyncio patterns that depend on this version.
Can MediaCrawler store data in both SQL and NoSQL databases simultaneously?
Yes, the architecture supports multiple backends through separate modules. You can configure SAVE_DATA_OPTION to select one or combine outputs—database/db.py handles SQLAlchemy connections while database/mongodb_store_base.py manages MongoDB, and both can be used in the same pipeline.
How does MediaCrawler handle JavaScript-rendered pages?
For pages requiring JavaScript execution, MediaCrawler uses Playwright CDP via tools/cdp_browser.py. This wrapper communicates with headless Chromium through the Chrome DevTools Protocol, enabling interaction with dynamic content while falling back to httpx for simpler requests.
Is the web UI mandatory for running MediaCrawler?
No, the web UI in webui/ is entirely optional. The core crawler runs via command line (python main.py) without any frontend dependencies. The Vite/TypeScript interface provides convenience for visualization but does not affect crawling functionality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →