MediaCrawler Architecture: A Complete Guide to Its Main Modules
MediaCrawler is organized around eight core modules—Base Engine, Site-Specific Crawlers, Data Storage, Database Layer, Proxy System, Caching, Configuration, and API/CLI Interfaces—that together provide a modular, extensible framework for scraping media from Chinese social platforms.
The NanmiCoder/MediaCrawler repository implements a layered architecture designed to separate concerns across platform-specific logic, infrastructure services, and user interfaces. Understanding these MediaCrawler architecture modules is essential for extending the system or troubleshooting crawl workflows. This guide breaks down each layer with concrete source references from the codebase.
Core Engine: The Base Crawler Abstraction
The foundation of MediaCrawler architecture is the BaseCrawler abstract class defined in main/base/base_crawler.py. This module establishes the common crawling contract that all platform implementations must follow.
Key responsibilities include:
- Orchestrating the request lifecycle (authentication, pagination, error handling)
- Providing hooks for site-specific login flows
- Managing concurrent request scheduling
- Standardizing retry and backoff strategies
All platform crawlers—whether scraping Zhihu, Weibo, Tieba, Douyin, Bilibili, Xiaohongshu, or Kuaishou—inherit from this base class. This inheritance pattern ensures consistent behavior across the MediaCrawler architecture while allowing each platform to override specific methods for API endpoint targeting and response parsing.
Site-Specific Crawlers: Platform Implementations
Each supported social platform has dedicated crawler logic residing in the main/ directory. These modules implement the concrete extraction rules for their respective APIs.
Notable implementations include:
- Zhihu crawler – Handles question/answer pagination and media attachment extraction
- Weibo crawler – Manages timeline scraping and video/image download orchestration
- Douyin/Xiaohongshu/Kuaishou – Short-video platform scrapers with signature generation for anti-bot circumvention
- Bilibili crawler – Danmaku (bullet comment) and video metadata extraction
The site-specific modules reference their corresponding store implementations (e.g., main/store/zhihu/_store_impl.py for Zhihu persistence logic) to hand off extracted data.
Data Storage: Platform-Specific Persistence
The Store Implementations module encapsulates file-level persistence for each platform. Located under main/store/, these modules handle:
- Writing media files (images, videos, audio) to disk with organized directory structures
- Serializing metadata to JSON format
- Managing filename generation and collision avoidance
Each platform has its own _store_impl.py file:
main/store/zhihu/_store_impl.pymain/store/weibo/_store_impl.pymain/store/tieba/_store_impl.pymain/store/douyin/_store_impl.pymain/store/bilibili/_store_impl.py
This separation allows platform-specific file organization conventions and media handling without cross-platform coupling.
Database Layer: MongoDB Integration
For structured data persistence, MediaCrawler architecture includes a generic MongoDB interface in main/database/mongodb_store_base.py. This base class provides:
- Connection pooling and health checking
- Collection management for user profiles, media records, and crawl logs
- Async-compatible CRUD operations
Platform crawlers optionally leverage this layer when configured to persist metadata to MongoDB rather than local JSON files. The abstraction keeps database concerns isolated from storage implementation details.
Proxy System: IP Rotation and Pool Management
The Proxy System module, located in main/proxy/, solves rate-limiting and IP-blocking challenges common in social media scraping.
Architecture components:
main/proxy/proxy_ip_pool.py– Maintains a rotating pool of healthy proxies with availability checkingmain/proxy/providers/*.py– Provider-specific adapters for third-party proxy services (e.g., Bright Data, 911 S5, custom lists)
During each request, the base crawler queries pool.get_one() to obtain a live proxy. Failed proxies are automatically deprioritized, ensuring crawl continuity across long-running jobs.
from main.proxy.proxy_ip_pool import ProxyIpPool
pool = ProxyIpPool()
proxy = pool.get_one()
print(f"Using proxy: {proxy.host}:{proxy.port}")
Caching: Redis and Local Cache Layers
To reduce redundant network requests, MediaCrawler architecture implements two caching strategies in main/cache/:
| Cache Type | Implementation | Use Case |
|---|---|---|
| Redis Cache | main/cache/redis_cache.py |
Distributed caching across multiple crawler instances; session sharing |
| Local Cache | main/cache/local_cache.py |
Single-process caching for development or lightweight deployments |
Cached objects include session cookies, API rate-limit status, and frequently accessed user profiles. The cache layer integrates transparently with the HTTP utility module (main/tools/httpx_util.py) to return cached responses when available.
Configuration: Centralized Settings Management
Configuration modules in main/config/ provide site-specific and global runtime parameters:
main/config/zhihu_config.py,main/config/weibo_config.py, etc. – Platform credentials, API endpoints, request headersmain/config/db_config.py– MongoDB connection strings and pool settings- Global options for timeouts, concurrency limits, and retry policies
This centralized approach allows environment-specific overrides without modifying crawler logic. The configuration classes are typically loaded at crawler instantiation and remain immutable during execution.
Command-Line and API Interfaces
MediaCrawler architecture exposes two control surfaces for executing crawls:
CLI Interface
The command-line entry point in main/main.py with argument parsing in main/cmd_arg/arg.py:
python -m main.main --platform zhihu --output ./data/zhihu
Arguments specify target platform, output directory, concurrency, and proxy preferences. The parser instantiates the appropriate crawler subclass and initiates execution.
FastAPI Service
For programmatic control, main/api/main.py exposes HTTP endpoints with business logic in main/api/services/crawler_manager.py:
import requests
response = requests.post(
"http://localhost:8000/crawl",
json={"platform": "zhihu", "output_dir": "/tmp/zhihu"}
)
print(response.json())
The API layer enables external automation, scheduled jobs, and integration with orchestration systems like Airflow or Prefect.
Utilities: Shared Helper Modules
Supporting the core modules, main/tools/ contains common utilities:
main/tools/httpx_util.py– Async HTTP client wrapper with retry logic, proxy injection, and header rotationmain/tools/async_file_writer.py– Non-blocking file I/O for high-throughput media downloadsmain/tools/file_header_manager.py– MIME type detection and file extension validation
These utilities are imported across crawler and storage modules to ensure consistent behavior.
Summary
The MediaCrawler architecture modules form a cohesive, extensible system for social media data extraction:
- Base Crawler (
main/base/base_crawler.py) provides the abstract workflow foundation - Site-Specific Crawlers implement platform extraction logic with inheritance-based customization
- Store Implementations handle file persistence per platform in
main/store/ - MongoDB Store Base offers optional structured data persistence
- Proxy System manages IP rotation through
main/proxy/proxy_ip_pool.pyand provider adapters - Caching Layer reduces network load via Redis or local cache
- Configuration Modules centralize credentials and runtime settings
- CLI and API Interfaces provide flexible execution models
- Utility Modules supply shared HTTP, async I/O, and file handling capabilities
Adding a new platform requires only a new crawler subclass and store implementation—proxy, cache, database, and API infrastructure remain reusable.
Frequently Asked Questions
What file contains the main crawling workflow logic in MediaCrawler?
The core workflow is defined in main/base/base_crawler.py. This abstract class specifies the sequence of login, request execution, pagination, and error handling that all platform crawlers inherit. Concrete implementations override methods like login() and parse_response() while keeping the orchestration logic intact.
How does MediaCrawler handle IP blocking and rate limiting?
The main/proxy/proxy_ip_pool.py module maintains a rotating pool of HTTP proxies with health checking. Before each request, crawlers call get_one() to obtain a live proxy. Failed requests trigger proxy deprioritization. Provider-specific adapters in main/proxy/providers/*.py abstract integration with commercial proxy services.
Can MediaCrawler run as a service instead of a command-line tool?
Yes. The main/api/main.py FastAPI application exposes HTTP endpoints for crawl management. The main/api/services/crawler_manager.py module handles request validation, crawler instantiation, and job tracking. This enables integration with schedulers, web dashboards, and automated pipelines.
Where is extracted data stored in MediaCrawler?
Platform-specific store modules in main/store/ handle local file persistence—each platform has its own _store_impl.py file. For database storage, main/database/mongodb_store_base.py provides a generic MongoDB interface. Storage mode is configurable per crawl via the configuration system.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →