MediaCrawler Architecture: A Complete Guide to Its Main Modules

MediaCrawler is organized around eight core modules—Base Engine, Site-Specific Crawlers, Data Storage, Database Layer, Proxy System, Caching, Configuration, and API/CLI Interfaces—that together provide a modular, extensible framework for scraping media from Chinese social platforms.

The NanmiCoder/MediaCrawler repository implements a layered architecture designed to separate concerns across platform-specific logic, infrastructure services, and user interfaces. Understanding these MediaCrawler architecture modules is essential for extending the system or troubleshooting crawl workflows. This guide breaks down each layer with concrete source references from the codebase.


Core Engine: The Base Crawler Abstraction

The foundation of MediaCrawler architecture is the BaseCrawler abstract class defined in main/base/base_crawler.py. This module establishes the common crawling contract that all platform implementations must follow.

Key responsibilities include:

  • Orchestrating the request lifecycle (authentication, pagination, error handling)
  • Providing hooks for site-specific login flows
  • Managing concurrent request scheduling
  • Standardizing retry and backoff strategies

All platform crawlers—whether scraping Zhihu, Weibo, Tieba, Douyin, Bilibili, Xiaohongshu, or Kuaishou—inherit from this base class. This inheritance pattern ensures consistent behavior across the MediaCrawler architecture while allowing each platform to override specific methods for API endpoint targeting and response parsing.


Site-Specific Crawlers: Platform Implementations

Each supported social platform has dedicated crawler logic residing in the main/ directory. These modules implement the concrete extraction rules for their respective APIs.

Notable implementations include:

  • Zhihu crawler – Handles question/answer pagination and media attachment extraction
  • Weibo crawler – Manages timeline scraping and video/image download orchestration
  • Douyin/Xiaohongshu/Kuaishou – Short-video platform scrapers with signature generation for anti-bot circumvention
  • Bilibili crawler – Danmaku (bullet comment) and video metadata extraction

The site-specific modules reference their corresponding store implementations (e.g., main/store/zhihu/_store_impl.py for Zhihu persistence logic) to hand off extracted data.


Data Storage: Platform-Specific Persistence

The Store Implementations module encapsulates file-level persistence for each platform. Located under main/store/, these modules handle:

  • Writing media files (images, videos, audio) to disk with organized directory structures
  • Serializing metadata to JSON format
  • Managing filename generation and collision avoidance

Each platform has its own _store_impl.py file:

This separation allows platform-specific file organization conventions and media handling without cross-platform coupling.


Database Layer: MongoDB Integration

For structured data persistence, MediaCrawler architecture includes a generic MongoDB interface in main/database/mongodb_store_base.py. This base class provides:

  • Connection pooling and health checking
  • Collection management for user profiles, media records, and crawl logs
  • Async-compatible CRUD operations

Platform crawlers optionally leverage this layer when configured to persist metadata to MongoDB rather than local JSON files. The abstraction keeps database concerns isolated from storage implementation details.


Proxy System: IP Rotation and Pool Management

The Proxy System module, located in main/proxy/, solves rate-limiting and IP-blocking challenges common in social media scraping.

Architecture components:

  • main/proxy/proxy_ip_pool.py – Maintains a rotating pool of healthy proxies with availability checking
  • main/proxy/providers/*.py – Provider-specific adapters for third-party proxy services (e.g., Bright Data, 911 S5, custom lists)

During each request, the base crawler queries pool.get_one() to obtain a live proxy. Failed proxies are automatically deprioritized, ensuring crawl continuity across long-running jobs.

from main.proxy.proxy_ip_pool import ProxyIpPool

pool = ProxyIpPool()
proxy = pool.get_one()
print(f"Using proxy: {proxy.host}:{proxy.port}")

Caching: Redis and Local Cache Layers

To reduce redundant network requests, MediaCrawler architecture implements two caching strategies in main/cache/:

Cache Type Implementation Use Case
Redis Cache main/cache/redis_cache.py Distributed caching across multiple crawler instances; session sharing
Local Cache main/cache/local_cache.py Single-process caching for development or lightweight deployments

Cached objects include session cookies, API rate-limit status, and frequently accessed user profiles. The cache layer integrates transparently with the HTTP utility module (main/tools/httpx_util.py) to return cached responses when available.


Configuration: Centralized Settings Management

Configuration modules in main/config/ provide site-specific and global runtime parameters:

This centralized approach allows environment-specific overrides without modifying crawler logic. The configuration classes are typically loaded at crawler instantiation and remain immutable during execution.


Command-Line and API Interfaces

MediaCrawler architecture exposes two control surfaces for executing crawls:

CLI Interface

The command-line entry point in main/main.py with argument parsing in main/cmd_arg/arg.py:

python -m main.main --platform zhihu --output ./data/zhihu

Arguments specify target platform, output directory, concurrency, and proxy preferences. The parser instantiates the appropriate crawler subclass and initiates execution.

FastAPI Service

For programmatic control, main/api/main.py exposes HTTP endpoints with business logic in main/api/services/crawler_manager.py:

import requests

response = requests.post(
    "http://localhost:8000/crawl",
    json={"platform": "zhihu", "output_dir": "/tmp/zhihu"}
)
print(response.json())

The API layer enables external automation, scheduled jobs, and integration with orchestration systems like Airflow or Prefect.


Utilities: Shared Helper Modules

Supporting the core modules, main/tools/ contains common utilities:

These utilities are imported across crawler and storage modules to ensure consistent behavior.


Summary

The MediaCrawler architecture modules form a cohesive, extensible system for social media data extraction:

  • Base Crawler (main/base/base_crawler.py) provides the abstract workflow foundation
  • Site-Specific Crawlers implement platform extraction logic with inheritance-based customization
  • Store Implementations handle file persistence per platform in main/store/
  • MongoDB Store Base offers optional structured data persistence
  • Proxy System manages IP rotation through main/proxy/proxy_ip_pool.py and provider adapters
  • Caching Layer reduces network load via Redis or local cache
  • Configuration Modules centralize credentials and runtime settings
  • CLI and API Interfaces provide flexible execution models
  • Utility Modules supply shared HTTP, async I/O, and file handling capabilities

Adding a new platform requires only a new crawler subclass and store implementation—proxy, cache, database, and API infrastructure remain reusable.


Frequently Asked Questions

What file contains the main crawling workflow logic in MediaCrawler?

The core workflow is defined in main/base/base_crawler.py. This abstract class specifies the sequence of login, request execution, pagination, and error handling that all platform crawlers inherit. Concrete implementations override methods like login() and parse_response() while keeping the orchestration logic intact.

How does MediaCrawler handle IP blocking and rate limiting?

The main/proxy/proxy_ip_pool.py module maintains a rotating pool of HTTP proxies with health checking. Before each request, crawlers call get_one() to obtain a live proxy. Failed requests trigger proxy deprioritization. Provider-specific adapters in main/proxy/providers/*.py abstract integration with commercial proxy services.

Can MediaCrawler run as a service instead of a command-line tool?

Yes. The main/api/main.py FastAPI application exposes HTTP endpoints for crawl management. The main/api/services/crawler_manager.py module handles request validation, crawler instantiation, and job tracking. This enables integration with schedulers, web dashboards, and automated pipelines.

Where is extracted data stored in MediaCrawler?

Platform-specific store modules in main/store/ handle local file persistence—each platform has its own _store_impl.py file. For database storage, main/database/mongodb_store_base.py provides a generic MongoDB interface. Storage mode is configurable per crawl via the configuration system.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →