Understanding Platform-Specific Features in NanmiCoder/MediaCrawler: Key Files and Architecture

The NanmiCoder/MediaCrawler repository organizes each social media platform into isolated packages under media_platform/ containing five standard files—core.py, client.py, login.py, field.py, and exception.py—that inherit from a shared abstract base in base/base_crawler.py to handle authentication, API communication, and data extraction.

MediaCrawler is a Python-based social media scraping framework that supports platforms including Weibo, Zhihu, Xiaohongshu, Douyin, Bilibili, Tieba, and Kuaishou through a modular plugin architecture. To understand how platform-specific features are implemented, you must examine the consistent file patterns within each media_platform/<platform>/ directory and the shared abstractions that standardize their behavior. This guide identifies the essential source files that define how each platform handles browser automation, API requests, and data persistence.

The Five-Layer Architecture Pattern

Each platform package follows an identical five-layer structure that separates concerns from authentication to data storage.

  • core.py: Orchestrates the high-level workflow including searching, detail fetching, comment crawling, and media downloads. This file contains the concrete crawler class (e.g., WeiboCrawler) that inherits from AbstractCrawler and implements methods like search() and get_specified_notes().

  • client.py: Wraps low-level HTTP and API interactions, handling headers, cookies, and platform-specific request logic using httpx or Playwright.

  • login.py: Manages authentication flows including QR-code scanning, cookie-based sessions, and mobile login methods.

  • field.py: Defines platform-specific constants, search-type enums (e.g., SearchType), and request parameter schemas.

  • exception.py: Declares custom error types like DataFetchError that bubble up to the core logic for centralized handling.

Abstract Base and Shared Infrastructure

The foundation of every platform implementation resides in base/base_crawler.py, which defines the AbstractCrawler interface. This abstract class mandates three core methods that all concrete crawlers must implement:

  • start(): Entry point for the crawling workflow.
  • search(): Handles query execution and result pagination.
  • launch_browser(): Manages browser initialization using Chromium instances, often coordinated through tools/cdp_browser.py when ENABLE_CDP_MODE is active in the configuration.

By inheriting from this base, each platform crawler gains access to shared helper methods while implementing platform-specific logic for content extraction.

Data Flow from Configuration to Persistence

Understanding the relationship between files requires tracing the execution flow through the system:

  1. Configuration → Crawler: config/base_config.py and platform-specific configs (e.g., config/weibo_config.py) set switches like ENABLE_CDP_MODE and WEIBO_SEARCH_TYPE, which initialize the crawler behavior.

  2. Crawler → Client: The core.py module instantiates a platform-specific client (e.g., create_weibo_client) with proper headers and authentication tokens.

  3. Client → API: The client issues HTTP requests to platform endpoints and returns raw JSON structures.

  4. Core → Store: Parsed data flows from core.py to platform-specific store functions (e.g., store/weibo/update_weibo_note) defined in store/__init__.py.

  5. Store → Persistence: Store implementations in store/<platform>/_store_impl.py map dictionaries to ORM models in database/models.py or write files directly via tools/async_file_writer.py.

Key Files by Functional Layer

Core Infrastructure

  • base/base_crawler.py: Defines the AbstractCrawler contract and shared utilities used by all platform implementations.

Weibo Implementation (Reference Platform)

The Weibo package exemplifies the standard structure:

Other Platform Packages

Each platform mirrors the Weibo structure:

Storage Layer

Configuration and Utilities

  • config/base_config.py: Global switches including proxy settings, CDP mode, and concurrency limits.

  • tools/cdp_browser.py: CDP-mode browser manager shared across all platforms when enabled.

  • tools/utils.py: Helper functions for user-agent generation, logging, and cookie conversion.

Practical Implementation Examples

The following examples demonstrate how to instantiate platform-specific crawlers using the abstract base and concrete implementations.

Running the Weibo Crawler

This example mirrors the execution pattern found in api/main.py, launching a search workflow programmatically:


# example: run_weibo.py

import asyncio
from media_platform.weibo.core import WeiboCrawler
from config import config

async def main():
    # Configure search parameters

    config.CRAWLER_TYPE = "search"
    config.KEYWORDS = "AI,机器学习"
    
    # Initialize and start crawler

    await WeiboCrawler().start()

if __name__ == "__main__":
    asyncio.run(main())

Execution flow: WeiboCrawler.start() → launch_browser (or CDP) → create_weibo_client → search() → weibo_store.update_weibo_note → persistence.

Implementing a Custom Crawler

To extend the framework with a new platform, inherit from the abstract base and implement the required methods:

from base.base_crawler import AbstractCrawler

class CustomCrawler(AbstractCrawler):
    async def start(self):
        # Platform-specific entry logic

        pass
        
    async def search(self):
        # Search implementation

        pass
        
    async def launch_browser(self, chromium, proxy, ua, headless=True):
        # Browser initialization logic

        return await super().launch_browser(chromium, proxy, ua, headless)

# Concrete implementations like WeiboCrawler provide full working examples

# of this pattern in production use.

Summary

  • MediaCrawler uses a plugin architecture where each platform resides in media_platform/<platform>/ with identical file structures.
  • Five standard files define each platform: core.py (workflow), client.py (API), login.py (auth), field.py (constants), and exception.py (errors).
  • All crawlers inherit from AbstractCrawler in base/base_crawler.py, ensuring consistent interfaces for start(), search(), and launch_browser().
  • Data flows through standardized stages: Configuration → Crawler → Client → API → Store → Persistence, with platform-specific logic isolated at each layer.
  • Storage implementations are platform-specific in store/<platform>/_store_impl.py to handle differing schemas (e.g., Weibo's pics vs. Zhihu's question_id).

Frequently Asked Questions

What is the purpose of core.py in each platform package?

The core.py file contains the concrete crawler class (e.g., WeiboCrawler) that orchestrates the entire extraction workflow. It inherits from AbstractCrawler and implements platform-specific methods for searching content, fetching details, downloading comments, and handling media files, while coordinating with the client.py module for HTTP requests and the store layer for persistence.

How does MediaCrawler handle different authentication methods across platforms?

Each platform implements its own login.py module containing platform-specific authentication logic such as QR-code scanning, cookie-based sessions, or mobile login flows. The core.py module calls these login handlers during initialization before commencing data extraction, allowing Weibo, Zhihu, and other platforms to use completely different authentication mechanisms while maintaining a consistent interface.

Where are platform-specific data schemas defined?

Platform-specific constants and enums are defined in each package's field.py file, including search types and API parameter mappings. The actual data persistence schemas are handled in the store layer, specifically in store/<platform>/_store_impl.py files, which map the raw JSON structures to database models in database/models.py or to file system outputs via tools/async_file_writer.py.

Can I add a new platform by copying an existing one?

Yes, the modular architecture enables rapid platform addition by creating a new directory under media_platform/ with the five standard files (core.py, client.py, login.py, field.py, exception.py) and inheriting from AbstractCrawler in base/base_crawler.py. You must also implement the store interface in store/<new_platform>/_store_impl.py and register the crawler in config/base_config.py to complete the integration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →