Understanding Platform-Specific Features in NanmiCoder/MediaCrawler: Key Files and Architecture
The NanmiCoder/MediaCrawler repository organizes each social media platform into isolated packages under media_platform/ containing five standard files—core.py, client.py, login.py, field.py, and exception.py—that inherit from a shared abstract base in base/base_crawler.py to handle authentication, API communication, and data extraction.
MediaCrawler is a Python-based social media scraping framework that supports platforms including Weibo, Zhihu, Xiaohongshu, Douyin, Bilibili, Tieba, and Kuaishou through a modular plugin architecture. To understand how platform-specific features are implemented, you must examine the consistent file patterns within each media_platform/<platform>/ directory and the shared abstractions that standardize their behavior. This guide identifies the essential source files that define how each platform handles browser automation, API requests, and data persistence.
The Five-Layer Architecture Pattern
Each platform package follows an identical five-layer structure that separates concerns from authentication to data storage.
-
core.py: Orchestrates the high-level workflow including searching, detail fetching, comment crawling, and media downloads. This file contains the concrete crawler class (e.g.,WeiboCrawler) that inherits fromAbstractCrawlerand implements methods likesearch()andget_specified_notes(). -
client.py: Wraps low-level HTTP and API interactions, handling headers, cookies, and platform-specific request logic usinghttpxor Playwright. -
login.py: Manages authentication flows including QR-code scanning, cookie-based sessions, and mobile login methods. -
field.py: Defines platform-specific constants, search-type enums (e.g.,SearchType), and request parameter schemas. -
exception.py: Declares custom error types likeDataFetchErrorthat bubble up to the core logic for centralized handling.
Abstract Base and Shared Infrastructure
The foundation of every platform implementation resides in base/base_crawler.py, which defines the AbstractCrawler interface. This abstract class mandates three core methods that all concrete crawlers must implement:
start(): Entry point for the crawling workflow.search(): Handles query execution and result pagination.launch_browser(): Manages browser initialization using Chromium instances, often coordinated throughtools/cdp_browser.pywhenENABLE_CDP_MODEis active in the configuration.
By inheriting from this base, each platform crawler gains access to shared helper methods while implementing platform-specific logic for content extraction.
Data Flow from Configuration to Persistence
Understanding the relationship between files requires tracing the execution flow through the system:
-
Configuration → Crawler:
config/base_config.pyand platform-specific configs (e.g.,config/weibo_config.py) set switches likeENABLE_CDP_MODEandWEIBO_SEARCH_TYPE, which initialize the crawler behavior. -
Crawler → Client: The
core.pymodule instantiates a platform-specific client (e.g.,create_weibo_client) with proper headers and authentication tokens. -
Client → API: The client issues HTTP requests to platform endpoints and returns raw JSON structures.
-
Core → Store: Parsed data flows from
core.pyto platform-specific store functions (e.g.,store/weibo/update_weibo_note) defined instore/__init__.py. -
Store → Persistence: Store implementations in
store/<platform>/_store_impl.pymap dictionaries to ORM models indatabase/models.pyor write files directly viatools/async_file_writer.py.
Key Files by Functional Layer
Core Infrastructure
base/base_crawler.py: Defines theAbstractCrawlercontract and shared utilities used by all platform implementations.
Weibo Implementation (Reference Platform)
The Weibo package exemplifies the standard structure:
-
media_platform/weibo/core.py: ContainsWeiboCrawlerwith complete workflow implementation including browser launch, search execution, and media handling. -
media_platform/weibo/client.py: Low-level API wrapper for Weibo endpoints, handling GET requests for notes, comments, and image downloads. -
media_platform/weibo/login.py: QR-code and cookie-based authentication flows. -
media_platform/weibo/field.py:SearchTypeenum and request parameter definitions. -
media_platform/weibo/exception.py: Custom exceptions includingDataFetchError.
Other Platform Packages
Each platform mirrors the Weibo structure:
- Zhihu:
media_platform/zhihu/core.pyandmedia_platform/zhihu/client.pyhandle answer and article extraction. - Xiaohongshu:
media_platform/xhs/core.pymanages XHS-specific crawling logic. - Douyin:
media_platform/douyin/core.pyfor video and comment crawling. - Bilibili:
media_platform/bilibili/core.pyfor user data and video extraction. - Tieba:
media_platform/tieba/core.pyfor post and comment crawling. - Kuaishou:
media_platform/kuaishou/core.pyutilizing GraphQL queries.
Storage Layer
-
store/__init__.py: Exposes platform-specific persistence functions likeupdate_weibo_noteandupdate_zhihu_content. -
store/weibo/_store_impl.py: Weibo-specific storage logic for MongoDB, SQLite, and image handling. -
store/zhihu/_store_impl.py: Zhihu-specific persistence implementations.
Configuration and Utilities
-
config/base_config.py: Global switches including proxy settings, CDP mode, and concurrency limits. -
tools/cdp_browser.py: CDP-mode browser manager shared across all platforms when enabled. -
tools/utils.py: Helper functions for user-agent generation, logging, and cookie conversion.
Practical Implementation Examples
The following examples demonstrate how to instantiate platform-specific crawlers using the abstract base and concrete implementations.
Running the Weibo Crawler
This example mirrors the execution pattern found in api/main.py, launching a search workflow programmatically:
# example: run_weibo.py
import asyncio
from media_platform.weibo.core import WeiboCrawler
from config import config
async def main():
# Configure search parameters
config.CRAWLER_TYPE = "search"
config.KEYWORDS = "AI,机器学习"
# Initialize and start crawler
await WeiboCrawler().start()
if __name__ == "__main__":
asyncio.run(main())
Execution flow: WeiboCrawler.start() → launch_browser (or CDP) → create_weibo_client → search() → weibo_store.update_weibo_note → persistence.
Implementing a Custom Crawler
To extend the framework with a new platform, inherit from the abstract base and implement the required methods:
from base.base_crawler import AbstractCrawler
class CustomCrawler(AbstractCrawler):
async def start(self):
# Platform-specific entry logic
pass
async def search(self):
# Search implementation
pass
async def launch_browser(self, chromium, proxy, ua, headless=True):
# Browser initialization logic
return await super().launch_browser(chromium, proxy, ua, headless)
# Concrete implementations like WeiboCrawler provide full working examples
# of this pattern in production use.
Summary
- MediaCrawler uses a plugin architecture where each platform resides in
media_platform/<platform>/with identical file structures. - Five standard files define each platform:
core.py(workflow),client.py(API),login.py(auth),field.py(constants), andexception.py(errors). - All crawlers inherit from
AbstractCrawlerinbase/base_crawler.py, ensuring consistent interfaces forstart(),search(), andlaunch_browser(). - Data flows through standardized stages: Configuration → Crawler → Client → API → Store → Persistence, with platform-specific logic isolated at each layer.
- Storage implementations are platform-specific in
store/<platform>/_store_impl.pyto handle differing schemas (e.g., Weibo'spicsvs. Zhihu'squestion_id).
Frequently Asked Questions
What is the purpose of core.py in each platform package?
The core.py file contains the concrete crawler class (e.g., WeiboCrawler) that orchestrates the entire extraction workflow. It inherits from AbstractCrawler and implements platform-specific methods for searching content, fetching details, downloading comments, and handling media files, while coordinating with the client.py module for HTTP requests and the store layer for persistence.
How does MediaCrawler handle different authentication methods across platforms?
Each platform implements its own login.py module containing platform-specific authentication logic such as QR-code scanning, cookie-based sessions, or mobile login flows. The core.py module calls these login handlers during initialization before commencing data extraction, allowing Weibo, Zhihu, and other platforms to use completely different authentication mechanisms while maintaining a consistent interface.
Where are platform-specific data schemas defined?
Platform-specific constants and enums are defined in each package's field.py file, including search types and API parameter mappings. The actual data persistence schemas are handled in the store layer, specifically in store/<platform>/_store_impl.py files, which map the raw JSON structures to database models in database/models.py or to file system outputs via tools/async_file_writer.py.
Can I add a new platform by copying an existing one?
Yes, the modular architecture enables rapid platform addition by creating a new directory under media_platform/ with the five standard files (core.py, client.py, login.py, field.py, exception.py) and inheriting from AbstractCrawler in base/base_crawler.py. You must also implement the store interface in store/<new_platform>/_store_impl.py and register the crawler in config/base_config.py to complete the integration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →