Which Platforms Does MediaCrawler Support for Web Scraping?
MediaCrawler supports seven major Chinese social media and content platforms out-of-the-box: Zhihu, Xiaohongshu (RED), Douyin, Bilibili, Kuaishou, Weibo, and Baidu Tieba.
The NanmiCoder/MediaCrawler repository implements a modular scraping framework where each platform maintains its own dedicated store package. This architecture enables consistent API access across diverse services while handling platform-specific authentication, pagination, and data extraction logic.
Supported Platforms and Implementation Details
MediaCrawler organizes platform support through dedicated sub-packages containing concrete crawling logic for HTTP requests, CDP browser automation, pagination, and data extraction. The following table identifies each supported platform and its corresponding implementation files.
Zhihu
The Zhihu implementation resides in store/zhihu/_store_impl.py, providing full support for the Chinese Q&A platform's article and answer scraping.
Xiaohongshu (RED)
Xiaohongshu support is implemented in store/xhs/_store_impl.py, handling the lifestyle platform's note extraction and image processing requirements.
Douyin (TikTok China)
Douyin scraping utilizes store/douyin/_store_impl.py, managing video metadata extraction and short-form content pagination specific to the Chinese TikTok variant.
Bilibili
The Bilibili store at store/bilibili/_store_impl.py supports video clip extraction and弹幕 (danmu) comment scraping for the popular video-sharing platform.
Kuaishou
Kuaishou integration leverages media_platform/kuaishou/login.py alongside store implementations, providing access to the short-video competitor's content ecosystem.
Weibo support utilizes model/m_weibo.py, which defines data models used by specialized crawlers for the microblogging service.
Baidu Tieba
The Tieba implementation uses model/m_baidu_tieba.py for forum thread and post extraction across Baidu's community platform.
How Platform Support Works
MediaCrawler achieves unified multi-platform support through a factory pattern and abstract base classes that orchestrate platform-specific modules.
The Store Factory Pattern
When users initialize a crawler, the store factory (store/__init__.py) instantiates the appropriate platform-specific implementation. This ensures a uniform API across all supported services while abstracting away underlying complexity.
BaseCrawler Orchestration
The BaseCrawler class in main/base/base_crawler.py defines the abstract crawling workflow that all platform implementations follow. Each store package provides concrete methods for authentication handling, request pagination, data parsing, and storage adapter integration.
Configuration and API Access
Platform-specific settings reside in main/config/<platform>_config.py files (e.g., zhihu_config.py, douyin_config.py), containing API keys, rate limits, and authentication credentials. The FastAPI endpoint at main/api/routers/crawler.py exposes these capabilities as web services, allowing remote orchestration of scraping tasks across all supported platforms.
Practical Usage Examples
Initialize the crawler for specific platforms using the unified interface:
# Scrape Zhihu articles
from media_crawler import MediaCrawler
crawler = MediaCrawler(platform="zhihu")
results = crawler.run(keywords="AI", max_pages=3)
print(results)
# Extract Douyin videos
crawler = MediaCrawler(platform="douyin")
videos = crawler.run(user_id="123456789", max_pages=5)
print(videos)
# Scrape Bilibili content
crawler = MediaCrawler(platform="bilibili")
clips = crawler.run(keyword="programming", max_pages=2)
print(clips)
Each call automatically loads the matching store implementation (e.g., store/zhihu/_store_impl.py for Zhihu) and executes the shared crawling pipeline.
Summary
- MediaCrawler supports seven major platforms: Zhihu, Xiaohongshu, Douyin, Bilibili, Kuaishou, Weibo, and Baidu Tieba according to the source code repository structure.
- Platform logic is modularized under
main/store/<platform>/ormodel/directories with concrete implementations like_store_impl.py. - The store factory (
store/__init__.py) and BaseCrawler (main/base/base_crawler.py) provide unified orchestration across all services. - Configuration files at
main/config/<platform>_config.pymanage platform-specific authentication and rate limiting.
Frequently Asked Questions
Which platforms does MediaCrawler support out of the box?
MediaCrawler supports Zhihu, Xiaohongshu (RED), Douyin, Bilibili, Kuaishou, Weibo, and Baidu Tieba. Each platform has dedicated store implementations under main/store/ or specialized models in the model/ directory, providing complete scraping capabilities for content, comments, and user data.
How does MediaCrawler handle different platform APIs?
The framework uses a store factory pattern (store/__init__.py) to instantiate platform-specific implementations while exposing a unified interface through the BaseCrawler class in main/base/base_crawler.py. This abstraction handles authentication, pagination, and data extraction differences transparently.
Can I extend MediaCrawler to support additional platforms?
Yes. Developers can create new platform support by implementing a store package under main/store/<new_platform>/ with a _store_impl.py file following the existing patterns, plus a corresponding configuration file in main/config/. The store factory will automatically register the new platform when referenced in crawler initialization.
Where are the platform-specific configurations stored?
Configuration files reside at main/config/<platform>_config.py (e.g., zhihu_config.py, douyin_config.py). These files contain API endpoints, rate limits, and authentication credentials required for each platform's scraping operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →