Supported Media Source Types in MediaCrawler: The Complete 7-Platform Guide

MediaCrawler supports seven major Chinese social media platforms—XiaoHongShu, Douyin, Kuaishou, Bilibili, Weibo, Baidu Tieba, and Zhihu—each accessible via short string identifiers defined in the CrawlerFactory mapping.

The open-source MediaCrawler project by NanmiCoder provides a unified scraping interface for China's dominant content platforms. Understanding the supported media source types is essential for configuring your crawling jobs correctly via command-line arguments or programmatic API calls.

Complete List of Supported Media Source Types

MediaCrawler identifies each platform through a short string code registered in the CrawlerFactory class within main.py. The following table lists all seven supported media source types, their identifiers, and their corresponding crawler implementations:

Identifier Platform Name Crawler Class
xhs XiaoHongShu (小红书) XiaoHongShuCrawler
dy Douyin (抖音) DouYinCrawler
ks Kuaishou (快手) KuaishouCrawler
bili Bilibili (哔哩哔哩) BilibiliCrawler
wb Weibo (微博) WeiboCrawler
tieba Baidu Tieba (百度贴吧) TieBaCrawler
zhihu Zhihu (知乎) ZhihuCrawler

These seven identifiers represent the complete set of valid values for the --platform command-line argument.

How Platform Selection Works in CrawlerFactory

The CrawlerFactory.create_crawler() method in main.py (lines 51-58) implements a factory pattern that maps each identifier to its specific crawler class. When instantiating a crawler, the factory validates the requested source type against its internal CRAWLERS dictionary.

Command-line usage requires the --platform flag:

python -m MediaCrawler.main --platform bili --save json

If an unsupported identifier is supplied, MediaCrawler raises a clear error at lines 65-66 in main.py listing the valid options rather than failing silently.

Platform Implementation Architecture

Each media source type maintains its implementation in a dedicated subdirectory under media_platform/. The __init__.py file in each directory exports the crawler class:

These modules handle platform-specific authentication flows, API endpoints, and rate-limiting logic.

Programmatic Usage Examples

Accessing supported media source types programmatically allows for dynamic crawler instantiation and platform discovery.

Creating a specific crawler instance:

from main import CrawlerFactory

# Instantiate a Douyin crawler using the 'dy' identifier

douyin_crawler = CrawlerFactory.create_crawler("dy")
await douyin_crawler.start()

Listing all available platforms dynamically:

from main import CrawlerFactory

# Retrieve all supported identifiers from the factory mapping

supported = sorted(CrawlerFactory.CRAWLERS.keys())
print("Supported media sources:", ", ".join(supported))

# Output: bili, dy, ks, tieba, wb, xhs, zhihu

The CrawlerFactory.CRAWLERS dictionary keys correspond directly to the valid --platform CLI values and the config.PLATFORM configuration setting.

Summary

  • MediaCrawler supports seven Chinese media platforms: XiaoHongShu (xhs), Douyin (dy), Kuaishou (ks), Bilibili (bili), Weibo (wb), Baidu Tieba (tieba), and Zhihu (zhihu)
  • Source identifiers are defined in the CrawlerFactory class mapping located in main.py at lines 51-58
  • Platform-specific implementations reside in media_platform/<identifier>/__init__.py directories
  • Use --platform <identifier> for CLI execution or CrawlerFactory.create_crawler() for Python API access
  • Invalid identifiers trigger explicit validation errors at lines 65-66 in main.py listing all valid options

Frequently Asked Questions

What happens if I use an unsupported platform identifier?

MediaCrawler validates identifiers against the CrawlerFactory.CRAWLERS mapping at initialization. If you provide an unsupported value, the application raises a runtime error at lines 65-66 of main.py that explicitly enumerates the seven valid platform identifiers, preventing silent failures and configuration mistakes.

Can I crawl multiple platforms simultaneously in one command?

The standard CLI interface accepts only one --platform argument per execution. To crawl multiple sources, you must either run separate command instances or create a Python wrapper that iterates through multiple CrawlerFactory.create_crawler() calls with different identifiers, handling each crawler's async lifecycle independently.

Where are the platform-specific crawling rules defined?

Each media source type implements its own scraping logic within dedicated directories. For example, Bilibili's crawler class resides in media_platform/bilibili/__init__.py, while Weibo's implementation is in media_platform/weibo/__init__.py. These modules contain platform-specific selectors, API clients, and authentication handlers that extend the base crawler interface.

How do I add support for a new media source type?

Extending MediaCrawler requires creating a new crawler class in media_platform/<new_source>/__init__.py following the existing pattern, then registering it in the CrawlerFactory.CRAWLERS dictionary in main.py with a unique string identifier. The new class must implement the start() method and handle platform-specific configuration via the config system.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →