What Kind of Media Can MediaCrawler Crawl? Complete Guide to Supported Platforms and Content Types

MediaCrawler can extract text posts, images, videos, and user comments from seven major Chinese social media platforms including Zhihu, XiaoHongShu (RED), Weibo, Tieba, Kuaishou, Douyin, and Bilibili.

MediaCrawler is an open-source social media scraping framework developed by NanmiCoder that provides a unified interface to crawl diverse content types across China's largest digital platforms. Whether you need to archive Q&A threads, download short-form videos, or analyze microblogging data, understanding what kind of media MediaCrawler can crawl helps you leverage its pluggable architecture effectively.

Supported Platforms and Media Types

MediaCrawler uses a factory pattern (CrawlerFactory) to map platform identifiers to concrete crawler implementations. Each platform resides in its own package under media_platform/ and extracts specific media formats:

Zhihu (Q&A and Articles)

The ZhihuCrawler class in media_platform/zhihu/core.py extracts:

  • Text articles and answers
  • User comments
  • Embedded images and videos

XiaoHongShu (RED) (Lifestyle Content)

The XiaoHongShuCrawler in media_platform/xhs/core.py handles:

  • Posts and notes
  • Image galleries
  • Short videos
  • User comments

Weibo (Microblogging)

The WeiboCrawler in media_platform/weibo/core.py captures:

  • Micro-blog posts
  • Pictures and galleries
  • Videos
  • Nested comment threads

Tieba (Forum Discussions)

The TieBaCrawler in media_platform/tieba/core.py scrapes:

  • Forum threads and posts
  • Image attachments
  • Reply chains

Kuaishou (Short Video)

The KuaishouCrawler in media_platform/kuaishou/core.py retrieves:

  • Short video clips
  • Thumbnail images
  • Associated metadata and captions

Douyin (TikTok China)

The DouYinCrawler in media_platform/douyin/core.py extracts:

  • Short videos
  • Cover images
  • Comment data and engagement metrics

Bilibili (Video Sharing)

The BilibiliCrawler in media_platform/bilibili/core.py handles:

  • Video uploads and user-generated clips
  • Image previews
  • Bullet-screen (danmu) comments

How MediaCrawler Processes Different Media Types

The architecture follows a consistent workflow across all platforms, defined in base/base_crawler.py through the AbstractCrawler interface:

  1. Login / Authentication – Each platform package includes a login.py module that manages credential acquisition or token generation.

  2. Client / API Wrapper – Platform-specific client.py modules encapsulate HTTP requests, using httpx or Playwright for dynamic content.

  3. Field Definitions – field.py files enumerate normalized data fields (title, content, media_url, timestamp) that standardize output across different media types.

  4. Parsing Logic – The concrete crawler in core.py extracts raw media and normalizes it into a unified JSON schema.

  5. Output – Results flow through the cache layer (cache/*) or return directly as structured dictionaries.

Because each crawler targets a specific platform, MediaCrawler can harvest text, images, and video content from all seven services while maintaining a consistent Python API.

Practical Code Examples

Instantiate a crawler for specific media extraction using CrawlerFactory defined in main/main.py:

from main.main import CrawlerFactory

# Supported platforms: zhihu, xhs, weibo, tieba, kuaishou, douyin, bilibili

platform = "zhihu"
crawler = CrawlerFactory.create(platform)

# Execute crawl - returns normalized media data

results = crawler.crawl_user(user_id="12345678")

# Results contain media-specific fields

for item in results:
    print(f"Type: {item['type']}, URL: {item.get('media_url')}")

To aggregate media across multiple platforms in a single operation:

from main.main import CrawlerFactory

platforms = ["weibo", "douyin", "bilibili"]
for platform in platforms:
    crawler = CrawlerFactory.create(platform)
    data = crawler.crawl_user(user_id="example_user")
    print(f"{platform.upper()}: Retrieved {len(data)} media items")

Summary

  • MediaCrawler supports seven Chinese social platforms: Zhihu, XiaoHongShu, Weibo, Tieba, Kuaishou, Douyin, and Bilibili.
  • CrawlerFactory in main/main.py instantiates platform-specific crawlers implementing the AbstractCrawler interface from base/base_crawler.py.
  • Each crawler extracts text, images, and videos through standardized field.py definitions and platform-specific core.py implementations.
  • The extensible architecture allows adding new media types by creating packages under media_platform/ with login.py, client.py, field.py, and core.py components.

Frequently Asked Questions

Can MediaCrawler download actual video files or just metadata?

MediaCrawler extracts actual video URLs and binary content from platforms like Douyin, Kuaishou, and Bilibili, not just metadata. The media_url field in returned dictionaries points to downloadable assets, while field.py definitions in each platform package specify additional metadata like resolution and duration.

How does MediaCrawler handle authentication for different platforms?

Each platform package includes a dedicated login.py module that manages authentication flows. According to the source code, these modules handle credential acquisition, cookie management, or token generation required before accessing media content, ensuring crawlers can reach restricted images and videos.

Is it possible to crawl multiple platforms simultaneously?

Yes. You can instantiate multiple crawlers by calling CrawlerFactory.create() with different platform identifiers (defined in cmd_arg/arg.py through CrawlerTypeEnum) within the same script. Each crawler operates independently through its own HTTP client implementation, allowing parallel extraction of text, images, and videos across different services.

What data format does MediaCrawler return?

Crawlers return Python dictionaries normalized according to each platform's field.py specifications. Typical fields include title, content, media_url, timestamp, and type (indicating whether the item is text, image, or video), creating a consistent JSON-like output structure regardless of the source platform.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →