# What Kind of Media Can MediaCrawler Crawl? Complete Guide to Supported Platforms and Content Types

> Discover what media MediaCrawler can crawl. Extract text posts, images, videos, and comments from 7 Chinese platforms like Weibo, Douyin, and Bilibili. Get the complete guide to supported content.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: getting-started
- Published: 2026-07-29

---

**MediaCrawler can extract text posts, images, videos, and user comments from seven major Chinese social media platforms including Zhihu, XiaoHongShu (RED), Weibo, Tieba, Kuaishou, Douyin, and Bilibili.**

MediaCrawler is an open-source social media scraping framework developed by NanmiCoder that provides a unified interface to crawl diverse content types across China's largest digital platforms. Whether you need to archive Q&A threads, download short-form videos, or analyze microblogging data, understanding what kind of media MediaCrawler can crawl helps you leverage its pluggable architecture effectively.

## Supported Platforms and Media Types

MediaCrawler uses a **factory pattern** (`CrawlerFactory`) to map platform identifiers to concrete crawler implementations. Each platform resides in its own package under `media_platform/` and extracts specific media formats:

### Zhihu (Q&A and Articles)

The `ZhihuCrawler` class in [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) extracts:

- Text articles and answers
- User comments
- Embedded images and videos

### XiaoHongShu (RED) (Lifestyle Content)

The `XiaoHongShuCrawler` in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) handles:

- Posts and notes
- Image galleries
- Short videos
- User comments

### Weibo (Microblogging)

The `WeiboCrawler` in [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py) captures:

- Micro-blog posts
- Pictures and galleries
- Videos
- Nested comment threads

### Tieba (Forum Discussions)

The `TieBaCrawler` in [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py) scrapes:

- Forum threads and posts
- Image attachments
- Reply chains

### Kuaishou (Short Video)

The `KuaishouCrawler` in [`media_platform/kuaishou/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/kuaishou/core.py) retrieves:

- Short video clips
- Thumbnail images
- Associated metadata and captions

### Douyin (TikTok China)

The `DouYinCrawler` in [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) extracts:

- Short videos
- Cover images
- Comment data and engagement metrics

### Bilibili (Video Sharing)

The `BilibiliCrawler` in [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) handles:

- Video uploads and user-generated clips
- Image previews
- Bullet-screen (danmu) comments

## How MediaCrawler Processes Different Media Types

The architecture follows a consistent workflow across all platforms, defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) through the `AbstractCrawler` interface:

1. **Login / Authentication** – Each platform package includes a [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) module that manages credential acquisition or token generation.

2. **Client / API Wrapper** – Platform-specific [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) modules encapsulate HTTP requests, using `httpx` or Playwright for dynamic content.

3. **Field Definitions** – [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py) files enumerate normalized data fields (`title`, `content`, `media_url`, `timestamp`) that standardize output across different media types.

4. **Parsing Logic** – The concrete crawler in [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) extracts raw media and normalizes it into a unified JSON schema.

5. **Output** – Results flow through the cache layer (`cache/*`) or return directly as structured dictionaries.

Because each crawler targets a specific platform, **MediaCrawler can harvest text, images, and video content from all seven services** while maintaining a consistent Python API.

## Practical Code Examples

Instantiate a crawler for specific media extraction using `CrawlerFactory` defined in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py):

```python
from main.main import CrawlerFactory

# Supported platforms: zhihu, xhs, weibo, tieba, kuaishou, douyin, bilibili

platform = "zhihu"
crawler = CrawlerFactory.create(platform)

# Execute crawl - returns normalized media data

results = crawler.crawl_user(user_id="12345678")

# Results contain media-specific fields

for item in results:
    print(f"Type: {item['type']}, URL: {item.get('media_url')}")

```

To aggregate media across multiple platforms in a single operation:

```python
from main.main import CrawlerFactory

platforms = ["weibo", "douyin", "bilibili"]
for platform in platforms:
    crawler = CrawlerFactory.create(platform)
    data = crawler.crawl_user(user_id="example_user")
    print(f"{platform.upper()}: Retrieved {len(data)} media items")

```

## Summary

- **MediaCrawler** supports seven Chinese social platforms: Zhihu, XiaoHongShu, Weibo, Tieba, Kuaishou, Douyin, and Bilibili.
- **CrawlerFactory** in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) instantiates platform-specific crawlers implementing the `AbstractCrawler` interface from [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py).
- Each crawler extracts text, images, and videos through standardized [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py) definitions and platform-specific [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) implementations.
- The extensible architecture allows adding new media types by creating packages under `media_platform/` with [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py), and [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) components.

## Frequently Asked Questions

### Can MediaCrawler download actual video files or just metadata?

MediaCrawler extracts actual video URLs and binary content from platforms like Douyin, Kuaishou, and Bilibili, not just metadata. The `media_url` field in returned dictionaries points to downloadable assets, while [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py) definitions in each platform package specify additional metadata like resolution and duration.

### How does MediaCrawler handle authentication for different platforms?

Each platform package includes a dedicated [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) module that manages authentication flows. According to the source code, these modules handle credential acquisition, cookie management, or token generation required before accessing media content, ensuring crawlers can reach restricted images and videos.

### Is it possible to crawl multiple platforms simultaneously?

Yes. You can instantiate multiple crawlers by calling `CrawlerFactory.create()` with different platform identifiers (defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) through `CrawlerTypeEnum`) within the same script. Each crawler operates independently through its own HTTP client implementation, allowing parallel extraction of text, images, and videos across different services.

### What data format does MediaCrawler return?

Crawlers return Python dictionaries normalized according to each platform's [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py) specifications. Typical fields include `title`, `content`, `media_url`, `timestamp`, and `type` (indicating whether the item is text, image, or video), creating a consistent JSON-like output structure regardless of the source platform.