What Kind of Media Can MediaCrawler Crawl? Complete Guide to Supported Platforms and Content Types
MediaCrawler can extract text posts, images, videos, and user comments from seven major Chinese social media platforms including Zhihu, XiaoHongShu (RED), Weibo, Tieba, Kuaishou, Douyin, and Bilibili.
MediaCrawler is an open-source social media scraping framework developed by NanmiCoder that provides a unified interface to crawl diverse content types across China's largest digital platforms. Whether you need to archive Q&A threads, download short-form videos, or analyze microblogging data, understanding what kind of media MediaCrawler can crawl helps you leverage its pluggable architecture effectively.
Supported Platforms and Media Types
MediaCrawler uses a factory pattern (CrawlerFactory) to map platform identifiers to concrete crawler implementations. Each platform resides in its own package under media_platform/ and extracts specific media formats:
Zhihu (Q&A and Articles)
The ZhihuCrawler class in media_platform/zhihu/core.py extracts:
- Text articles and answers
- User comments
- Embedded images and videos
XiaoHongShu (RED) (Lifestyle Content)
The XiaoHongShuCrawler in media_platform/xhs/core.py handles:
- Posts and notes
- Image galleries
- Short videos
- User comments
Weibo (Microblogging)
The WeiboCrawler in media_platform/weibo/core.py captures:
- Micro-blog posts
- Pictures and galleries
- Videos
- Nested comment threads
Tieba (Forum Discussions)
The TieBaCrawler in media_platform/tieba/core.py scrapes:
- Forum threads and posts
- Image attachments
- Reply chains
Kuaishou (Short Video)
The KuaishouCrawler in media_platform/kuaishou/core.py retrieves:
- Short video clips
- Thumbnail images
- Associated metadata and captions
Douyin (TikTok China)
The DouYinCrawler in media_platform/douyin/core.py extracts:
- Short videos
- Cover images
- Comment data and engagement metrics
Bilibili (Video Sharing)
The BilibiliCrawler in media_platform/bilibili/core.py handles:
- Video uploads and user-generated clips
- Image previews
- Bullet-screen (danmu) comments
How MediaCrawler Processes Different Media Types
The architecture follows a consistent workflow across all platforms, defined in base/base_crawler.py through the AbstractCrawler interface:
-
Login / Authentication – Each platform package includes a
login.pymodule that manages credential acquisition or token generation. -
Client / API Wrapper – Platform-specific
client.pymodules encapsulate HTTP requests, usinghttpxor Playwright for dynamic content. -
Field Definitions –
field.pyfiles enumerate normalized data fields (title,content,media_url,timestamp) that standardize output across different media types. -
Parsing Logic – The concrete crawler in
core.pyextracts raw media and normalizes it into a unified JSON schema. -
Output – Results flow through the cache layer (
cache/*) or return directly as structured dictionaries.
Because each crawler targets a specific platform, MediaCrawler can harvest text, images, and video content from all seven services while maintaining a consistent Python API.
Practical Code Examples
Instantiate a crawler for specific media extraction using CrawlerFactory defined in main/main.py:
from main.main import CrawlerFactory
# Supported platforms: zhihu, xhs, weibo, tieba, kuaishou, douyin, bilibili
platform = "zhihu"
crawler = CrawlerFactory.create(platform)
# Execute crawl - returns normalized media data
results = crawler.crawl_user(user_id="12345678")
# Results contain media-specific fields
for item in results:
print(f"Type: {item['type']}, URL: {item.get('media_url')}")
To aggregate media across multiple platforms in a single operation:
from main.main import CrawlerFactory
platforms = ["weibo", "douyin", "bilibili"]
for platform in platforms:
crawler = CrawlerFactory.create(platform)
data = crawler.crawl_user(user_id="example_user")
print(f"{platform.upper()}: Retrieved {len(data)} media items")
Summary
- MediaCrawler supports seven Chinese social platforms: Zhihu, XiaoHongShu, Weibo, Tieba, Kuaishou, Douyin, and Bilibili.
- CrawlerFactory in
main/main.pyinstantiates platform-specific crawlers implementing theAbstractCrawlerinterface frombase/base_crawler.py. - Each crawler extracts text, images, and videos through standardized
field.pydefinitions and platform-specificcore.pyimplementations. - The extensible architecture allows adding new media types by creating packages under
media_platform/withlogin.py,client.py,field.py, andcore.pycomponents.
Frequently Asked Questions
Can MediaCrawler download actual video files or just metadata?
MediaCrawler extracts actual video URLs and binary content from platforms like Douyin, Kuaishou, and Bilibili, not just metadata. The media_url field in returned dictionaries points to downloadable assets, while field.py definitions in each platform package specify additional metadata like resolution and duration.
How does MediaCrawler handle authentication for different platforms?
Each platform package includes a dedicated login.py module that manages authentication flows. According to the source code, these modules handle credential acquisition, cookie management, or token generation required before accessing media content, ensuring crawlers can reach restricted images and videos.
Is it possible to crawl multiple platforms simultaneously?
Yes. You can instantiate multiple crawlers by calling CrawlerFactory.create() with different platform identifiers (defined in cmd_arg/arg.py through CrawlerTypeEnum) within the same script. Each crawler operates independently through its own HTTP client implementation, allowing parallel extraction of text, images, and videos across different services.
What data format does MediaCrawler return?
Crawlers return Python dictionaries normalized according to each platform's field.py specifications. Typical fields include title, content, media_url, timestamp, and type (indicating whether the item is text, image, or video), creating a consistent JSON-like output structure regardless of the source platform.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →