How MediaCrawler Handles Images, Videos, and Text: A Complete Technical Guide
MediaCrawler handles different media types by implementing platform-specific dispatch methods in the AbstractCrawler subclass that detect content type from API responses and delegate downloads to specialized storage helpers while treating text as inherent metadata in the JSON response.
The open-source MediaCrawler repository (NanmiCoder/MediaCrawler) provides a unified framework for scraping content from Chinese social media platforms. Understanding how it handles different media types requires examining the platform-specific implementations that inherit from the base AbstractCrawler class and their respective media-dispatch methods.
Architecture Overview: The AbstractCrawler Pattern
Every supported platform in MediaCrawler—XiaoHongShu (XHS), Douyin, and Bilibili—is implemented as a concrete crawler class inheriting from AbstractCrawler. Each platform implements a media-dispatch method that inspects API responses to determine whether the current item contains images, videos, or text-only content.
The core flow follows three distinct phases: detection, delegation, and storage. First, the crawler parses the platform-specific JSON response to identify media fields. Then it delegates binary downloads to HTTP client utilities. Finally, it persists data through platform-specific store modules that abstract filesystem or database operations.
Platform-Specific Media Handling
XiaoHongShu (XHS) Image and Video Extraction
In media_platform/xhs/core.py, the XiaoHongShuCrawler class implements get_notice_media as its primary dispatch method. This method differentiates between media types by checking for the presence of image_list and video URL arrays in the note JSON.
For image handling, the crawler loops over the image_list field, normalizes URLs from url_default to url, and fetches binary data using self.xhs_client.get_note_media. The implementation at lines 62-70 stores results via xhs_store.update_xhs_note_image:
# Simplified flow from xhs/core.py
image_list = note_detail.get("image_list", [])
for img in image_list:
url = img.get("url") or img.get("url_default")
content = await self.xhs_client.get_note_media(url)
await xhs_store.update_xhs_note_image(note_id, content, file_name)
For video handling, the process mirrors image extraction. The crawler builds a list from xhs_store.get_video_url_arr, downloads each URL using the same HTTP client method, and persists via xhs_store.update_xhs_note_video (lines 99-108). Text content is handled implicitly—the note's title and description exist within the note_detail JSON and are persisted by xhs_store.update_xhs_note without additional binary processing.
Douyin Multi-Format Detection
The Douyin implementation in media_platform/douyin/core.py uses get_aweme_media to determine content type through helper extraction functions. The douyin_store._extract_note_image_list method returns a list of image URLs; if empty, the content is classified as video.
Image extraction (lines 18-24) occurs when _extract_note_image_list returns a non-empty list. The crawler iterates through URLs, downloads via self.dy_client.get_aweme_media, and stores via douyin_store.update_dy_aweme_image.
Video extraction (lines 48-57) activates when no images are detected. The crawler retrieves the download URL from douyin_store._extract_video_download_url, downloads the single video (or audio) file, and stores via douyin_store.update_dy_aweme_video:
# From douyin/core.py - simplified dispatch logic
image_list = douyin_store._extract_note_image_list(aweme_item)
if image_list:
await self.get_aweme_images(aweme_item, image_list)
else:
await self.get_aweme_video(aweme_item)
Textual metadata (title, caption) remains part of the aweme JSON structure and is saved alongside media references through the store layer's update methods.
Bilibili Video Processing
The Bilibili crawler in media_platform/bilibili/core.py focuses primarily on video content through get_bilibili_video. Following API response parsing, the implementation selects the highest-resolution video URL using max_size logic, then downloads via self.bili_client.get_video_media (lines 92-99):
# From bilibili/core.py
video_url = bilibili_store._extract_video_download_url(item)
content = await self.bili_client.get_video_media(video_url)
await bilibili_store.store_video(aid, content, "video.mp4")
While Bilibili does not currently expose separate image-only posts in this crawler, the architecture supports adding image extraction following the same pattern used for XHS and Douyin. Textual metadata (title, description) is captured from API responses and stored alongside media files.
Common Patterns Across Platforms
Configuration and Concurrency Control
All three platforms share a unified configuration system. The boolean flag config.ENABLE_GET_MEIDAS (preserving the source code typo) toggles whether media crawling executes at all. Concurrency is controlled via config.MAX_CONCURRENCY_NUM, which initializes a semaphore capping simultaneous download tasks to prevent rate limiting.
# Configuration example
config.ENABLE_GET_MEIDAS = True # Enable image/video downloads
config.MAX_CONCURRENCY_NUM = 8 # Maximum 8 parallel downloads
Storage Abstraction Layer
Each platform implements a dedicated store module (xhs_store, douyin_store, bilibili_store) providing uniform APIs for persistence. These modules abstract whether data is written to filesystem or database, allowing the crawler core to remain agnostic to storage implementation details.
Implementation Examples
To run a complete crawl including media downloads:
# Enable media crawling in configuration
config.ENABLE_GET_MEIDAS = True
config.MAX_CONCURRENCY_NUM = 8
# Execute XiaoHongShu search crawl
# Images and videos are saved automatically via internal dispatch
await XiaoHongShuCrawler().start()
For manual media extraction on specific items:
# Douyin: Process a specific aweme item
aweme_item = await dy_client.get_aweme_detail(aweme_id)
await DouYinCrawler().get_aweme_media(aweme_item)
# Bilibili: Direct video download
video_url = bilibili_store._extract_video_download_url(video_item)
content = await BiliClient().get_video_media(video_url)
await bilibili_store.store_video(aid, content, f"{aid}.mp4")
Summary
- MediaCrawler uses an inheritance pattern where each platform extends
AbstractCrawlerand implements dispatch methods to detect content type from API responses. - XiaoHongShu handles images and videos through
get_notice_mediainmedia_platform/xhs/core.py, checking forimage_listand video URL arrays. - Douyin determines content type via
get_aweme_mediainmedia_platform/douyin/core.pyusing helper extraction functions to distinguish between image galleries and videos. - Bilibili focuses on video processing through
get_bilibili_videoinmedia_platform/bilibili/core.py, selecting highest-resolution streams. - Text handling is implicit across all platforms—metadata exists in JSON responses and is persisted without binary conversion through store layer methods.
- Global configuration controls media downloading via
ENABLE_GET_MEIDASand concurrency viaMAX_CONCURRENCY_NUM.
Frequently Asked Questions
How does MediaCrawler distinguish between image and video content on Douyin?
According to the source code in media_platform/douyin/core.py, the get_aweme_media method calls douyin_store._extract_note_image_list to check for image URLs. If this returns a non-empty list, the content is treated as an image gallery; otherwise, the crawler treats it as video content and extracts the download URL via _extract_video_download_url.
Where is text content stored if there is no dedicated text extraction method?
Textual data (titles, descriptions, captions) is handled implicitly because it already exists as structured fields within the platform API JSON responses. The store layer methods (update_xhs_note, update_dy_aweme, etc.) persist this metadata alongside media references without requiring separate download or conversion steps.
What limits the number of simultaneous media downloads in MediaCrawler?
The concurrency is controlled by the config.MAX_CONCURRENCY_NUM setting, which creates a semaphore that caps parallel download tasks. This prevents overwhelming target servers and reduces the risk of IP bans during high-volume crawling operations across images and videos.
Can I disable media downloads and crawl only text metadata?
Yes. Set config.ENABLE_GET_MEIDAS = False (note the typo matching the source code) to disable the media download phase entirely. When disabled, the crawler skips the binary download delegation and only persists the JSON metadata containing text fields through the standard store update methods.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →