# How MediaCrawler Handles Images, Videos, and Text: A Complete Technical Guide

> Learn how MediaCrawler handles images, videos, and text with platform-specific dispatch methods and specialized storage helpers. A complete technical guide to the NanmiCoder/MediaCrawler repo.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: deep-dive
- Published: 2026-06-28

---

**MediaCrawler handles different media types by implementing platform-specific dispatch methods in the `AbstractCrawler` subclass that detect content type from API responses and delegate downloads to specialized storage helpers while treating text as inherent metadata in the JSON response.**

The open-source MediaCrawler repository (NanmiCoder/MediaCrawler) provides a unified framework for scraping content from Chinese social media platforms. Understanding how it handles different media types requires examining the platform-specific implementations that inherit from the base `AbstractCrawler` class and their respective media-dispatch methods.

## Architecture Overview: The AbstractCrawler Pattern

Every supported platform in MediaCrawler—XiaoHongShu (XHS), Douyin, and Bilibili—is implemented as a concrete crawler class inheriting from `AbstractCrawler`. Each platform implements a **media-dispatch method** that inspects API responses to determine whether the current item contains images, videos, or text-only content.

The core flow follows three distinct phases: detection, delegation, and storage. First, the crawler parses the platform-specific JSON response to identify media fields. Then it delegates binary downloads to HTTP client utilities. Finally, it persists data through platform-specific store modules that abstract filesystem or database operations.

## Platform-Specific Media Handling

### XiaoHongShu (XHS) Image and Video Extraction

In [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py), the `XiaoHongShuCrawler` class implements `get_notice_media` as its primary dispatch method. This method differentiates between media types by checking for the presence of `image_list` and video URL arrays in the note JSON.

For **image handling**, the crawler loops over the `image_list` field, normalizes URLs from `url_default` to `url`, and fetches binary data using `self.xhs_client.get_note_media`. The implementation at lines 62-70 stores results via `xhs_store.update_xhs_note_image`:

```python

# Simplified flow from xhs/core.py

image_list = note_detail.get("image_list", [])
for img in image_list:
    url = img.get("url") or img.get("url_default")
    content = await self.xhs_client.get_note_media(url)
    await xhs_store.update_xhs_note_image(note_id, content, file_name)

```

For **video handling**, the process mirrors image extraction. The crawler builds a list from `xhs_store.get_video_url_arr`, downloads each URL using the same HTTP client method, and persists via `xhs_store.update_xhs_note_video` (lines 99-108). Text content is handled implicitly—the note's title and description exist within the `note_detail` JSON and are persisted by `xhs_store.update_xhs_note` without additional binary processing.

### Douyin Multi-Format Detection

The Douyin implementation in [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) uses `get_aweme_media` to determine content type through helper extraction functions. The `douyin_store._extract_note_image_list` method returns a list of image URLs; if empty, the content is classified as video.

**Image extraction** (lines 18-24) occurs when `_extract_note_image_list` returns a non-empty list. The crawler iterates through URLs, downloads via `self.dy_client.get_aweme_media`, and stores via `douyin_store.update_dy_aweme_image`.

**Video extraction** (lines 48-57) activates when no images are detected. The crawler retrieves the download URL from `douyin_store._extract_video_download_url`, downloads the single video (or audio) file, and stores via `douyin_store.update_dy_aweme_video`:

```python

# From douyin/core.py - simplified dispatch logic

image_list = douyin_store._extract_note_image_list(aweme_item)
if image_list:
    await self.get_aweme_images(aweme_item, image_list)
else:
    await self.get_aweme_video(aweme_item)

```

Textual metadata (title, caption) remains part of the aweme JSON structure and is saved alongside media references through the store layer's update methods.

### Bilibili Video Processing

The Bilibili crawler in [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) focuses primarily on video content through `get_bilibili_video`. Following API response parsing, the implementation selects the highest-resolution video URL using `max_size` logic, then downloads via `self.bili_client.get_video_media` (lines 92-99):

```python

# From bilibili/core.py

video_url = bilibili_store._extract_video_download_url(item)
content = await self.bili_client.get_video_media(video_url)
await bilibili_store.store_video(aid, content, "video.mp4")

```

While Bilibili does not currently expose separate image-only posts in this crawler, the architecture supports adding image extraction following the same pattern used for XHS and Douyin. Textual metadata (title, description) is captured from API responses and stored alongside media files.

## Common Patterns Across Platforms

### Configuration and Concurrency Control

All three platforms share a unified configuration system. The boolean flag `config.ENABLE_GET_MEIDAS` (preserving the source code typo) toggles whether media crawling executes at all. Concurrency is controlled via `config.MAX_CONCURRENCY_NUM`, which initializes a semaphore capping simultaneous download tasks to prevent rate limiting.

```python

# Configuration example

config.ENABLE_GET_MEIDAS = True        # Enable image/video downloads

config.MAX_CONCURRENCY_NUM = 8         # Maximum 8 parallel downloads

```

### Storage Abstraction Layer

Each platform implements a dedicated store module (`xhs_store`, `douyin_store`, `bilibili_store`) providing uniform APIs for persistence. These modules abstract whether data is written to filesystem or database, allowing the crawler core to remain agnostic to storage implementation details.

## Implementation Examples

To run a complete crawl including media downloads:

```python

# Enable media crawling in configuration

config.ENABLE_GET_MEIDAS = True
config.MAX_CONCURRENCY_NUM = 8

# Execute XiaoHongShu search crawl

# Images and videos are saved automatically via internal dispatch

await XiaoHongShuCrawler().start()

```

For manual media extraction on specific items:

```python

# Douyin: Process a specific aweme item

aweme_item = await dy_client.get_aweme_detail(aweme_id)
await DouYinCrawler().get_aweme_media(aweme_item)

# Bilibili: Direct video download

video_url = bilibili_store._extract_video_download_url(video_item)
content = await BiliClient().get_video_media(video_url)
await bilibili_store.store_video(aid, content, f"{aid}.mp4")

```

## Summary

- **MediaCrawler** uses an inheritance pattern where each platform extends `AbstractCrawler` and implements dispatch methods to detect content type from API responses.
- **XiaoHongShu** handles images and videos through `get_notice_media` in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py), checking for `image_list` and video URL arrays.
- **Douyin** determines content type via `get_aweme_media` in [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) using helper extraction functions to distinguish between image galleries and videos.
- **Bilibili** focuses on video processing through `get_bilibili_video` in [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py), selecting highest-resolution streams.
- **Text handling** is implicit across all platforms—metadata exists in JSON responses and is persisted without binary conversion through store layer methods.
- **Global configuration** controls media downloading via `ENABLE_GET_MEIDAS` and concurrency via `MAX_CONCURRENCY_NUM`.

## Frequently Asked Questions

### How does MediaCrawler distinguish between image and video content on Douyin?

According to the source code in [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py), the `get_aweme_media` method calls `douyin_store._extract_note_image_list` to check for image URLs. If this returns a non-empty list, the content is treated as an image gallery; otherwise, the crawler treats it as video content and extracts the download URL via `_extract_video_download_url`.

### Where is text content stored if there is no dedicated text extraction method?

Textual data (titles, descriptions, captions) is handled implicitly because it already exists as structured fields within the platform API JSON responses. The store layer methods (`update_xhs_note`, `update_dy_aweme`, etc.) persist this metadata alongside media references without requiring separate download or conversion steps.

### What limits the number of simultaneous media downloads in MediaCrawler?

The concurrency is controlled by the `config.MAX_CONCURRENCY_NUM` setting, which creates a semaphore that caps parallel download tasks. This prevents overwhelming target servers and reduces the risk of IP bans during high-volume crawling operations across images and videos.

### Can I disable media downloads and crawl only text metadata?

Yes. Set `config.ENABLE_GET_MEIDAS = False` (note the typo matching the source code) to disable the media download phase entirely. When disabled, the crawler skips the binary download delegation and only persists the JSON metadata containing text fields through the standard store update methods.