# Enabling Media Downloading During MediaCrawler Crawling: Images, Videos, and Music

> Easily enable media downloading for images, videos, and music with MediaCrawler. Its CDP integration automatically extracts and saves media URLs during crawling on platforms like Douyin and Bilibili. Get started now.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**MediaCrawler automatically enables media downloads through its Chrome DevTools Protocol (CDP) browser integration, extracting and storing direct URLs for videos, images, and music in the database while crawling platforms like Douyin, Bilibili, and Weibo.**

MediaCrawler is an open-source multi-platform crawler that supports scraping content from Douyin, Bilibili, Zhihu, Weibo, Kuaishou, and other social media platforms. When enabling media downloading during MediaCrawler crawling, the system leverages headless browser automation to capture high-quality, non-watermarked media files and persists their URLs for later retrieval via a FastAPI endpoint.

## How Media Downloading Works

MediaCrawler implements a five-stage pipeline that captures media URLs during the crawling process without requiring manual intervention.

### Stage 1: CDP Browser Initialization

The **Chrome DevTools Protocol (CDP) Browser** launches a headless Chromium instance with downloads automatically enabled. In [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py), the launcher configuration includes `{"accept_downloads": True}`, which instructs the browser to automatically accept any download request triggered by the target page.

### Stage 2: Platform Core Execution

Each platform's core module (located in `media_platform/*/core.py`) opens the CDP browser through `cdp_browser.launch_browser()` and inherits the `accept_downloads` setting. For example, [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) at line 343 and [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) at line 443 instantiate the browser with these capabilities, ensuring any download button or API call issued by the site is captured.

### Stage 3: Media URL Extraction

Platform-specific **store modules** parse raw response JSON to extract direct media URLs. In [`store/douyin/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py), three key functions handle this extraction:

- `_extract_video_download_url` – Returns clean, non-watermarked video URLs
- `_extract_music_download_url` – Extracts background music download links
- `_extract_note_image_list` – Retrieves lists of image URLs from note posts

### Stage 4: Database Persistence

Extracted URLs are saved into the **`Media`** table defined in [`database/models.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/models.py). The relevant columns include:

- `video_download_url` – Direct link to video files
- `music_download_url` – Audio track download URLs
- `note_download_url` – Comma-separated list of image URLs

This schema makes the URLs queryable after the crawl finishes.

### Stage 5: API Retrieval

The FastAPI application exposes a generic download endpoint in [`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py). The route `GET /download/{file_path:path}` streams the file located at the stored URL (or a locally cached copy) to the client.

## Key Configuration Files

### CDP Browser Settings

The `accept_downloads` flag is hard-coded to `True` in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py), meaning media downloading is **enabled by default** for all supported platforms. To disable downloads, you would modify this launcher option.

### Database Schema

The `Media` model in [`database/models.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/models.py) stores the extracted URLs persistently. Each media type has its own dedicated column to prevent data mixing and enable efficient querying.

### Store Functions

Platform-specific store modules contain the extraction logic. For Douyin content, [`store/douyin/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py) processes the raw API responses to sanitize and return usable download links.

## Practical Implementation Examples

### Running a Crawl that Captures Media URLs

Use the `CrawlerManager` service to start a crawl. The media download feature requires no additional configuration—it is active by default.

```python
from api.services.crawler_manager import CrawlerManager
from api.schemas.crawler import CrawlerStartRequest, CrawlerTypeEnum

# Create a start request (the "download" feature is already enabled)

request = CrawlerStartRequest(
    crawler_type=CrawlerTypeEnum.DOUYIN,
    keyword="pandas dance",
    max_pages=5,
)

# Launch the crawler

manager = CrawlerManager()
await manager.start_crawler(request)

```

The crawl automatically accepts download prompts and stores the extracted media URLs in the database.

### Accessing Stored Media URLs

Query the `Media` table to retrieve the URLs persisted during crawling:

```python
from sqlalchemy.orm import Session
from database.models import Media

def get_media_urls(session: Session, media_id: int):
    row = session.query(Media).filter_by(id=media_id).one()
    return {
        "video": row.video_download_url,
        "music": row.music_download_url,
        "note_images": row.note_download_url.split(","),  # list of image URLs

    }

```

### Downloading Media via the API

With the API server running, retrieve files using the stored URLs:

```bash

# Assume the API is running on localhost:8000

curl -OJ http://localhost:8000/download/http%3A%2F%2Fx%2Fvideo_h264.mp4

```

The endpoint streams the file located at the decoded URL (`http://x/video_h264.mp4` in this example).

## Platform-Specific Integration Points

Each platform integrates the CDP browser at specific execution points:

- **Douyin**: [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) line 343 initiates the browser instance
- **Zhihu**: [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) line 443 handles browser launch for article crawling
- **Other platforms**: All platform cores follow the same pattern, calling `cdp_browser.launch_browser()` with the pre-configured `accept_downloads=True` parameter

## Summary

- **MediaCrawler enables downloads by default** through the `accept_downloads: True` setting in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py)
- **Platform store modules** extract clean URLs using functions like `_extract_video_download_url` and `_extract_note_image_list` in [`store/douyin/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py)
- **Database storage** persists URLs in the `Media` table columns `video_download_url`, `music_download_url`, and `note_download_url`
- **FastAPI endpoint** at `/download/{file_path:path}` in [`api/routers/data.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/data.py) serves files for post-crawl retrieval
- **No configuration required** to enable media downloading—simply run the crawler normally and query the database or API for results

## Frequently Asked Questions

### Is media downloading enabled by default in MediaCrawler?

Yes. According to the source code in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py), the `accept_downloads` parameter is hard-coded to `True` in the CDP browser launcher options. This means all crawls automatically accept download requests and extract media URLs without requiring explicit configuration.

### Where are the downloaded media files stored?

MediaCrawler stores the **download URLs** rather than the binary files themselves (unless cached). These URLs are saved in the `Media` table in [`database/models.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/models.py) under columns `video_download_url`, `music_download_url`, and `note_download_url`. You can retrieve the actual files later via the `/download/{file_path:path}` API endpoint or download them directly using the stored URLs.

### How does MediaCrawler extract non-watermarked videos from Douyin?

The extraction logic resides in [`store/douyin/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py). The function `_extract_video_download_url` parses the raw JSON response from Douyin's API to identify and return the highest quality, non-watermarked video URL before the platform applies overlays or compression artifacts.

### Can I disable media downloading if I only need text content?

To disable downloads, modify [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) and change `{"accept_downloads": True}` to `False` in the browser launcher options. Alternatively, you can modify the specific platform core files (e.g., [`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py)) to skip the store function calls that extract media URLs, though changing the CDP browser configuration is the most straightforward approach.