Enabling Media Downloading During MediaCrawler Crawling: Images, Videos, and Music

MediaCrawler automatically enables media downloads through its Chrome DevTools Protocol (CDP) browser integration, extracting and storing direct URLs for videos, images, and music in the database while crawling platforms like Douyin, Bilibili, and Weibo.

MediaCrawler is an open-source multi-platform crawler that supports scraping content from Douyin, Bilibili, Zhihu, Weibo, Kuaishou, and other social media platforms. When enabling media downloading during MediaCrawler crawling, the system leverages headless browser automation to capture high-quality, non-watermarked media files and persists their URLs for later retrieval via a FastAPI endpoint.

How Media Downloading Works

MediaCrawler implements a five-stage pipeline that captures media URLs during the crawling process without requiring manual intervention.

Stage 1: CDP Browser Initialization

The Chrome DevTools Protocol (CDP) Browser launches a headless Chromium instance with downloads automatically enabled. In tools/cdp_browser.py, the launcher configuration includes {"accept_downloads": True}, which instructs the browser to automatically accept any download request triggered by the target page.

Stage 2: Platform Core Execution

Each platform's core module (located in media_platform/*/core.py) opens the CDP browser through cdp_browser.launch_browser() and inherits the accept_downloads setting. For example, media_platform/douyin/core.py at line 343 and media_platform/zhihu/core.py at line 443 instantiate the browser with these capabilities, ensuring any download button or API call issued by the site is captured.

Stage 3: Media URL Extraction

Platform-specific store modules parse raw response JSON to extract direct media URLs. In store/douyin/__init__.py, three key functions handle this extraction:

  • _extract_video_download_url – Returns clean, non-watermarked video URLs
  • _extract_music_download_url – Extracts background music download links
  • _extract_note_image_list – Retrieves lists of image URLs from note posts

Stage 4: Database Persistence

Extracted URLs are saved into the Media table defined in database/models.py. The relevant columns include:

  • video_download_url – Direct link to video files
  • music_download_url – Audio track download URLs
  • note_download_url – Comma-separated list of image URLs

This schema makes the URLs queryable after the crawl finishes.

Stage 5: API Retrieval

The FastAPI application exposes a generic download endpoint in api/routers/data.py. The route GET /download/{file_path:path} streams the file located at the stored URL (or a locally cached copy) to the client.

Key Configuration Files

CDP Browser Settings

The accept_downloads flag is hard-coded to True in tools/cdp_browser.py, meaning media downloading is enabled by default for all supported platforms. To disable downloads, you would modify this launcher option.

Database Schema

The Media model in database/models.py stores the extracted URLs persistently. Each media type has its own dedicated column to prevent data mixing and enable efficient querying.

Store Functions

Platform-specific store modules contain the extraction logic. For Douyin content, store/douyin/__init__.py processes the raw API responses to sanitize and return usable download links.

Practical Implementation Examples

Running a Crawl that Captures Media URLs

Use the CrawlerManager service to start a crawl. The media download feature requires no additional configuration—it is active by default.

from api.services.crawler_manager import CrawlerManager
from api.schemas.crawler import CrawlerStartRequest, CrawlerTypeEnum

# Create a start request (the "download" feature is already enabled)

request = CrawlerStartRequest(
    crawler_type=CrawlerTypeEnum.DOUYIN,
    keyword="pandas dance",
    max_pages=5,
)

# Launch the crawler

manager = CrawlerManager()
await manager.start_crawler(request)

The crawl automatically accepts download prompts and stores the extracted media URLs in the database.

Accessing Stored Media URLs

Query the Media table to retrieve the URLs persisted during crawling:

from sqlalchemy.orm import Session
from database.models import Media

def get_media_urls(session: Session, media_id: int):
    row = session.query(Media).filter_by(id=media_id).one()
    return {
        "video": row.video_download_url,
        "music": row.music_download_url,
        "note_images": row.note_download_url.split(","),  # list of image URLs

    }

Downloading Media via the API

With the API server running, retrieve files using the stored URLs:


# Assume the API is running on localhost:8000

curl -OJ http://localhost:8000/download/http%3A%2F%2Fx%2Fvideo_h264.mp4

The endpoint streams the file located at the decoded URL (http://x/video_h264.mp4 in this example).

Platform-Specific Integration Points

Each platform integrates the CDP browser at specific execution points:

  • Douyin: media_platform/douyin/core.py line 343 initiates the browser instance
  • Zhihu: media_platform/zhihu/core.py line 443 handles browser launch for article crawling
  • Other platforms: All platform cores follow the same pattern, calling cdp_browser.launch_browser() with the pre-configured accept_downloads=True parameter

Summary

  • MediaCrawler enables downloads by default through the accept_downloads: True setting in tools/cdp_browser.py
  • Platform store modules extract clean URLs using functions like _extract_video_download_url and _extract_note_image_list in store/douyin/__init__.py
  • Database storage persists URLs in the Media table columns video_download_url, music_download_url, and note_download_url
  • FastAPI endpoint at /download/{file_path:path} in api/routers/data.py serves files for post-crawl retrieval
  • No configuration required to enable media downloading—simply run the crawler normally and query the database or API for results

Frequently Asked Questions

Is media downloading enabled by default in MediaCrawler?

Yes. According to the source code in tools/cdp_browser.py, the accept_downloads parameter is hard-coded to True in the CDP browser launcher options. This means all crawls automatically accept download requests and extract media URLs without requiring explicit configuration.

Where are the downloaded media files stored?

MediaCrawler stores the download URLs rather than the binary files themselves (unless cached). These URLs are saved in the Media table in database/models.py under columns video_download_url, music_download_url, and note_download_url. You can retrieve the actual files later via the /download/{file_path:path} API endpoint or download them directly using the stored URLs.

How does MediaCrawler extract non-watermarked videos from Douyin?

The extraction logic resides in store/douyin/__init__.py. The function _extract_video_download_url parses the raw JSON response from Douyin's API to identify and return the highest quality, non-watermarked video URL before the platform applies overlays or compression artifacts.

Can I disable media downloading if I only need text content?

To disable downloads, modify tools/cdp_browser.py and change {"accept_downloads": True} to False in the browser launcher options. Alternatively, you can modify the specific platform core files (e.g., media_platform/douyin/core.py) to skip the store function calls that extract media URLs, though changing the CDP browser configuration is the most straightforward approach.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →