Enabling Media Downloading During MediaCrawler Crawling: Images, Videos, and Music
MediaCrawler automatically enables media downloads through its Chrome DevTools Protocol (CDP) browser integration, extracting and storing direct URLs for videos, images, and music in the database while crawling platforms like Douyin, Bilibili, and Weibo.
MediaCrawler is an open-source multi-platform crawler that supports scraping content from Douyin, Bilibili, Zhihu, Weibo, Kuaishou, and other social media platforms. When enabling media downloading during MediaCrawler crawling, the system leverages headless browser automation to capture high-quality, non-watermarked media files and persists their URLs for later retrieval via a FastAPI endpoint.
How Media Downloading Works
MediaCrawler implements a five-stage pipeline that captures media URLs during the crawling process without requiring manual intervention.
Stage 1: CDP Browser Initialization
The Chrome DevTools Protocol (CDP) Browser launches a headless Chromium instance with downloads automatically enabled. In tools/cdp_browser.py, the launcher configuration includes {"accept_downloads": True}, which instructs the browser to automatically accept any download request triggered by the target page.
Stage 2: Platform Core Execution
Each platform's core module (located in media_platform/*/core.py) opens the CDP browser through cdp_browser.launch_browser() and inherits the accept_downloads setting. For example, media_platform/douyin/core.py at line 343 and media_platform/zhihu/core.py at line 443 instantiate the browser with these capabilities, ensuring any download button or API call issued by the site is captured.
Stage 3: Media URL Extraction
Platform-specific store modules parse raw response JSON to extract direct media URLs. In store/douyin/__init__.py, three key functions handle this extraction:
_extract_video_download_url– Returns clean, non-watermarked video URLs_extract_music_download_url– Extracts background music download links_extract_note_image_list– Retrieves lists of image URLs from note posts
Stage 4: Database Persistence
Extracted URLs are saved into the Media table defined in database/models.py. The relevant columns include:
video_download_url– Direct link to video filesmusic_download_url– Audio track download URLsnote_download_url– Comma-separated list of image URLs
This schema makes the URLs queryable after the crawl finishes.
Stage 5: API Retrieval
The FastAPI application exposes a generic download endpoint in api/routers/data.py. The route GET /download/{file_path:path} streams the file located at the stored URL (or a locally cached copy) to the client.
Key Configuration Files
CDP Browser Settings
The accept_downloads flag is hard-coded to True in tools/cdp_browser.py, meaning media downloading is enabled by default for all supported platforms. To disable downloads, you would modify this launcher option.
Database Schema
The Media model in database/models.py stores the extracted URLs persistently. Each media type has its own dedicated column to prevent data mixing and enable efficient querying.
Store Functions
Platform-specific store modules contain the extraction logic. For Douyin content, store/douyin/__init__.py processes the raw API responses to sanitize and return usable download links.
Practical Implementation Examples
Running a Crawl that Captures Media URLs
Use the CrawlerManager service to start a crawl. The media download feature requires no additional configuration—it is active by default.
from api.services.crawler_manager import CrawlerManager
from api.schemas.crawler import CrawlerStartRequest, CrawlerTypeEnum
# Create a start request (the "download" feature is already enabled)
request = CrawlerStartRequest(
crawler_type=CrawlerTypeEnum.DOUYIN,
keyword="pandas dance",
max_pages=5,
)
# Launch the crawler
manager = CrawlerManager()
await manager.start_crawler(request)
The crawl automatically accepts download prompts and stores the extracted media URLs in the database.
Accessing Stored Media URLs
Query the Media table to retrieve the URLs persisted during crawling:
from sqlalchemy.orm import Session
from database.models import Media
def get_media_urls(session: Session, media_id: int):
row = session.query(Media).filter_by(id=media_id).one()
return {
"video": row.video_download_url,
"music": row.music_download_url,
"note_images": row.note_download_url.split(","), # list of image URLs
}
Downloading Media via the API
With the API server running, retrieve files using the stored URLs:
# Assume the API is running on localhost:8000
curl -OJ http://localhost:8000/download/http%3A%2F%2Fx%2Fvideo_h264.mp4
The endpoint streams the file located at the decoded URL (http://x/video_h264.mp4 in this example).
Platform-Specific Integration Points
Each platform integrates the CDP browser at specific execution points:
- Douyin:
media_platform/douyin/core.pyline 343 initiates the browser instance - Zhihu:
media_platform/zhihu/core.pyline 443 handles browser launch for article crawling - Other platforms: All platform cores follow the same pattern, calling
cdp_browser.launch_browser()with the pre-configuredaccept_downloads=Trueparameter
Summary
- MediaCrawler enables downloads by default through the
accept_downloads: Truesetting intools/cdp_browser.py - Platform store modules extract clean URLs using functions like
_extract_video_download_urland_extract_note_image_listinstore/douyin/__init__.py - Database storage persists URLs in the
Mediatable columnsvideo_download_url,music_download_url, andnote_download_url - FastAPI endpoint at
/download/{file_path:path}inapi/routers/data.pyserves files for post-crawl retrieval - No configuration required to enable media downloading—simply run the crawler normally and query the database or API for results
Frequently Asked Questions
Is media downloading enabled by default in MediaCrawler?
Yes. According to the source code in tools/cdp_browser.py, the accept_downloads parameter is hard-coded to True in the CDP browser launcher options. This means all crawls automatically accept download requests and extract media URLs without requiring explicit configuration.
Where are the downloaded media files stored?
MediaCrawler stores the download URLs rather than the binary files themselves (unless cached). These URLs are saved in the Media table in database/models.py under columns video_download_url, music_download_url, and note_download_url. You can retrieve the actual files later via the /download/{file_path:path} API endpoint or download them directly using the stored URLs.
How does MediaCrawler extract non-watermarked videos from Douyin?
The extraction logic resides in store/douyin/__init__.py. The function _extract_video_download_url parses the raw JSON response from Douyin's API to identify and return the highest quality, non-watermarked video URL before the platform applies overlays or compression artifacts.
Can I disable media downloading if I only need text content?
To disable downloads, modify tools/cdp_browser.py and change {"accept_downloads": True} to False in the browser launcher options. Alternatively, you can modify the specific platform core files (e.g., media_platform/douyin/core.py) to skip the store function calls that extract media URLs, though changing the CDP browser configuration is the most straightforward approach.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →