How to Use MediaCrawler for Downloading Videos: A Complete Guide
MediaCrawler provides a factory-based architecture that lets you download videos from platforms like Douyin, Bilibili, and Kuaishou via CLI commands or Python async APIs.
MediaCrawler is an open-source crawling framework by NanmiCoder that abstracts platform-specific APIs into a unified interface for extracting media URLs. Whether you need to archive content programmatically or automate bulk downloads, understanding the layered architecture—from the CrawlerFactory in main.py to the store helpers in store/douyin/__init__.py—ensures reliable video extraction.
MediaCrawler Architecture Overview
The codebase follows a clean separation of concerns that isolates platform logic from data persistence:
- Entry Point:
main.pyparses CLI arguments and launches the crawl flow. - CrawlerFactory: Maps platform identifiers (
dy,bili,ks) to concrete implementations. - AbstractCrawler: Defined in
base/base_crawler.py, this base class establishes thecrawl()andclose()interface. - Platform Crawlers: Located in
media_platform/douyin/core.py,media_platform/bilibili/core.py, andmedia_platform/kuaishou/core.py, these modules handle API authentication and raw JSON extraction. - Store Layer:
store/douyin/__init__.pynormalizes payloads and extracts direct download URLs via helper functions like_extract_video_download_url.
When you execute a download, the flow follows four stages: argument parsing, factory instantiation, API crawling, and storage—governed by config.SAVE_DATA_OPTION settings from config/dy_config.py.
Prerequisites and Installation
Before downloading videos, install the required dependencies:
git clone https://github.com/NanmiCoder/MediaCrawler.git
cd MediaCrawler
pip install -r requirements.txt
Verify that cmd_arg/arg.py contains the platform flags you intend to use, as this module defines the CLI interface mapping user inputs to internal crawler configurations.
Command-Line Usage for Video Downloads
The fastest way to use MediaCrawler for downloading videos is through the CLI entry point in main.py. Pass the platform identifier, target URL, and storage format:
python -m main \
--platform dy \
--url https://www.douyin.com/video/1234567890 \
--save-data-option json
Supported --platform values include:
dyfor Douyinbilifor Bilibiliksfor Kuaishou
The --save-data-option flag accepts json, csv, excel, or mongodb, routing output through the store layer defined in config/dy_config.py.
Programmatic Video Extraction
For custom pipelines, import the factory and crawler classes directly from main.py to orchestrate downloads within Python.
Initializing the Crawler
Use CrawlerFactory.create_crawler() to instantiate platform-specific implementations:
import asyncio
from main import CrawlerFactory
from config import dy_config as cfg
# Configure storage backend
cfg.SAVE_DATA_OPTION = "json"
async def download_video():
# Factory returns Douyin crawler for identifier "dy"
crawler = CrawlerFactory.create_crawler("dy")
# Crawl returns dict containing video_download_url
result = await crawler.crawl("https://www.douyin.com/video/1234567890")
print("Direct download link:", result["video_download_url"])
await crawler.close()
asyncio.run(download_video())
Extracting Direct Download URLs
To isolate URL extraction logic without running the full crawler, import helper functions from the store module. In store/douyin/__init__.py, the _extract_video_download_url function parses the raw API response:
from store.douyin import _extract_video_download_url
# Raw payload from Douyin API (aweme_detail structure)
aweme_detail = {
"video": {
"play_addr": {"url_list": ["https://example-cdn.com/video_h264.mp4"]},
"download_addr": {"url_list": ["https://example-cdn.com/video_h264.mp4"]},
}
}
video_url = _extract_video_download_url(aweme_detail)
# Returns: https://example-cdn.com/video_h264.mp4
This approach is useful when you have cached API responses and need to regenerate download links without re-crawling.
Configuration and Storage Options
MediaCrawler behavior is controlled via dy_config.py (for Douyin) and equivalent config modules for other platforms. Key parameters include:
- SAVE_DATA_OPTION: Determines output format (
json,csv,excel,mongodb). - ENABLE_GET_WORDCLOUD: When set to
True, triggersAsyncFileWriterfromtools/async_file_writer.pyto generate comment word-clouds post-crawl.
To flush Excel buffers after programmatic crawling:
from config import dy_config as cfg
from store.excel_store_base import ExcelStoreBase
cfg.SAVE_DATA_OPTION = "excel"
# After crawl completion:
ExcelStoreBase.flush_all()
Platform-Specific Implementation Details
Each platform crawler in media_platform/ inherits from AbstractCrawler and implements the crawl() method to retrieve raw JSON. For Douyin, media_platform/douyin/core.py handles anti-detection headers and rate limiting, then passes the aweme_detail object to store/douyin/__init__.py for URL normalization.
The extraction logic remains consistent across platforms: the crawler retrieves the payload, the store layer extracts media URLs, and main.py coordinates the flow based on cmd_arg/arg.py inputs.
Summary
- MediaCrawler uses a factory pattern (
CrawlerFactoryinmain.py) to map platform identifiers to specific crawlers. - Video URLs are extracted in the store layer—specifically via
_extract_video_download_urlinstore/douyin/__init__.py. - You can trigger downloads via CLI (
python -m main --platform dy --url <URL>) or async Python by instantiating crawlers directly. - Storage formats are controlled through
config/dy_config.pyusing theSAVE_DATA_OPTIONparameter.
Frequently Asked Questions
How do I choose between JSON and Excel output formats?
Set SAVE_DATA_OPTION in your configuration before running the crawler. Valid options defined in config/dy_config.py include json, csv, excel, and mongodb. For Excel, explicitly call ExcelStoreBase.flush_all() after crawling to ensure all rows are written to disk.
Can I download videos from Bilibili using the same commands?
Yes. Replace the --platform argument with bili when using the CLI, or pass "bili" to CrawlerFactory.create_crawler(). The media_platform/bilibili/core.py module handles Bilibili's specific API requirements while maintaining the same interface as the Douyin implementation.
Where does the actual video download happen in the codebase?
The crawler extracts the direct download URL from the platform API, but the actual file download to your local storage depends on your implementation. The crawl() method returns metadata including video_download_url, which you can pass to an HTTP client like aiohttp or requests to save the binary data.
Why does my crawl return metadata but no video file?
MediaCrawler extracts the download URL rather than automatically saving the binary file. Check that config.SAVE_DATA_OPTION is set correctly and that you are accessing the video_download_url key in the returned result dictionary. For automatic downloading, implement HTTP GET logic after receiving the URL from the crawler.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →