How to Configure Media File (Image/Video) Download and Storage in MediaCrawler
Set SAVE_DATA_OPTION in config/base_config.py to control metadata storage, enable accept_downloads=True in tools/cdp_browser.py for automatic downloads, and customize DEFAULT_DOWNLOAD_PATH in var.py to change where media files are saved.
MediaCrawler is an open-source multi-platform crawler that extracts and stores media from sites like Douyin, Xiaohongshu, and others. Understanding how to configure media file download and storage in MediaCrawler lets you control whether images and videos are fetched as binary files or just stored as URLs, where they land on disk, and how their metadata is persisted.
How Media Download Works in MediaCrawler
MediaCrawler follows a five-stage pipeline for handling media files:
| Stage | Action | Source File |
|---|---|---|
| URL extraction | Platform modules parse API responses and build download URL dictionaries | store/<platform>/__init__.py (e.g., [store/douyin/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py)) |
| Download permission | CDP/Playwright browser launches with accept_downloads=True |
tools/cdp_browser.py#L380 |
| File retrieval | Async HTTP client fetches binary content | [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) |
| Data persistence | Records (including media URLs) saved via configured backend | config/base_config.py#L90 |
| Path resolution | Download directory determined by global constant | [var.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) |
Configure Metadata Storage Format
The SAVE_DATA_OPTION variable in config/base_config.py determines how crawled data—including media URLs—is persisted. This does not affect where binary media files are stored, only the metadata format.
# config/base_config.py
SAVE_DATA_OPTION = "jsonl" # Options: csv, db, json, jsonl, sqlite, excel, postgres
"jsonl"(default): Line-delimited JSON, efficient for large datasets"csv": Spreadsheet-compatible format with limited nesting"excel": Native Excel workbook with sheets"sqlite"/"postgres": Relational database backends"json": Pretty-printed JSON array
Media URLs appear as fields like video_download_url, music_download_url, and note_download_url in whichever format you select.
Enable or Disable Automatic Media Downloads
Binary file download depends on the Playwright browser context configuration in tools/cdp_browser.py. The accept_downloads parameter controls whether the crawler fetches actual files or only extracts URLs.
# tools/cdp_browser.py (excerpt around line 380)
browser = await playwright.chromium.launch_persistent_context(
user_data_dir=user_data_dir,
headless=headless,
accept_downloads=True, # <-- Set to True to auto-download media files
)
accept_downloads=True: Media files are downloaded automatically when URLs are requestedaccept_downloads=False: Only metadata and URLs are collected, no binary download occurs
To skip downloads entirely and save disk space, change this value to False or remove the parameter.
Set the Media Download Directory
The physical location where media files are written is controlled by DEFAULT_DOWNLOAD_PATH in var.py.
# var.py
from pathlib import Path
DEFAULT_DOWNLOAD_PATH = Path.cwd() / "download"
Modify this to any writable path:
# Custom media storage location
DEFAULT_DOWNLOAD_PATH = Path("/mnt/media_archive") # Absolute path
DEFAULT_DOWNLOAD_PATH = Path.home() / "MediaCrawler/files" # User home directory
The AsyncFileWriter utility respects this path and may create subdirectories per platform depending on the store implementation.
Complete Configuration Example
This example switches to Excel metadata storage with a custom media directory:
# Step 1: config/base_config.py
SAVE_DATA_OPTION = "excel"
# Step 2: var.py
from pathlib import Path
DEFAULT_DOWNLOAD_PATH = Path("/data/media_pool")
With these settings, a crawl will:
- Extract media URLs from each platform's API response in
store/<platform>/__init__.py - Download binary files to
/data/media_pool/viaAsyncFileWriter - Persist all records—including
video_download_url,music_download_url, etc.—to an Excel workbook
Manual Media Download for Debugging
To fetch a single URL outside the normal crawl flow, reuse the project's async HTTP pattern:
import httpx
import asyncio
from pathlib import Path
async def download_single_media(url: str, destination: Path) -> None:
"""Download a media file using the same pattern as MediaCrawler."""
async with httpx.AsyncClient(timeout=30) as client:
response = await client.get(url)
response.raise_for_status()
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_bytes(response.content)
print(f"Saved {len(response.content)} bytes to {destination}")
# Example usage
asyncio.run(download_single_media(
"https://example.com/video.mp4",
Path("/data/media_pool/debug_video.mp4")
))
This mirrors the behavior in [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) without invoking the full crawler.
Key Configuration Reference
| File | Purpose | Key Symbol |
|---|---|---|
[config/base_config.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) |
Storage backend selection | SAVE_DATA_OPTION |
[var.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) |
Download directory path | DEFAULT_DOWNLOAD_PATH |
[tools/cdp_browser.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py#L380) |
Browser download permission | accept_downloads |
[store/douyin/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py) |
Platform-specific URL extraction | video_download_url, music_download_url |
[tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) |
Binary file I/O utility | AsyncFileWriter class |
Summary
- Control metadata format: Edit
SAVE_DATA_OPTIONinconfig/base_config.py(jsonl, csv, excel, sqlite, postgres) - Toggle file downloads: Set
accept_downloads=TrueorFalseintools/cdp_browser.py - Change save location: Modify
DEFAULT_DOWNLOAD_PATHinvar.pyto any absolute or relative path - URL extraction happens automatically in each platform's store module—no configuration needed
- Actual media files are never stored in databases; only metadata and URLs are persisted there
Frequently Asked Questions
Does MediaCrawler store images and videos inside the database?
No. According to the NanmiCoder/MediaCrawler source code, only media URLs are stored in the database, CSV, or Excel file. The binary files are written to disk separately under DEFAULT_DOWNLOAD_PATH. Set SAVE_DATA_OPTION to configure where metadata goes, and edit var.py to change where media files land.
How do I stop MediaCrawler from downloading video files and only save URLs?
Set accept_downloads=False in tools/cdp_browser.py around line 380. This prevents the Playwright browser from fetching binary content. The crawler will still extract video_download_url and other fields in the store modules, but no files will be written to DEFAULT_DOWNLOAD_PATH.
Can I use a network-attached storage (NAS) path for media downloads?
Yes. Since DEFAULT_DOWNLOAD_PATH in var.py accepts any pathlib.Path, you can mount your NAS and point to it directly: DEFAULT_DOWNLOAD_PATH = Path("/mnt/nas/media_crawler"). Ensure the running user has write permissions. The AsyncFileWriter utility handles directory creation automatically.
Where does MediaCrawler extract the actual download URLs from?
Each platform's store module extracts URLs from API responses. For example, in [store/douyin/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py), the JSON returned by Douyin's API is parsed to build dictionaries containing video_download_url, music_download_url, and note_download_url keys. Similar extraction logic exists for other platforms in their respective store/ subdirectories.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →