How to Configure Media File (Image/Video) Download and Storage in MediaCrawler

Set SAVE_DATA_OPTION in config/base_config.py to control metadata storage, enable accept_downloads=True in tools/cdp_browser.py for automatic downloads, and customize DEFAULT_DOWNLOAD_PATH in var.py to change where media files are saved.

MediaCrawler is an open-source multi-platform crawler that extracts and stores media from sites like Douyin, Xiaohongshu, and others. Understanding how to configure media file download and storage in MediaCrawler lets you control whether images and videos are fetched as binary files or just stored as URLs, where they land on disk, and how their metadata is persisted.

How Media Download Works in MediaCrawler

MediaCrawler follows a five-stage pipeline for handling media files:

Stage Action Source File
URL extraction Platform modules parse API responses and build download URL dictionaries store/<platform>/__init__.py (e.g., [store/douyin/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py))
Download permission CDP/Playwright browser launches with accept_downloads=True tools/cdp_browser.py#L380
File retrieval Async HTTP client fetches binary content [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py)
Data persistence Records (including media URLs) saved via configured backend config/base_config.py#L90
Path resolution Download directory determined by global constant [var.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py)

Configure Metadata Storage Format

The SAVE_DATA_OPTION variable in config/base_config.py determines how crawled data—including media URLs—is persisted. This does not affect where binary media files are stored, only the metadata format.


# config/base_config.py

SAVE_DATA_OPTION = "jsonl"   # Options: csv, db, json, jsonl, sqlite, excel, postgres
  • "jsonl" (default): Line-delimited JSON, efficient for large datasets
  • "csv": Spreadsheet-compatible format with limited nesting
  • "excel": Native Excel workbook with sheets
  • "sqlite"/"postgres": Relational database backends
  • "json": Pretty-printed JSON array

Media URLs appear as fields like video_download_url, music_download_url, and note_download_url in whichever format you select.

Enable or Disable Automatic Media Downloads

Binary file download depends on the Playwright browser context configuration in tools/cdp_browser.py. The accept_downloads parameter controls whether the crawler fetches actual files or only extracts URLs.


# tools/cdp_browser.py (excerpt around line 380)

browser = await playwright.chromium.launch_persistent_context(
    user_data_dir=user_data_dir,
    headless=headless,
    accept_downloads=True,   # <-- Set to True to auto-download media files

)
  • accept_downloads=True: Media files are downloaded automatically when URLs are requested
  • accept_downloads=False: Only metadata and URLs are collected, no binary download occurs

To skip downloads entirely and save disk space, change this value to False or remove the parameter.

Set the Media Download Directory

The physical location where media files are written is controlled by DEFAULT_DOWNLOAD_PATH in var.py.


# var.py

from pathlib import Path

DEFAULT_DOWNLOAD_PATH = Path.cwd() / "download"

Modify this to any writable path:


# Custom media storage location

DEFAULT_DOWNLOAD_PATH = Path("/mnt/media_archive")           # Absolute path

DEFAULT_DOWNLOAD_PATH = Path.home() / "MediaCrawler/files"   # User home directory

The AsyncFileWriter utility respects this path and may create subdirectories per platform depending on the store implementation.

Complete Configuration Example

This example switches to Excel metadata storage with a custom media directory:


# Step 1: config/base_config.py

SAVE_DATA_OPTION = "excel"

# Step 2: var.py

from pathlib import Path
DEFAULT_DOWNLOAD_PATH = Path("/data/media_pool")

With these settings, a crawl will:

  1. Extract media URLs from each platform's API response in store/<platform>/__init__.py
  2. Download binary files to /data/media_pool/ via AsyncFileWriter
  3. Persist all records—including video_download_url, music_download_url, etc.—to an Excel workbook

Manual Media Download for Debugging

To fetch a single URL outside the normal crawl flow, reuse the project's async HTTP pattern:

import httpx
import asyncio
from pathlib import Path

async def download_single_media(url: str, destination: Path) -> None:
    """Download a media file using the same pattern as MediaCrawler."""
    async with httpx.AsyncClient(timeout=30) as client:
        response = await client.get(url)
        response.raise_for_status()
        
        destination.parent.mkdir(parents=True, exist_ok=True)
        destination.write_bytes(response.content)
        print(f"Saved {len(response.content)} bytes to {destination}")

# Example usage

asyncio.run(download_single_media(
    "https://example.com/video.mp4",
    Path("/data/media_pool/debug_video.mp4")
))

This mirrors the behavior in [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) without invoking the full crawler.

Key Configuration Reference

File Purpose Key Symbol
[config/base_config.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) Storage backend selection SAVE_DATA_OPTION
[var.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) Download directory path DEFAULT_DOWNLOAD_PATH
[tools/cdp_browser.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py#L380) Browser download permission accept_downloads
[store/douyin/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py) Platform-specific URL extraction video_download_url, music_download_url
[tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) Binary file I/O utility AsyncFileWriter class

Summary

  • Control metadata format: Edit SAVE_DATA_OPTION in config/base_config.py (jsonl, csv, excel, sqlite, postgres)
  • Toggle file downloads: Set accept_downloads=True or False in tools/cdp_browser.py
  • Change save location: Modify DEFAULT_DOWNLOAD_PATH in var.py to any absolute or relative path
  • URL extraction happens automatically in each platform's store module—no configuration needed
  • Actual media files are never stored in databases; only metadata and URLs are persisted there

Frequently Asked Questions

Does MediaCrawler store images and videos inside the database?

No. According to the NanmiCoder/MediaCrawler source code, only media URLs are stored in the database, CSV, or Excel file. The binary files are written to disk separately under DEFAULT_DOWNLOAD_PATH. Set SAVE_DATA_OPTION to configure where metadata goes, and edit var.py to change where media files land.

How do I stop MediaCrawler from downloading video files and only save URLs?

Set accept_downloads=False in tools/cdp_browser.py around line 380. This prevents the Playwright browser from fetching binary content. The crawler will still extract video_download_url and other fields in the store modules, but no files will be written to DEFAULT_DOWNLOAD_PATH.

Can I use a network-attached storage (NAS) path for media downloads?

Yes. Since DEFAULT_DOWNLOAD_PATH in var.py accepts any pathlib.Path, you can mount your NAS and point to it directly: DEFAULT_DOWNLOAD_PATH = Path("/mnt/nas/media_crawler"). Ensure the running user has write permissions. The AsyncFileWriter utility handles directory creation automatically.

Where does MediaCrawler extract the actual download URLs from?

Each platform's store module extracts URLs from API responses. For example, in [store/douyin/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py), the JSON returned by Douyin's API is parsed to build dictionaries containing video_download_url, music_download_url, and note_download_url keys. Similar extraction logic exists for other platforms in their respective store/ subdirectories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →