# How to Configure Media File (Image/Video) Download and Storage in MediaCrawler

> Learn how to configure media file downloads and storage in MediaCrawler. Easily set save options, enable downloads, and customize storage paths for your images and videos.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-14

---

**Set `SAVE_DATA_OPTION` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to control metadata storage, enable `accept_downloads=True` in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) for automatic downloads, and customize `DEFAULT_DOWNLOAD_PATH` in [`var.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) to change where media files are saved.**

MediaCrawler is an open-source multi-platform crawler that extracts and stores media from sites like Douyin, Xiaohongshu, and others. Understanding how to **configure media file download and storage in MediaCrawler** lets you control whether images and videos are fetched as binary files or just stored as URLs, where they land on disk, and how their metadata is persisted.

## How Media Download Works in MediaCrawler

MediaCrawler follows a five-stage pipeline for handling media files:

| Stage | Action | Source File |
|-------|--------|-------------|
| URL extraction | Platform modules parse API responses and build download URL dictionaries | `store/<platform>/__init__.py` (e.g., [[`store/douyin/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py)) |
| Download permission | CDP/Playwright browser launches with `accept_downloads=True` | [`tools/cdp_browser.py#L380`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py#L380) |
| File retrieval | Async HTTP client fetches binary content | [[`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) |
| Data persistence | Records (including media URLs) saved via configured backend | [`config/base_config.py#L90`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py#L90) |
| Path resolution | Download directory determined by global constant | [[`var.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) |

## Configure Metadata Storage Format

The `SAVE_DATA_OPTION` variable in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) determines how crawled data—including media URLs—is persisted. **This does not affect where binary media files are stored**, only the metadata format.

```python

# config/base_config.py

SAVE_DATA_OPTION = "jsonl"   # Options: csv, db, json, jsonl, sqlite, excel, postgres

```

- **`"jsonl"`** (default): Line-delimited JSON, efficient for large datasets
- **`"csv"`**: Spreadsheet-compatible format with limited nesting
- **`"excel"`**: Native Excel workbook with sheets
- **`"sqlite"`/`"postgres"`**: Relational database backends
- **`"json"`**: Pretty-printed JSON array

Media URLs appear as fields like `video_download_url`, `music_download_url`, and `note_download_url` in whichever format you select.

## Enable or Disable Automatic Media Downloads

Binary file download depends on the Playwright browser context configuration in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py). The `accept_downloads` parameter controls whether the crawler fetches actual files or only extracts URLs.

```python

# tools/cdp_browser.py (excerpt around line 380)

browser = await playwright.chromium.launch_persistent_context(
    user_data_dir=user_data_dir,
    headless=headless,
    accept_downloads=True,   # <-- Set to True to auto-download media files

)

```

- **`accept_downloads=True`**: Media files are downloaded automatically when URLs are requested
- **`accept_downloads=False`**: Only metadata and URLs are collected, no binary download occurs

To skip downloads entirely and save disk space, change this value to `False` or remove the parameter.

## Set the Media Download Directory

The physical location where media files are written is controlled by `DEFAULT_DOWNLOAD_PATH` in [`var.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py).

```python

# var.py

from pathlib import Path

DEFAULT_DOWNLOAD_PATH = Path.cwd() / "download"

```

Modify this to any writable path:

```python

# Custom media storage location

DEFAULT_DOWNLOAD_PATH = Path("/mnt/media_archive")           # Absolute path

DEFAULT_DOWNLOAD_PATH = Path.home() / "MediaCrawler/files"   # User home directory

```

The [`AsyncFileWriter`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) utility respects this path and may create subdirectories per platform depending on the store implementation.

## Complete Configuration Example

This example switches to Excel metadata storage with a custom media directory:

```python

# Step 1: config/base_config.py

SAVE_DATA_OPTION = "excel"

# Step 2: var.py

from pathlib import Path
DEFAULT_DOWNLOAD_PATH = Path("/data/media_pool")

```

With these settings, a crawl will:
1. Extract media URLs from each platform's API response in `store/<platform>/__init__.py`
2. Download binary files to `/data/media_pool/` via `AsyncFileWriter`
3. Persist all records—including `video_download_url`, `music_download_url`, etc.—to an Excel workbook

## Manual Media Download for Debugging

To fetch a single URL outside the normal crawl flow, reuse the project's async HTTP pattern:

```python
import httpx
import asyncio
from pathlib import Path

async def download_single_media(url: str, destination: Path) -> None:
    """Download a media file using the same pattern as MediaCrawler."""
    async with httpx.AsyncClient(timeout=30) as client:
        response = await client.get(url)
        response.raise_for_status()
        
        destination.parent.mkdir(parents=True, exist_ok=True)
        destination.write_bytes(response.content)
        print(f"Saved {len(response.content)} bytes to {destination}")

# Example usage

asyncio.run(download_single_media(
    "https://example.com/video.mp4",
    Path("/data/media_pool/debug_video.mp4")
))

```

This mirrors the behavior in [[`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) without invoking the full crawler.

## Key Configuration Reference

| File | Purpose | Key Symbol |
|------|---------|------------|
| [[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) | Storage backend selection | `SAVE_DATA_OPTION` |
| [[`var.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) | Download directory path | `DEFAULT_DOWNLOAD_PATH` |
| [[`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py#L380) | Browser download permission | `accept_downloads` |
| [[`store/douyin/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py) | Platform-specific URL extraction | `video_download_url`, `music_download_url` |
| [[`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) | Binary file I/O utility | `AsyncFileWriter` class |

## Summary

- **Control metadata format**: Edit `SAVE_DATA_OPTION` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) (jsonl, csv, excel, sqlite, postgres)
- **Toggle file downloads**: Set `accept_downloads=True` or `False` in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py)
- **Change save location**: Modify `DEFAULT_DOWNLOAD_PATH` in [`var.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) to any absolute or relative path
- **URL extraction happens automatically** in each platform's store module—no configuration needed
- **Actual media files** are never stored in databases; only metadata and URLs are persisted there

## Frequently Asked Questions

### Does MediaCrawler store images and videos inside the database?

No. According to the `NanmiCoder/MediaCrawler` source code, **only media URLs are stored in the database, CSV, or Excel file**. The binary files are written to disk separately under `DEFAULT_DOWNLOAD_PATH`. Set `SAVE_DATA_OPTION` to configure where metadata goes, and edit [`var.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) to change where media files land.

### How do I stop MediaCrawler from downloading video files and only save URLs?

Set `accept_downloads=False` in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) around line 380. This prevents the Playwright browser from fetching binary content. The crawler will still extract `video_download_url` and other fields in the store modules, but no files will be written to `DEFAULT_DOWNLOAD_PATH`.

### Can I use a network-attached storage (NAS) path for media downloads?

Yes. Since `DEFAULT_DOWNLOAD_PATH` in [`var.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/var.py) accepts any `pathlib.Path`, you can mount your NAS and point to it directly: `DEFAULT_DOWNLOAD_PATH = Path("/mnt/nas/media_crawler")`. Ensure the running user has write permissions. The `AsyncFileWriter` utility handles directory creation automatically.

### Where does MediaCrawler extract the actual download URLs from?

Each platform's store module extracts URLs from API responses. For example, in [[`store/douyin/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/__init__.py), the JSON returned by Douyin's API is parsed to build dictionaries containing `video_download_url`, `music_download_url`, and `note_download_url` keys. Similar extraction logic exists for other platforms in their respective `store/` subdirectories.