MediaCrawler Storage Layer Configuration: SQLite, MySQL, MongoDB, JSON/CSV/Excel Options Explained

MediaCrawler supports eight storage backends—csv, json, jsonl, sqlite, mysql/db, postgres, mongodb, and excel—configurable via --save_data_option or SAVE_DATA_OPTION in config/base_config.py.

MediaCrawler's modular storage layer allows you to persist crawl results to different backends without modifying core crawler code. This guide covers all storage layer configuration options in the NanmiCoder/MediaCrawler repository, including setup commands, connection parameters, and platform-specific implementations.

Supported Storage Backends

MediaCrawler implements a registry pattern where each platform's store/{platform}/__init__.py maps string keys to concrete storage classes. The following options are available across all supported platforms (Zhihu, Xiaohongshu, Weibo, Bilibili, Douyin, Kuaishou):

Option Backend Use Case
csv CSV files Simple, human-readable exports; spreadsheet compatibility
json JSON files Structured data with nested fields; API consumption
jsonl JSON Lines files Streaming large datasets; line-by-line processing
sqlite SQLite database Local, serverless relational storage; zero configuration
mysql / db MySQL database Production relational workloads; concurrent access
postgres PostgreSQL database Advanced SQL features; enterprise deployments
mongodb MongoDB collection Flexible schema; document-oriented storage
excel Excel workbook (XLSX) Business reporting; non-technical stakeholders

The authoritative list appears in config/base_config.py at lines 89–90: "csv, db, json, jsonl, sqlite, excel, postgres".

Configuration Methods

Command-Line Selection

Use --save_data_option when running main.py:

uv run main.py --platform zhihu --type search --save_data_option sqlite

Configuration File Default

Set the default backend in config/base_config.py:

SAVE_DATA_OPTION = "sqlite"  # csv, json, jsonl, mysql, postgres, mongodb, excel

Changes to base_config.py affect all subsequent runs unless overridden by CLI arguments.

Database-Specific Setup

SQLite Configuration

SQLite requires no external service. The database path is configured in config/db_config.py at line 50:


# config/db_config.py

SQLITE_DB_PATH = "data/mediacrawler.db"

Initialize the schema before first use:

uv run main.py --init_db sqlite

Then run crawls with SQLite persistence:

uv run main.py --platform zhihu --lt qrcode --type search --save_data_option sqlite

MySQL / PostgreSQL Configuration

Relational databases require credentials in config/db_config.py (line 30 for MySQL):


# config/db_config.py

MYSQL_HOST = "localhost"
MYSQL_PORT = 3306
MYSQL_USER = "mediacrawler"
MYSQL_PASSWORD = "your_password"
MYSQL_DB = "mediacrawler_db"

Initialize schema and crawl:

uv run main.py --init_db mysql
uv run main.py --platform weibo --lt hot --type search --save_data_option mysql

Note: PostgreSQL uses the same --init_db flag with postgres as the argument, handled through the generic DB implementation.

MongoDB Configuration

MongoDB connection parameters reside in db_config.py. Unlike SQL backends, MongoDB does not require --init_db for schema creation due to its schemaless nature:

uv run main.py --platform xhs --type search --save_data_option mongodb

File-Based Storage Usage

JSON and JSONL

Both formats suit downstream processing pipelines. JSON produces a single array object; JSONL writes one document per line for streaming:


# Single structured file

uv run main.py --platform bilibili --type search --save_data_option json

# Line-delimited for large datasets

uv run main.py --platform douyin --type search --save_data_option jsonl

CSV Export

CSV provides immediate spreadsheet compatibility but flattens nested structures:

uv run main.py --platform kuaishou --type search --save_data_option csv

Excel Workbooks

Excel output generates .xlsx files ideal for business stakeholders:

uv run main.py --platform zhihu --type search --save_data_option excel

Platform-Specific Implementation Details

Each crawler platform registers storage implementations in its store/{platform}/__init__.py. For example, in store/zhihu/__init__.py:

  • "csv" → CSV storage class
  • "json" → JSON storage class
  • "jsonl" → JSONL storage class
  • "sqlite" → SQLite storage class
  • "db" / "mysql" → MySQL storage class
  • "postgres" → PostgreSQL storage class (via DB implementation)
  • "mongodb" → MongoDB storage class
  • "excel" → Excel storage class

Equivalent registrations exist in:

The main.py entry point parses --save_data_option and wires the selected implementation through this registry.

Initialization Workflow

Per docs/data_storage_guide.md (lines 20–21, 38–39), the recommended workflow is:

  1. Initialize the target backend's schema (SQL databases only):
uv run main.py --init_db sqlite    # or mysql, postgres
  1. Execute crawls with the selected backend:
uv run main.py --platform zhihu --type search --save_data_option sqlite

Skip --init_db for file-based formats (csv, json, jsonl, excel) and MongoDB.

Key Configuration Files

File Purpose
config/base_config.py Default SAVE_DATA_OPTION; supported backend list
config/db_config.py Connection credentials for SQLite, MySQL, PostgreSQL, MongoDB
store/{platform}/__init__.py Per-platform storage class registry
docs/data_storage_guide.md User documentation for storage initialization
main.py CLI argument parsing and storage wiring

Summary

  • MediaCrawler offers eight storage backends: csv, json, jsonl, sqlite, mysql, postgres, mongodb, excel
  • Configure via --save_data_option CLI flag or SAVE_DATA_OPTION in config/base_config.py
  • SQL databases (SQLite, MySQL, PostgreSQL) require --init_db for schema creation
  • Connection parameters live in config/db_config.py
  • Platform-specific implementations are registered in store/{platform}/__init__.py

Frequently Asked Questions

How do I switch from SQLite to MySQL in MediaCrawler?

Update SAVE_DATA_OPTION = "mysql" in config/base_config.py, add MySQL credentials to config/db_config.py, run uv run main.py --init_db mysql to create tables, then execute crawls. The --save_data_option mysql CLI flag overrides the config file for single runs.

Does MediaCrawler support Excel export without code changes?

Yes. Pass --save_data_option excel when running main.py. The Excel implementation in store/{platform}/__init__.py handles workbook generation automatically. No initialization step is required.

Where is the complete list of supported storage options defined?

The authoritative list appears in config/base_config.py at lines 89–90: "csv, db, json, jsonl, sqlite, excel, postgres". MongoDB is additionally supported via the "mongodb" key in each platform's storage registry.

What is the difference between --init_db and --save_data_option?

--init_db creates database schemas for SQL backends (sqlite, mysql, postgres) and should run once before crawling. --save_data_option selects where crawl results are written during execution. File-based formats and MongoDB do not require --init_db.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →