MediaCrawler Storage Layer Configuration: SQLite, MySQL, MongoDB, JSON/CSV/Excel Options Explained
MediaCrawler supports eight storage backends—csv, json, jsonl, sqlite, mysql/db, postgres, mongodb, and excel—configurable via --save_data_option or SAVE_DATA_OPTION in config/base_config.py.
MediaCrawler's modular storage layer allows you to persist crawl results to different backends without modifying core crawler code. This guide covers all storage layer configuration options in the NanmiCoder/MediaCrawler repository, including setup commands, connection parameters, and platform-specific implementations.
Supported Storage Backends
MediaCrawler implements a registry pattern where each platform's store/{platform}/__init__.py maps string keys to concrete storage classes. The following options are available across all supported platforms (Zhihu, Xiaohongshu, Weibo, Bilibili, Douyin, Kuaishou):
| Option | Backend | Use Case |
|---|---|---|
| csv | CSV files | Simple, human-readable exports; spreadsheet compatibility |
| json | JSON files | Structured data with nested fields; API consumption |
| jsonl | JSON Lines files | Streaming large datasets; line-by-line processing |
| sqlite | SQLite database | Local, serverless relational storage; zero configuration |
| mysql / db | MySQL database | Production relational workloads; concurrent access |
| postgres | PostgreSQL database | Advanced SQL features; enterprise deployments |
| mongodb | MongoDB collection | Flexible schema; document-oriented storage |
| excel | Excel workbook (XLSX) | Business reporting; non-technical stakeholders |
The authoritative list appears in config/base_config.py at lines 89–90: "csv, db, json, jsonl, sqlite, excel, postgres".
Configuration Methods
Command-Line Selection
Use --save_data_option when running main.py:
uv run main.py --platform zhihu --type search --save_data_option sqlite
Configuration File Default
Set the default backend in config/base_config.py:
SAVE_DATA_OPTION = "sqlite" # csv, json, jsonl, mysql, postgres, mongodb, excel
Changes to base_config.py affect all subsequent runs unless overridden by CLI arguments.
Database-Specific Setup
SQLite Configuration
SQLite requires no external service. The database path is configured in config/db_config.py at line 50:
# config/db_config.py
SQLITE_DB_PATH = "data/mediacrawler.db"
Initialize the schema before first use:
uv run main.py --init_db sqlite
Then run crawls with SQLite persistence:
uv run main.py --platform zhihu --lt qrcode --type search --save_data_option sqlite
MySQL / PostgreSQL Configuration
Relational databases require credentials in config/db_config.py (line 30 for MySQL):
# config/db_config.py
MYSQL_HOST = "localhost"
MYSQL_PORT = 3306
MYSQL_USER = "mediacrawler"
MYSQL_PASSWORD = "your_password"
MYSQL_DB = "mediacrawler_db"
Initialize schema and crawl:
uv run main.py --init_db mysql
uv run main.py --platform weibo --lt hot --type search --save_data_option mysql
Note: PostgreSQL uses the same --init_db flag with postgres as the argument, handled through the generic DB implementation.
MongoDB Configuration
MongoDB connection parameters reside in db_config.py. Unlike SQL backends, MongoDB does not require --init_db for schema creation due to its schemaless nature:
uv run main.py --platform xhs --type search --save_data_option mongodb
File-Based Storage Usage
JSON and JSONL
Both formats suit downstream processing pipelines. JSON produces a single array object; JSONL writes one document per line for streaming:
# Single structured file
uv run main.py --platform bilibili --type search --save_data_option json
# Line-delimited for large datasets
uv run main.py --platform douyin --type search --save_data_option jsonl
CSV Export
CSV provides immediate spreadsheet compatibility but flattens nested structures:
uv run main.py --platform kuaishou --type search --save_data_option csv
Excel Workbooks
Excel output generates .xlsx files ideal for business stakeholders:
uv run main.py --platform zhihu --type search --save_data_option excel
Platform-Specific Implementation Details
Each crawler platform registers storage implementations in its store/{platform}/__init__.py. For example, in store/zhihu/__init__.py:
"csv"→ CSV storage class"json"→ JSON storage class"jsonl"→ JSONL storage class"sqlite"→ SQLite storage class"db"/"mysql"→ MySQL storage class"postgres"→ PostgreSQL storage class (via DB implementation)"mongodb"→ MongoDB storage class"excel"→ Excel storage class
Equivalent registrations exist in:
store/xhs/__init__.py(Xiaohongshu)store/weibo/__init__.py(Weibo)store/bilibili/__init__.py(Bilibili)store/douyin/__init__.py(Douyin)store/kuaishou/__init__.py(Kuaishou)
The main.py entry point parses --save_data_option and wires the selected implementation through this registry.
Initialization Workflow
Per docs/data_storage_guide.md (lines 20–21, 38–39), the recommended workflow is:
- Initialize the target backend's schema (SQL databases only):
uv run main.py --init_db sqlite # or mysql, postgres
- Execute crawls with the selected backend:
uv run main.py --platform zhihu --type search --save_data_option sqlite
Skip --init_db for file-based formats (csv, json, jsonl, excel) and MongoDB.
Key Configuration Files
| File | Purpose |
|---|---|
config/base_config.py |
Default SAVE_DATA_OPTION; supported backend list |
config/db_config.py |
Connection credentials for SQLite, MySQL, PostgreSQL, MongoDB |
store/{platform}/__init__.py |
Per-platform storage class registry |
docs/data_storage_guide.md |
User documentation for storage initialization |
main.py |
CLI argument parsing and storage wiring |
Summary
- MediaCrawler offers eight storage backends: csv, json, jsonl, sqlite, mysql, postgres, mongodb, excel
- Configure via
--save_data_optionCLI flag orSAVE_DATA_OPTIONinconfig/base_config.py - SQL databases (SQLite, MySQL, PostgreSQL) require
--init_dbfor schema creation - Connection parameters live in
config/db_config.py - Platform-specific implementations are registered in
store/{platform}/__init__.py
Frequently Asked Questions
How do I switch from SQLite to MySQL in MediaCrawler?
Update SAVE_DATA_OPTION = "mysql" in config/base_config.py, add MySQL credentials to config/db_config.py, run uv run main.py --init_db mysql to create tables, then execute crawls. The --save_data_option mysql CLI flag overrides the config file for single runs.
Does MediaCrawler support Excel export without code changes?
Yes. Pass --save_data_option excel when running main.py. The Excel implementation in store/{platform}/__init__.py handles workbook generation automatically. No initialization step is required.
Where is the complete list of supported storage options defined?
The authoritative list appears in config/base_config.py at lines 89–90: "csv, db, json, jsonl, sqlite, excel, postgres". MongoDB is additionally supported via the "mongodb" key in each platform's storage registry.
What is the difference between --init_db and --save_data_option?
--init_db creates database schemas for SQL backends (sqlite, mysql, postgres) and should run once before crawling. --save_data_option selects where crawl results are written during execution. File-based formats and MongoDB do not require --init_db.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →