What Data Formats Does MediaCrawler Output? Complete Guide to File and Database Export Options
MediaCrawler outputs data to seven formats: CSV, JSON, JSONL (default), Excel, SQLite, MySQL, and PostgreSQL, configured via the --save_data_option CLI flag.
MediaCrawler is a popular open-source scraping framework for social media platforms. Understanding its data output formats is essential for integrating scraped content into your analytics pipeline or downstream workflows. This guide covers all supported serialization options based on the source code analysis of NanmiCoder/MediaCrawler.
Flat File Output Formats
MediaCrawler supports four file-based output formats for local storage. These are ideal for quick exports, version control, or feeding data into spreadsheet tools.
JSONL (Default Format)
JSON Lines is the default output format when no --save_data_option is specified. Each scraped record writes as a single JSON object per line, enabling efficient append operations without rewriting entire files.
# JSONL output (default behavior)
uv run main.py --platform xhs --lt qrcode --type search
Output files land in data/*.jsonl with one valid JSON object per line.
JSON (Single Document Mode)
For a traditional JSON array structure, specify --save_data_option json. This wraps all records into a single JSON array, suitable for JavaScript-based consumers or API responses.
# Single JSON document export
uv run main.py --platform xhs --lt qrcode --type search --save_data_option json
File location: data/*.json.
CSV
Comma-separated values work seamlessly with Excel, pandas, and BI tools. MediaCrawler flattens nested structures into CSV columns.
# CSV export
uv run main.py --platform xhs --lt qrcode --type search --save_data_option csv
File location: data/*.csv.
Excel Workbooks
The Excel exporter generates .xlsx files with multiple styled sheets—typically separating content, comments, and creator metadata. This is the preferred format for stakeholder reports or manual review.
# Excel with styled multi-sheet output
uv run main.py --platform xhs --lt qrcode --type search --save_data_option excel
File location: data/*.xlsx.
Database Output Formats
For production pipelines, MediaCrawler supports three relational database engines. Database output requires a two-step process: initialization, then selection.
SQLite (Local File Database)
SQLite requires no external server—data persists to a local .sqlite file ideal for single-node deployments or development environments.
# Initialize SQLite database (one-time)
uv run main.py --init_db sqlite
# Store results in SQLite
uv run main.py --platform xhs --lt qrcode --type search --save_data_option sqlite
MySQL (Remote Server)
For multi-user or high-throughput scenarios, MySQL persists data to a remote server.
# Initialize MySQL connection
uv run main.py --init_db mysql
# Store results (historical flag uses "db")
uv run main.py --platform xhs --lt qrcode --type search --save_data_option db
PostgreSQL (Enterprise Grade)
PostgreSQL support is available for advanced analytics workloads requiring complex queries or JSONB storage.
# Initialize PostgreSQL connection
uv run main.py --init_db postgres
# Store results in PostgreSQL
uv run main.py --platform xhs --lt qrcode --type search --save_data_option postgres
Key Source Files and Implementation Details
The output format system is implemented across these critical paths:
docs/data_storage_guide.md— Documents all supported formats and CLI flags per lines 10-28README.md— Summarizes storage capabilities in the project overview (lines 82-86)main.py— Entry point parsing--save_data_optionand--init_dbargumentsdatabase/*— Contains SQLite, MySQL, and PostgreSQL persistence layer implementationsutils/exporter.py— Handles file-based serialization to CSV, JSON, JSONL, and Excel
According to the data_storage_guide.md source, the --save_data_option flag accepts: csv, json, jsonl, excel, sqlite, db (MySQL), or postgres. The README confirms: "MediaCrawler 支持多种数据存储方式,包括 CSV、JSON、JSONL、Excel、SQLite 和 MySQL 数据库."
Choosing the Right Output Format
| Use Case | Recommended Format | Reason |
|---|---|---|
| Log streaming, incremental processing | JSONL | Append-friendly, minimal overhead |
| API integration, JavaScript consumers | JSON | Native array structure |
| Excel analysis, stakeholder reports | Excel | Multi-sheet, styled formatting |
| Data science with pandas | CSV or JSONL | Direct read_csv() or read_json(lines=True) |
| Lightweight local storage | SQLite | Serverless, file-based, SQL queryable |
| Production multi-node deployment | MySQL or PostgreSQL | Concurrent access, ACID compliance |
Summary
- MediaCrawler outputs to CSV, JSON, JSONL (default), Excel, SQLite, MySQL, and PostgreSQL
- File formats use
--save_data_option <format>directly:csv,json,jsonl,excel - Database formats require
--init_db <engine>first, then--save_data_option <engine> - JSONL is the default when no flag is provided—optimal for append-heavy scraping
- Excel export includes styled multi-sheet workbooks for content, comments, and creators
- Source documentation lives in
docs/data_storage_guide.mdandREADME.mdlines 82-86
Frequently Asked Questions
How do I change the default output format from JSONL to CSV?
Run your command with --save_data_option csv. The default JSONL behavior activates only when this flag is omitted entirely. All flat-file options bypass database initialization.
Can I export to multiple formats in a single run?
No—MediaCrawler processes one --save_data_option at a time per the main.py argument parser. To generate multiple formats, execute separate runs or post-process JSONL output through conversion tools.
Where does SQLite store the database file?
The SQLite implementation in database/* creates the file in your working directory with a .sqlite extension. The exact path depends on your configuration; check database/sqlite.py for the default naming convention.
Is PostgreSQL supported in all MediaCrawler versions?
PostgreSQL support was added after MySQL. Verify your installation includes recent database/postgres.py code if --save_data_option postgres fails—older releases may lack this exporter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →