MediaCrawler Data Storage Backends: JSONL, SQLite, MySQL, and PostgreSQL Configuration Guide
MediaCrawler supports four primary data storage backends—JSONL (line-delimited JSON), SQLite, MySQL, and PostgreSQL—configurable via the SAVE_DATA_OPTION setting in config/base_config.py and implemented through platform-specific factory classes in the store/ directory.
The open-source MediaCrawler repository by NanmiCoder provides flexible data persistence options for scraped content across multiple social media platforms. Understanding these MediaCrawler data storage backends is essential for configuring the appropriate persistence layer based on your infrastructure requirements, query complexity, and scalability needs.
Supported Storage Backends in MediaCrawler
MediaCrawler implements a pluggable storage architecture where each platform (Zhihu, XiaoHongShu, Weibo, etc.) registers concrete store implementations through factory classes. The system supports four primary backends for data persistence.
JSONL (Line-Delimited JSON)
JSONL serves as the default storage backend. When SAVE_DATA_OPTION = "jsonl" is set in config/base_config.py, MediaCrawler persists scraped items as newline-delimited JSON records.
- Configuration value:
"jsonl" - Implementation class:
*JsonlStoreImplement(e.g.,ZhihuJsonlStoreImplement,XhsJsonlStoreImplement) - Source location:
store/<platform>/_store_impl.py
This backend requires no external database infrastructure and creates human-readable files that append one JSON object per line.
SQLite
SQLite provides a serverless, file-based relational database option suitable for local development and moderate data volumes.
- Configuration value:
"sqlite" - Implementation class:
*SqliteStoreImplement(e.g.,ZhihuSqliteStoreImplement) - Storage: Local
.dbfiles managed through SQLAlchemy ORM
The SQLite backend automatically handles table creation and schema management through the models defined in database/models.py.
MySQL and PostgreSQL
Both MySQL and PostgreSQL are supported through a unified database abstraction layer. Unlike SQLite, these require running database server instances and connection URLs.
- Configuration values:
"db"(MySQL) or"postgres"(PostgreSQL) - Implementation class:
*DbStoreImplement(e.g.,WeiboDbStoreImplement,XhsDbStoreImplement) - Engine selection: The
database/db_session.pymodule selects the appropriate SQLAlchemy driver based on the database URL supplied via environment variables (MYSQL_URLorPOSTGRES_URL)
Both backends map to the same implementation class, with the specific dialect determined at runtime from the connection string.
Configuration and Implementation Architecture
Central Configuration
The supported storage options are defined centrally in config/base_config.py. According to the source code at lines 89-90, the available options include CSV, DB, JSON, JSONL, SQLite, Excel, and Postgres:
# Data saving type option configuration, supports: csv, db, json, jsonl, sqlite, excel, postgres.
SAVE_DATA_OPTION = "jsonl"
👉 View source: config/base_config.py#L89-L90
Factory Pattern Registration
Each platform implements a store factory that maps configuration strings to concrete implementation classes. For example, the Zhihu factory in store/zhihu/__init__.py (lines 38-48) registers all available backends:
class ZhihuStoreFactory:
STORES = {
"csv": ZhihuCsvStoreImplement,
"db": ZhihuDbStoreImplement,
"postgres": ZhihuDbStoreImplement,
"json": ZhihuJsonStoreImplement,
"jsonl": ZhihuJsonlStoreImplement,
"sqlite": ZhihuSqliteStoreImplement,
"mongodb": ZhihuMongoStoreImplement,
"excel": ZhihuExcelStoreImplement,
}
👉 View source: store/zhihu/init.py#L38-L48
Similar factories exist for other platforms including store/xhs/__init__.py, store/weibo/__init__.py, store/tieba/__init__.py, store/kuaishou/__init__.py, store/douyin/__init__.py, and store/bilibili/__init__.py.
Database Session Management
The database/db_session.py file handles SQLAlchemy engine creation and session management. It supports SQLite, MySQL, and PostgreSQL by parsing the database URL and selecting the appropriate driver and dialect.
Code Examples
Persisting Data to JSONL (Default)
# Example: Persist Zhihu content using the JSONL backend (default)
from store.zhihu import ZhihuStoreFactory
from model.m_zhihu import ZhihuContent
async def save_content(item: ZhihuContent):
# The factory returns a JsonlStoreImplement because SAVE_DATA_OPTION = "jsonl"
store = ZhihuStoreFactory.create_store()
await store.store_content(item.model_dump())
Switching to SQLite Storage
# Example: Switch to SQLite (set the config flag)
import config
config.SAVE_DATA_OPTION = "sqlite" # change at runtime or via env
from store.xhs import XhsStoreFactory
from model.m_xiaohongshu import XhsContent
async def save_to_sqlite(item: XhsContent):
store = XhsStoreFactory.create_store() # returns XhsSqliteStoreImplement
await store.store_content(item.model_dump())
Using MySQL or PostgreSQL
# Example: Use MySQL / PostgreSQL (both map to the generic DB store)
import config
config.SAVE_DATA_OPTION = "db" # or "postgres"
from store.weibo import WeiboStoreFactory
from model.m_weibo import WeiboContent
async def save_to_db(item: WeiboContent):
store = WeiboStoreFactory.create_store() # returns WeiboDbStoreImplement
await store.store_content(item.model_dump())
Key Implementation Files
| File | Role |
|---|---|
config/base_config.py |
Central configuration, lists all supported storage options and defaults to JSONL. |
store/<platform>/__init__.py (e.g., store/zhihu/__init__.py) |
Factory that maps SAVE_DATA_OPTION strings to concrete store implementations. |
store/<platform>/_store_impl.py |
Contains the actual implementation classes (*JsonlStoreImplement, *SqliteStoreImplement, *DbStoreImplement, etc.). |
database/db_session.py |
Handles creation of SQLAlchemy engines for SQLite, MySQL, and PostgreSQL. |
database/models.py |
ORM models used by the DB store implementations. |
Summary
- JSONL is the default storage backend, requiring no external dependencies and storing data as newline-delimited JSON files.
- SQLite provides a file-based relational database option, implemented via
*SqliteStoreImplementclasses. - MySQL and PostgreSQL share the same implementation (
*DbStoreImplement) and are selected using"db"or"postgres"configuration values, with the driver determined by the connection URL. - Configuration is centralized in
config/base_config.pyvia theSAVE_DATA_OPTIONvariable, while platform-specific factories instore/<platform>/__init__.pyhandle instantiation of the correct storage class.
Frequently Asked Questions
What is the default storage backend in MediaCrawler?
JSONL is the default. According to the source code in config/base_config.py line 90, the SAVE_DATA_OPTION variable is set to "jsonl" by default, which triggers the *JsonlStoreImplement classes across all platform stores. This backend writes each scraped item as a single line of JSON, making it ideal for logging and debugging without requiring database setup.
Can I use PostgreSQL with MediaCrawler?
Yes. Set SAVE_DATA_OPTION = "postgres" in your configuration. This maps to the same *DbStoreImplement class used for MySQL, but the underlying SQLAlchemy engine selects the PostgreSQL dialect based on your POSTGRES_URL environment variable. The database session manager in database/db_session.py handles the specific driver selection automatically.
How do I switch from JSONL to SQLite at runtime?
Modify the SAVE_DATA_OPTION variable before importing the store factory. You can either change the value in config/base_config.py permanently or override it at runtime as shown in the code examples above. Once set to "sqlite", calling StoreFactory.create_store() will return a *SqliteStoreImplement instance that persists data to a local .db file using the SQLAlchemy ORM models defined in database/models.py.
Where are the database models defined for relational storage?
In database/models.py. This file contains the SQLAlchemy ORM class definitions used by both SQLite and the shared MySQL/PostgreSQL implementations (*DbStoreImplement). When using any relational backend, these models define the table schemas for scraped content from platforms like Zhihu, Weibo, XiaoHongShu, and others.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →