MediaCrawler Command-Line Arguments: Complete CLI Reference
MediaCrawler accepts over 20 command-line arguments via a Typer-based interface defined in cmd_arg/arg.py to control platform selection, login methods, crawling scope, storage backends, and proxy configuration.
NanmiCoder/MediaCrawler is an open-source Python framework for scraping content from major Chinese social media platforms. The entire application is driven by a single entry point that normalizes raw CLI input through the parse_cmd coroutine, updates the global config module, and returns a SimpleNamespace containing the final runtime configuration.
Platform and Authentication Arguments
The crawler supports seven social media platforms and three authentication methods, configured via the primary flags defined in cmd_arg/arg.py.
Platform Selection (--platform)
The --platform argument accepts a PlatformEnum value to select the target site. Supported platforms include XiaoHongShu (xhs), Douyin (dy), Kuaishou (ks), Bilibili (bili), Weibo (wb), Tieba (tieba), and Zhihu (zhihu). This flag is defined at line 61 and determines which crawler implementation is instantiated.
Login Methods (--lt and --cookies)
Authentication is controlled by --lt (LoginTypeEnum) at line 70, supporting qrcode, phone, or cookie modes. When using cookie-based authentication, pass the raw cookie string via --cookies (line 144). For QR code or phone login, the crawler initiates an interactive browser session unless --headless is enabled.
Crawling Mode and Scope
MediaCrawler operates in three distinct modes controlled by the --type argument (CrawlerTypeEnum, line 78), each requiring different identifier flags.
Crawler Types (--type)
- search: Keyword-based crawling using
--keywords(line 94). Accepts comma-separated search terms. - detail: Specific post extraction using
--specified_id(line 152), which accepts comma-separated post IDs or full URLs. - creator: Author-centric crawling using
--creator_id(line 160), accepting comma-separated creator IDs or profile URLs.
Pagination and Limits
Control crawl volume using:
--start(line 86): Starting page number for paginated requests.--crawler_max_notes_count(line 178): Maximum total items (posts/videos) to fetch.--max_comments_count_singlenotes(line 170): Upper limit of first-level comments per item.--max_concurrency_num(line 186): Maximum parallel crawler instances for concurrent processing.
Data Extraction Options
Fine-tune what content is extracted from each target using boolean flags and headless browser settings.
Comment Harvesting (--get_comment and --get_sub_comment)
Enable first-level comment extraction with --get_comment (line 102) and second-level (reply) comments with --get_sub_comment (line 110). Both accept boolean strings (yes/true/t/y/1 or no/false/f/n/0).
Browser Configuration (--headless)
The --headless flag (line 118) controls whether Playwright or CDP browsers run in headless mode, essential for running the crawler on servers without GUI support.
Storage and Output Configuration
MediaCrawler supports multiple output formats and database backends, configured via storage-specific arguments.
Output Formats (--save_data_option)
The --save_data_option flag (SaveDataOptionEnum, line 126) selects the destination format:
- CSV (
csv) - JSON (
json) - JSONL (
jsonl) - SQLite (
sqlite) - MongoDB (
mongodb) - PostgreSQL (
postgres) - Excel (
excel) - MySQL (
db)
Database Initialization (--init_db)
When using database storage, initialize the schema using --init_db (InitDbOptionEnum, line 134). This optional flag defaults to sqlite when supplied without a value, but supports mysql and postgres explicit values.
File Paths (--save_data_path)
Specify the output directory with --save_data_path (line 194). Defaults to the project's data/ folder if not provided.
Proxy and Network Settings
For high-volume crawling, MediaCrawler supports configurable proxy pools to manage IP rotation and rate limiting.
Proxy Enablement (--enable_ip_proxy)
Toggle proxy usage with --enable_ip_proxy (line 202), accepting standard boolean string values.
Proxy Configuration
When proxies are enabled, configure the pool using:
--ip_proxy_provider_name(line 218): Select the provider (kuaidaili,wandouhttp, orstatic).--ip_proxy_pool_count(line 210): Number of proxy instances to maintain.--static_proxy_url(line 226): Full URL for static proxy configuration (e.g.,http://user:pass@host:port).
Practical Usage Examples
Basic Keyword Search on Douyin
Search for keywords and store results as JSONL:
python -m MediaCrawler.main \
--platform dy \
--lt qrcode \
--type search \
--keywords "funny cats, travel vlog" \
--save_data_option jsonl
Detail Mode with Comments
Crawl specific Weibo posts with comment extraction in headless mode:
python -m MediaCrawler.main \
--platform wb \
--lt cookie \
--type detail \
--specified_id "https://weibo.com/1234567890,9876543210" \
--get_comment yes \
--get_sub_comment no \
--headless yes \
--save_data_option csv
Creator Mode with Database Initialization
Initialize SQLite schema before crawling Bilibili creators:
python -m MediaCrawler.main \
--platform bili \
--lt qrcode \
--type creator \
--creator_id "u12345,u67890" \
--init_db \
--save_data_option sqlite
High-Concurrency Crawling with Static Proxy
Enable static proxy for high-volume Tieba searches:
python -m MediaCrawler.main \
--platform tieba \
--lt phone \
--type search \
--keywords "python tutorials" \
--max_concurrency_num 12 \
--enable_ip_proxy yes \
--ip_proxy_provider_name static \
--static_proxy_url "http://user:pass@proxy.example.com:8080"
Summary
- Entry Point: All arguments are defined in
cmd_arg/arg.pyand parsed by theparse_cmdcoroutine. - Platform Support: Seven platforms (XiaoHongShu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu) via
--platform. - Authentication: Flexible login via
--lt(qrcode/phone/cookie) with raw cookie support via--cookies. - Crawl Modes: Three distinct types (search, detail, creator) with specific ID and keyword flags.
- Data Control: Granular limits for items, comments, and concurrency levels.
- Storage: Eight output formats including SQL databases, NoSQL, and flat files.
- Networking: Built-in proxy support with configurable providers and static URL options.
Frequently Asked Questions
What is the default login method if I don't specify --lt?
The crawler requires explicit authentication configuration. According to the source code in cmd_arg/arg.py, the --lt flag has no default value and must be provided as either qrcode, phone, or cookie. Attempting to run without this argument will trigger a Typer validation error.
Can I use multiple keywords in a single search command?
Yes. The --keywords argument accepts a comma-separated string of search terms. For example, --keywords "python, golang, rust" will search for all three terms sequentially. This parameter is consumed by the search crawler type defined in the base crawler implementation.
How do I store data in PostgreSQL instead of the default format?
Specify --save_data_option postgres and initialize the database schema using --init_db postgres before running the crawler. The PostgreSQL implementation is located in store/*/_store_impl.py, which handles connection pooling and table creation based on the global configuration updated by these CLI flags.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →