# MediaCrawler Command-Line Arguments: Complete CLI Reference

> Explore MediaCrawler command-line arguments for platform selection, login, crawling scope, storage, and proxy settings. Get the complete CLI reference for efficient media crawling.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: api-reference
- Published: 2026-07-01

---

**MediaCrawler accepts over 20 command-line arguments via a Typer-based interface defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) to control platform selection, login methods, crawling scope, storage backends, and proxy configuration.**

NanmiCoder/MediaCrawler is an open-source Python framework for scraping content from major Chinese social media platforms. The entire application is driven by a single entry point that normalizes raw CLI input through the `parse_cmd` coroutine, updates the global `config` module, and returns a `SimpleNamespace` containing the final runtime configuration.

## Platform and Authentication Arguments

The crawler supports seven social media platforms and three authentication methods, configured via the primary flags defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py).

### Platform Selection (`--platform`)

The `--platform` argument accepts a `PlatformEnum` value to select the target site. Supported platforms include XiaoHongShu (`xhs`), Douyin (`dy`), Kuaishou (`ks`), Bilibili (`bili`), Weibo (`wb`), Tieba (`tieba`), and Zhihu (`zhihu`). This flag is defined at line 61 and determines which crawler implementation is instantiated.

### Login Methods (`--lt` and `--cookies`)

Authentication is controlled by `--lt` (LoginTypeEnum) at line 70, supporting `qrcode`, `phone`, or `cookie` modes. When using cookie-based authentication, pass the raw cookie string via `--cookies` (line 144). For QR code or phone login, the crawler initiates an interactive browser session unless `--headless` is enabled.

## Crawling Mode and Scope

MediaCrawler operates in three distinct modes controlled by the `--type` argument (CrawlerTypeEnum, line 78), each requiring different identifier flags.

### Crawler Types (`--type`)

- **search**: Keyword-based crawling using `--keywords` (line 94). Accepts comma-separated search terms.
- **detail**: Specific post extraction using `--specified_id` (line 152), which accepts comma-separated post IDs or full URLs.
- **creator**: Author-centric crawling using `--creator_id` (line 160), accepting comma-separated creator IDs or profile URLs.

### Pagination and Limits

Control crawl volume using:
- `--start` (line 86): Starting page number for paginated requests.
- `--crawler_max_notes_count` (line 178): Maximum total items (posts/videos) to fetch.
- `--max_comments_count_singlenotes` (line 170): Upper limit of first-level comments per item.
- `--max_concurrency_num` (line 186): Maximum parallel crawler instances for concurrent processing.

## Data Extraction Options

Fine-tune what content is extracted from each target using boolean flags and headless browser settings.

### Comment Harvesting (`--get_comment` and `--get_sub_comment`)

Enable first-level comment extraction with `--get_comment` (line 102) and second-level (reply) comments with `--get_sub_comment` (line 110). Both accept boolean strings (`yes/true/t/y/1` or `no/false/f/n/0`).

### Browser Configuration (`--headless`)

The `--headless` flag (line 118) controls whether Playwright or CDP browsers run in headless mode, essential for running the crawler on servers without GUI support.

## Storage and Output Configuration

MediaCrawler supports multiple output formats and database backends, configured via storage-specific arguments.

### Output Formats (`--save_data_option`)

The `--save_data_option` flag (SaveDataOptionEnum, line 126) selects the destination format:
- **CSV** (`csv`)
- **JSON** (`json`)
- **JSONL** (`jsonl`)
- **SQLite** (`sqlite`)
- **MongoDB** (`mongodb`)
- **PostgreSQL** (`postgres`)
- **Excel** (`excel`)
- **MySQL** (`db`)

### Database Initialization (`--init_db`)

When using database storage, initialize the schema using `--init_db` (InitDbOptionEnum, line 134). This optional flag defaults to `sqlite` when supplied without a value, but supports `mysql` and `postgres` explicit values.

### File Paths (`--save_data_path`)

Specify the output directory with `--save_data_path` (line 194). Defaults to the project's `data/` folder if not provided.

## Proxy and Network Settings

For high-volume crawling, MediaCrawler supports configurable proxy pools to manage IP rotation and rate limiting.

### Proxy Enablement (`--enable_ip_proxy`)

Toggle proxy usage with `--enable_ip_proxy` (line 202), accepting standard boolean string values.

### Proxy Configuration

When proxies are enabled, configure the pool using:
- `--ip_proxy_provider_name` (line 218): Select the provider (`kuaidaili`, `wandouhttp`, or `static`).
- `--ip_proxy_pool_count` (line 210): Number of proxy instances to maintain.
- `--static_proxy_url` (line 226): Full URL for static proxy configuration (e.g., `http://user:pass@host:port`).

## Practical Usage Examples

### Basic Keyword Search on Douyin

Search for keywords and store results as JSONL:

```bash
python -m MediaCrawler.main \
  --platform dy \
  --lt qrcode \
  --type search \
  --keywords "funny cats, travel vlog" \
  --save_data_option jsonl

```

### Detail Mode with Comments

Crawl specific Weibo posts with comment extraction in headless mode:

```bash
python -m MediaCrawler.main \
  --platform wb \
  --lt cookie \
  --type detail \
  --specified_id "https://weibo.com/1234567890,9876543210" \
  --get_comment yes \
  --get_sub_comment no \
  --headless yes \
  --save_data_option csv

```

### Creator Mode with Database Initialization

Initialize SQLite schema before crawling Bilibili creators:

```bash
python -m MediaCrawler.main \
  --platform bili \
  --lt qrcode \
  --type creator \
  --creator_id "u12345,u67890" \
  --init_db \
  --save_data_option sqlite

```

### High-Concurrency Crawling with Static Proxy

Enable static proxy for high-volume Tieba searches:

```bash
python -m MediaCrawler.main \
  --platform tieba \
  --lt phone \
  --type search \
  --keywords "python tutorials" \
  --max_concurrency_num 12 \
  --enable_ip_proxy yes \
  --ip_proxy_provider_name static \
  --static_proxy_url "http://user:pass@proxy.example.com:8080"

```

## Summary

- **Entry Point**: All arguments are defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) and parsed by the `parse_cmd` coroutine.
- **Platform Support**: Seven platforms (XiaoHongShu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu) via `--platform`.
- **Authentication**: Flexible login via `--lt` (qrcode/phone/cookie) with raw cookie support via `--cookies`.
- **Crawl Modes**: Three distinct types (search, detail, creator) with specific ID and keyword flags.
- **Data Control**: Granular limits for items, comments, and concurrency levels.
- **Storage**: Eight output formats including SQL databases, NoSQL, and flat files.
- **Networking**: Built-in proxy support with configurable providers and static URL options.

## Frequently Asked Questions

### What is the default login method if I don't specify `--lt`?

The crawler requires explicit authentication configuration. According to the source code in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py), the `--lt` flag has no default value and must be provided as either `qrcode`, `phone`, or `cookie`. Attempting to run without this argument will trigger a Typer validation error.

### Can I use multiple keywords in a single search command?

Yes. The `--keywords` argument accepts a comma-separated string of search terms. For example, `--keywords "python, golang, rust"` will search for all three terms sequentially. This parameter is consumed by the search crawler type defined in the base crawler implementation.

### How do I store data in PostgreSQL instead of the default format?

Specify `--save_data_option postgres` and initialize the database schema using `--init_db postgres` before running the crawler. The PostgreSQL implementation is located in `store/*/_store_impl.py`, which handles connection pooling and table creation based on the global configuration updated by these CLI flags.