# How to Configure MediaCrawler Settings: Complete Guide to Python Config Files and CLI Options

> Learn how to configure MediaCrawler settings easily by editing Python config files or using CLI options. Customize your media crawling experience without altering source code.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-12

---

**Configure MediaCrawler by editing Python modules in the `config/` directory—primarily [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) for global settings and platform-specific files like [`xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/xhs_config.py) for URL lists—without modifying any source code.**

MediaCrawler is an open-source, multi-platform social media scraper that centralizes all behavior in configuration files located under `config/`. As implemented in NanmiCoder/MediaCrawler, the framework imports these modules at runtime through [`config/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/__init__.py), which re-exports every constant from [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) and [`db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/db_config.py). This architecture lets you customize crawling behavior, authentication, proxies, and data output formats by changing values in a few Python files.

## Core Global Configuration (config/base_config.py)

The [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) file controls platform selection, authentication, browser behavior, proxy settings, and data persistence. All constants are exposed via `from .base_config import *` in [`config/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/__init__.py), making them available project-wide.

### Platform and Crawl Target Settings

| Setting | Description | Default Value |
|---------|-------------|---------------|
| **`PLATFORM`** | Target platform: `xhs`, `dy`, `ks`, `bili`, `wb`, `tieba`, `zhihu` | `"xhs"` |
| **`XHS_INTERNATIONAL`** | Use overseas Xiaohongshu (`rednote.com`) | `False` |
| **`KEYWORDS`** | Comma-separated search terms | `"编程副业,编程兼职"` |
| **`CRAWLER_TYPE`** | Crawl mode: `search`, `detail`, or `creator` | `"search"` |

In [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) line 21, `PLATFORM` defaults to `"xhs"` but can be switched to any supported platform identifier.

### Authentication Configuration

MediaCrawler supports three login methods configured via `LOGIN_TYPE` (line 29):

- **`qrcode`** – Scan QR code with mobile app (default)
- **`phone`** – SMS verification login
- **`cookie`** – Use existing session via `COOKIES` string (line 30)

```python

# config/base_config.py

LOGIN_TYPE = "cookie"
COOKIES = "acw_tc=xxx; session_id=yyy; ..."  # Paste full cookie string

```

When `SAVE_LOGIN_STATE = True` (line 53), the crawler persists cookies between runs to avoid repeated authentication.

### Browser and CDP Settings

The framework uses Chrome DevTools Protocol (CDP) for anti-detection. Key settings from lines 59-86:

```python

# config/base_config.py

ENABLE_CDP_MODE = True          # Use real browser instead of Playwright

CDP_DEBUG_PORT = 9222           # Chrome remote debugging port

HEADLESS = False                # Show browser UI (False) or hide it (True)

CDP_HEADLESS = False            # Headless inside CDP (may trigger detection)

CDP_CONNECT_EXISTING = True     # Attach to already-running Chrome

CUSTOM_BROWSER_PATH = ""        # Path to Chrome/Edge binary (empty = auto-detect)

BROWSER_LAUNCH_TIMEOUT = 60     # Seconds to wait for browser startup

AUTO_CLOSE_BROWSER = True       # Close browser when script ends

```

To connect to an existing Chrome instance, start it with remote debugging and match the port:

```bash
google-chrome --remote-debugging-port=9222

```

### Proxy Configuration

Enable IP rotation via `ENABLE_IP_PROXY` (line 33) with three provider options:

```python

# config/base_config.py

ENABLE_IP_PROXY = True
IP_PROXY_POOL_COUNT = 2
IP_PROXY_PROVIDER_NAME = "kuaidaili"  # Alternatives: "wandouhttp", "static"

# For static proxy:

# IP_PROXY_PROVIDER_NAME = "static"

# STATIC_PROXY_URL = "http://myproxy.example.com:3128"

```

### Data Output Format

Control persistence via `SAVE_DATA_OPTION` (line 90):

```python

# config/base_config.py

SAVE_DATA_OPTION = "jsonl"      # Options: csv, db, json, jsonl, sqlite, excel, postgres

SAVE_DATA_PATH = ""             # Empty defaults to data/ directory

```

The [`store_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store_factory.py) module (in `store/`) instantiates the appropriate storage class based on this constant.

## Platform-Specific Configuration

Each supported platform has dedicated config files for URL lists and special parameters. These are imported alongside [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) at runtime.

### Xiaohongshu (config/xhs_config.py)

Define specific notes or creators to crawl when `CRAWLER_TYPE` is `detail` or `creator`:

```python

# config/xhs_config.py

XHS_SPECIFIED_NOTE_URL_LIST = [
    "https://www.xiaohongshu.com/explore/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
]

XHS_CREATOR_ID_LIST = [
    "https://www.xiaohongshu.com/user/profile/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
]

```

URLs **must include** the `xsec_token` query parameter for authentication validation (lines 26-37).

### Zhihu (config/zhihu_config.py)

Configure user profiles, questions, answers, and videos:

```python

# config/zhihu_config.py

ZHIHU_CREATOR_URL_LIST = [
    "https://www.zhihu.com/people/example-user",
]

ZHIHU_SPECIFIED_ID_LIST = [
    "https://www.zhihu.com/question/123456789",
    "https://www.zhihu.com/zvideo/987654321",
]

```

Other platforms follow identical patterns: [`dy_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/dy_config.py) (Douyin), [`ks_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/ks_config.py) (Kuaishou), [`bilibili_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/bilibili_config.py), [`tieba_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tieba_config.py), and [`weibo_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/weibo_config.py).

## Command-Line Overrides

The [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) module defines enumerations that map directly to configuration constants, enabling runtime overrides without code changes:

```bash
python main.py \
    --platform zhihu \
    --login-type phone \
    --crawler-type creator \
    --save-data-option csv \
    --headless true

```

Available enum mappings (lines 40-68):

- `PlatformEnum` → `PLATFORM`
- `LoginTypeEnum` → `LOGIN_TYPE`
- `CrawlerTypeEnum` → `CRAWLER_TYPE`
- `SaveDataOptionEnum` → `SAVE_DATA_OPTION`

CLI arguments take precedence over file-based configuration.

## Database Configuration (config/db_config.py)

Database connectivity for MySQL, PostgreSQL, SQLite, and MongoDB is declared in [`db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/db_config.py). This file is imported by [`database/db_session.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db_session.py) and storage factories, so credential changes propagate across all persistence operations.

## Environment Variables for Secrets

Sensitive values can be externalized to a `.env` file (optional, requires `python-dotenv`):

```bash

# .env in project root

COOKIES=your_session_cookie_string
STATIC_PROXY_URL=http://auth:pass@proxy.example.com:8080

```

Reference these in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py):

```python
import os
COOKIES = os.getenv("COOKIES", "")

```

## Configuration Execution Flow

When you run `python main.py`, the following sequence executes (as documented in `docs/项目架构文档.md`):

1. **Entry point** ([`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)) parses CLI arguments
2. **Config import** loads `config/` package → [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py), [`db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/db_config.py), platform configs
3. **CrawlerFactory** (`media_platform/`) instantiates platform-specific crawler using `PLATFORM`
4. **Crawler execution** uses authentication, proxy, and browser settings from config
5. **StoreFactory** (`store/`) writes data according to `SAVE_DATA_OPTION`

No hard-coded values remain after import—all behavior is externally configurable.

## Practical Configuration Examples

### Switch to Douyin with Cookie Login

```python

# config/base_config.py

PLATFORM = "dy"
LOGIN_TYPE = "cookie"
COOKIES = "sessionid=xxx; sid_guard=xxx"
CRAWLER_TYPE = "search"
KEYWORDS = "美食,旅游"

```

### Enable SQLite Output with Custom Path

```python

# config/base_config.py

SAVE_DATA_OPTION = "sqlite"
SAVE_DATA_PATH = "/var/media_crawler/db"

```

### Crawl Specific Weibo Posts

```python

# config/weibo_config.py

WEIBO_SEARCH_TYPE = "1"  # Comprehensive search

WEIBO_SPECIFIED_ID_LIST = ["4892012345678901", "4892012345678902"]

```

### Connect to Remote Debug Chrome

```bash

# Terminal 1: Start Chrome

/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome \
    --remote-debugging-port=9222 \
    --user-data-dir=/tmp/chrome_debug_profile

# Terminal 2: Run crawler

python main.py --cdp-connect-existing true --headless false

```

## Summary

- **All configuration lives in `config/`** — edit [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) for global settings and `{platform}_config.py` for URL lists
- **Runtime overrides** via CLI flags defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) take precedence over file values
- **Authentication** supports QR code, phone, or cookie-based login with optional persistence
- **Proxy support** includes Kuaidaili, Wandou, or static proxy configurations
- **Data output** spans 7 formats: `csv`, `json`, `jsonl`, `sqlite`, `excel`, `postgres`, `db`
- **CDP mode** uses real Chrome/Edge for anti-detection with configurable headless and connection options

## Frequently Asked Questions

### Can I configure MediaCrawler without editing Python files?

Yes. Use command-line arguments to override any [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) setting. For example: `python main.py --platform dy --login-type cookie --save-data-option csv`. CLI flags map directly to configuration constants via the enums in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py).

### Where do I set cookies for authentication?

Paste the complete cookie string into `COOKIES` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) (line 30), or use the `COOKIES` environment variable in a `.env` file. Then set `LOGIN_TYPE = "cookie"`. The crawler will use this session for authenticated requests.

### Why does my Xiaohongshu URL fail with authentication errors?

Xiaohongshu URLs require the `xsec_token` query parameter for validation. Ensure every URL in `XHS_SPECIFIED_NOTE_URL_LIST` or `XHS_CREATOR_ID_LIST` includes this token, which you can extract from your browser's URL bar when viewing the content.

### How do I switch from JSONL to database storage?

Change `SAVE_DATA_OPTION` to your preferred format: `sqlite`, `postgres`, `db` (MySQL), or `mongo`. Update [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) with connection credentials if using a server database. The `StoreFactory` automatically instantiates the correct storage class.