# How to Configure MediaCrawler: A Complete Guide to Base, Platform-Specific, and Database Settings

> Configure MediaCrawler effectively with this comprehensive guide. Master base, platform, and database settings without touching source code. Customize your media management now.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**MediaCrawler behavior is controlled entirely through Python configuration modules in the `config/` package, allowing you to customize platforms, login methods, proxies, and storage formats without modifying source code.**

The NanmiCoder/MediaCrawler repository organizes all tunable parameters into modular config files. At runtime, [`config/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/__init__.py) re-exports constants from [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) and [`db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/db_config.py), while platform-specific modules like [`xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/xhs_config.py) handle unique URL lists and switches. This architecture decouples crawler logic from deployment settings.

## Core Configuration ([`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py))

The [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) file contains global constants that affect every platform. These settings dictate target platforms, authentication strategies, network proxies, browser behavior, and output formats.

### Platform Selection and Login Methods

The **`PLATFORM`** constant (line 21) selects your target: `"xhs"` (Xiaohongshu/Rednote), `"dy"` (Douyin), `"ks"` (Kuaishou), `"bili"` (Bilibili), `"wb"` (Weibo), `"tieba"`, or `"zhihu"`.

Authentication is controlled by **`LOGIN_TYPE`** (line 29), supporting `"qrcode"`, `"phone"`, or `"cookie"`. When using cookie-based login, populate the **`COOKIES`** string (line 30). To persist sessions between runs, enable **`SAVE_LOGIN_STATE`** (line 53, default `True`).

For Xiaohongshu specifically, set **`XHS_INTERNATIONAL`** (line 25) to `True` to crawl the overseas version (`rednote.com`) instead of the domestic site.

Define search queries in **`KEYWORDS`** (line 27) as a comma-separated string:

```python

# config/base_config.py

PLATFORM = "xhs"
KEYWORDS = "编程副业,编程兼职"
LOGIN_TYPE = "qrcode"

```

### Proxy and Network Settings

Enable proxy pools by setting **`ENABLE_IP_PROXY`** (line 33) to `True`. Configure the provider via **`IP_PROXY_PROVIDER_NAME`** (line 40), choosing from `"kuaidaili"`, `"wandouhttp"`, or `"static"`.

For static proxies, specify the URL in **`STATIC_PROXY_URL`** (line 44). Parallel proxy pools are controlled by **`IP_PROXY_POOL_COUNT`** (line 37, default `2`).

```python

# config/base_config.py

ENABLE_IP_PROXY = True
IP_PROXY_PROVIDER_NAME = "static"
STATIC_PROXY_URL = "http://myproxy.example.com:3128"

```

### Browser and CDP Configuration

MediaCrawler uses Chrome DevTools Protocol (CDP) for anti-detection. Key settings include:

- **`ENABLE_CDP_MODE`** (line 59): Set to `True` to use real browser automation instead of Playwright (default `True`).
- **`CDP_DEBUG_PORT`** (line 63): CDP communication port (default `9222`).
- **`CDP_CONNECT_EXISTING`** (line 81): Attach to an already-running Chrome instance (default `True`).
- **`CUSTOM_BROWSER_PATH`** (line 69): Absolute path to Chrome/Edge binary; leave empty for auto-detection.
- **`HEADLESS`** (line 50): Run browser without UI (default `False`).
- **`BROWSER_LAUNCH_TIMEOUT`** (line 76): Seconds to wait for browser startup (default `60`).

To connect to an existing Chrome instance:

```bash

# Start Chrome with remote debugging

google-chrome --remote-debugging-port=9222

# Then run the crawler

python main.py --headless false --cdp-connect-existing true

```

### Data Persistence Options

The **`SAVE_DATA_OPTION`** constant (line 90) determines output format: `"csv"`, `"db"`, `"json"`, `"jsonl"`, `"sqlite"`, `"excel"`, or `"postgres"`. The export directory is set via **`SAVE_DATA_PATH`** (line 93); if empty, files write to `data/`.

```python

# config/base_config.py

SAVE_DATA_OPTION = "jsonl"
SAVE_DATA_PATH = "/var/media_crawler/output"

```

## Platform-Specific Configuration

Each platform has dedicated config files under `config/` that define URL lists and specialized toggles.

### Xiaohongshu (XHS) Settings

In [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 26-37), define specific content to crawl:

- **`XHS_SPECIFIED_NOTE_URL_LIST`**: List of note URLs requiring the `xsec_token` query parameter.
- **`XHS_CREATOR_ID_LIST`**: Creator profile URLs for user-centric crawling.

```python

# config/xhs_config.py

XHS_SPECIFIED_NOTE_URL_LIST = [
    "https://www.xiaohongshu.com/explore/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
    "https://www.xiaohongshu.com/explore/64b962a2000000000c0345a9?xsec_token=YOUR_TOKEN",
]

```

### Zhihu, Weibo, and Other Platforms

Platform modules follow identical patterns:

- **Zhihu** ([`config/zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py), lines 24-34): Use `ZHIHU_CREATOR_URL_LIST` and `ZHIHU_SPECIFIED_ID_LIST` for user pages and Q&A content.
- **Weibo** ([`config/weibo_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/weibo_config.py), lines 23-28): Set `WEIBO_SEARCH_TYPE` for strategy selection and `WEIBO_SPECIFIED_ID_LIST` for concrete post IDs.
- **Others** ([`dy_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/dy_config.py), [`ks_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/ks_config.py), [`bilibili_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/bilibili_config.py), [`tieba_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tieba_config.py)): Each exports analogous URL lists and platform-specific constants.

## Database Configuration ([`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py))

Database connectivity for MySQL, PostgreSQL, MongoDB, and SQLite is declared in [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py). This file is imported by [`database/db_session.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db_session.py) and storage factories, ensuring credential changes propagate across the entire persistence layer. Update connection strings and pool settings here to switch between local SQLite and production PostgreSQL instances.

## Command-Line Overrides

The [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) module defines enumerations that map directly to configuration constants, allowing runtime overrides without file edits:

- `PlatformEnum` (line 40): Overrides `PLATFORM`.
- `LoginTypeEnum` (line 52): Overrides `LOGIN_TYPE`.
- `CrawlerTypeEnum` (line 60): Overrides `CRAWLER_TYPE` (`"search"`, `"detail"`, or `"creator"`).
- `SaveDataOptionEnum` (line 68): Overrides `SAVE_DATA_OPTION`.

Example CLI invocation:

```bash
python main.py \
    --platform zhihu \
    --login-type phone \
    --crawler-type creator \
    --save-data-option csv \
    --headless true

```

## Practical Configuration Examples

### Switching to SQLite Output

```python

# config/base_config.py

SAVE_DATA_OPTION = "sqlite"
SAVE_DATA_PATH = "/var/media_crawler/sqlite"

```

### Enabling Douyin with Phone Login

```python

# config/base_config.py

PLATFORM = "dy"
LOGIN_TYPE = "phone"
CRAWLER_TYPE = "search"

```

### Using Environment Variables for Secrets

Create a `.env` file in the project root (requires `python-dotenv`):

```text
COOKIES=your_cookie_string_here
STATIC_PROXY_URL=http://myproxy.example.com:8080

```

Reference these in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) via `os.getenv` to keep credentials out of version control.

## Summary

- **Global settings** reside in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and control platforms, login methods, proxies, CDP behavior, and storage formats.
- **Platform-specific modules** ([`xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/xhs_config.py), [`zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/zhihu_config.py), etc.) manage URL lists and unique parameters.
- **Database credentials** are centralized in [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) and consumed by [`database/db_session.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db_session.py).
- **CLI arguments** defined in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) allow temporary overrides without code changes.
- All constants are re-exported through [`config/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/__init__.py), ensuring the runtime remains decoupled from configuration files.

## Frequently Asked Questions

### How do I switch from Xiaohongshu to Douyin without editing code?

Use the command-line interface: `python main.py --platform dy`. The `PlatformEnum` in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) (line 40) maps the CLI value to the `PLATFORM` constant in [`base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_config.py) at runtime.

### What is the difference between `CRAWLER_TYPE` options?

The **`CRAWLER_TYPE`** setting (line 31) accepts three values: `"search"` for keyword-based crawling, `"detail"` for specific post/note URLs, and `"creator"` for user profile scraping. Each type triggers different logic in the platform-specific crawler implementations.

### Why should I use CDP mode instead of standard Playwright?

**`ENABLE_CDP_MODE`** (line 59, default `True`) uses Chrome DevTools Protocol with a real browser instance, significantly reducing detection rates by anti-bot systems. Standard Playwright operates in a more detectable automation environment.

### Where do I store sensitive data like cookies and proxy URLs?

Place sensitive values in a `.env` file at the project root and reference them via `os.getenv()` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). Alternatively, pass cookies via the **`COOKIES`** constant (line 30) if your deployment environment supports secure configuration injection.