How to Configure MediaCrawler: A Complete Guide to Base, Platform-Specific, and Database Settings

MediaCrawler behavior is controlled entirely through Python configuration modules in the config/ package, allowing you to customize platforms, login methods, proxies, and storage formats without modifying source code.

The NanmiCoder/MediaCrawler repository organizes all tunable parameters into modular config files. At runtime, config/__init__.py re-exports constants from base_config.py and db_config.py, while platform-specific modules like xhs_config.py handle unique URL lists and switches. This architecture decouples crawler logic from deployment settings.

Core Configuration (config/base_config.py)

The config/base_config.py file contains global constants that affect every platform. These settings dictate target platforms, authentication strategies, network proxies, browser behavior, and output formats.

Platform Selection and Login Methods

The PLATFORM constant (line 21) selects your target: "xhs" (Xiaohongshu/Rednote), "dy" (Douyin), "ks" (Kuaishou), "bili" (Bilibili), "wb" (Weibo), "tieba", or "zhihu".

Authentication is controlled by LOGIN_TYPE (line 29), supporting "qrcode", "phone", or "cookie". When using cookie-based login, populate the COOKIES string (line 30). To persist sessions between runs, enable SAVE_LOGIN_STATE (line 53, default True).

For Xiaohongshu specifically, set XHS_INTERNATIONAL (line 25) to True to crawl the overseas version (rednote.com) instead of the domestic site.

Define search queries in KEYWORDS (line 27) as a comma-separated string:


# config/base_config.py

PLATFORM = "xhs"
KEYWORDS = "编程副业,编程兼职"
LOGIN_TYPE = "qrcode"

Proxy and Network Settings

Enable proxy pools by setting ENABLE_IP_PROXY (line 33) to True. Configure the provider via IP_PROXY_PROVIDER_NAME (line 40), choosing from "kuaidaili", "wandouhttp", or "static".

For static proxies, specify the URL in STATIC_PROXY_URL (line 44). Parallel proxy pools are controlled by IP_PROXY_POOL_COUNT (line 37, default 2).


# config/base_config.py

ENABLE_IP_PROXY = True
IP_PROXY_PROVIDER_NAME = "static"
STATIC_PROXY_URL = "http://myproxy.example.com:3128"

Browser and CDP Configuration

MediaCrawler uses Chrome DevTools Protocol (CDP) for anti-detection. Key settings include:

  • ENABLE_CDP_MODE (line 59): Set to True to use real browser automation instead of Playwright (default True).
  • CDP_DEBUG_PORT (line 63): CDP communication port (default 9222).
  • CDP_CONNECT_EXISTING (line 81): Attach to an already-running Chrome instance (default True).
  • CUSTOM_BROWSER_PATH (line 69): Absolute path to Chrome/Edge binary; leave empty for auto-detection.
  • HEADLESS (line 50): Run browser without UI (default False).
  • BROWSER_LAUNCH_TIMEOUT (line 76): Seconds to wait for browser startup (default 60).

To connect to an existing Chrome instance:


# Start Chrome with remote debugging

google-chrome --remote-debugging-port=9222

# Then run the crawler

python main.py --headless false --cdp-connect-existing true

Data Persistence Options

The SAVE_DATA_OPTION constant (line 90) determines output format: "csv", "db", "json", "jsonl", "sqlite", "excel", or "postgres". The export directory is set via SAVE_DATA_PATH (line 93); if empty, files write to data/.


# config/base_config.py

SAVE_DATA_OPTION = "jsonl"
SAVE_DATA_PATH = "/var/media_crawler/output"

Platform-Specific Configuration

Each platform has dedicated config files under config/ that define URL lists and specialized toggles.

Xiaohongshu (XHS) Settings

In config/xhs_config.py (lines 26-37), define specific content to crawl:

  • XHS_SPECIFIED_NOTE_URL_LIST: List of note URLs requiring the xsec_token query parameter.
  • XHS_CREATOR_ID_LIST: Creator profile URLs for user-centric crawling.

# config/xhs_config.py

XHS_SPECIFIED_NOTE_URL_LIST = [
    "https://www.xiaohongshu.com/explore/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
    "https://www.xiaohongshu.com/explore/64b962a2000000000c0345a9?xsec_token=YOUR_TOKEN",
]

Zhihu, Weibo, and Other Platforms

Platform modules follow identical patterns:

Database Configuration (config/db_config.py)

Database connectivity for MySQL, PostgreSQL, MongoDB, and SQLite is declared in config/db_config.py. This file is imported by database/db_session.py and storage factories, ensuring credential changes propagate across the entire persistence layer. Update connection strings and pool settings here to switch between local SQLite and production PostgreSQL instances.

Command-Line Overrides

The cmd_arg/arg.py module defines enumerations that map directly to configuration constants, allowing runtime overrides without file edits:

  • PlatformEnum (line 40): Overrides PLATFORM.
  • LoginTypeEnum (line 52): Overrides LOGIN_TYPE.
  • CrawlerTypeEnum (line 60): Overrides CRAWLER_TYPE ("search", "detail", or "creator").
  • SaveDataOptionEnum (line 68): Overrides SAVE_DATA_OPTION.

Example CLI invocation:

python main.py \
    --platform zhihu \
    --login-type phone \
    --crawler-type creator \
    --save-data-option csv \
    --headless true

Practical Configuration Examples

Switching to SQLite Output


# config/base_config.py

SAVE_DATA_OPTION = "sqlite"
SAVE_DATA_PATH = "/var/media_crawler/sqlite"

Enabling Douyin with Phone Login


# config/base_config.py

PLATFORM = "dy"
LOGIN_TYPE = "phone"
CRAWLER_TYPE = "search"

Using Environment Variables for Secrets

Create a .env file in the project root (requires python-dotenv):

COOKIES=your_cookie_string_here
STATIC_PROXY_URL=http://myproxy.example.com:8080

Reference these in base_config.py via os.getenv to keep credentials out of version control.

Summary

Frequently Asked Questions

How do I switch from Xiaohongshu to Douyin without editing code?

Use the command-line interface: python main.py --platform dy. The PlatformEnum in cmd_arg/arg.py (line 40) maps the CLI value to the PLATFORM constant in base_config.py at runtime.

What is the difference between CRAWLER_TYPE options?

The CRAWLER_TYPE setting (line 31) accepts three values: "search" for keyword-based crawling, "detail" for specific post/note URLs, and "creator" for user profile scraping. Each type triggers different logic in the platform-specific crawler implementations.

Why should I use CDP mode instead of standard Playwright?

ENABLE_CDP_MODE (line 59, default True) uses Chrome DevTools Protocol with a real browser instance, significantly reducing detection rates by anti-bot systems. Standard Playwright operates in a more detectable automation environment.

Where do I store sensitive data like cookies and proxy URLs?

Place sensitive values in a .env file at the project root and reference them via os.getenv() in config/base_config.py. Alternatively, pass cookies via the COOKIES constant (line 30) if your deployment environment supports secure configuration injection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →