How to Configure MediaCrawler Settings: Complete Guide to Python Config Files and CLI Options

Configure MediaCrawler by editing Python modules in the config/ directory—primarily base_config.py for global settings and platform-specific files like xhs_config.py for URL lists—without modifying any source code.

MediaCrawler is an open-source, multi-platform social media scraper that centralizes all behavior in configuration files located under config/. As implemented in NanmiCoder/MediaCrawler, the framework imports these modules at runtime through config/__init__.py, which re-exports every constant from base_config.py and db_config.py. This architecture lets you customize crawling behavior, authentication, proxies, and data output formats by changing values in a few Python files.

Core Global Configuration (config/base_config.py)

The config/base_config.py file controls platform selection, authentication, browser behavior, proxy settings, and data persistence. All constants are exposed via from .base_config import * in config/__init__.py, making them available project-wide.

Platform and Crawl Target Settings

Setting Description Default Value
PLATFORM Target platform: xhs, dy, ks, bili, wb, tieba, zhihu "xhs"
XHS_INTERNATIONAL Use overseas Xiaohongshu (rednote.com) False
KEYWORDS Comma-separated search terms "编程副业,编程兼职"
CRAWLER_TYPE Crawl mode: search, detail, or creator "search"

In base_config.py line 21, PLATFORM defaults to "xhs" but can be switched to any supported platform identifier.

Authentication Configuration

MediaCrawler supports three login methods configured via LOGIN_TYPE (line 29):

  • qrcode – Scan QR code with mobile app (default)
  • phone – SMS verification login
  • cookie – Use existing session via COOKIES string (line 30)

# config/base_config.py

LOGIN_TYPE = "cookie"
COOKIES = "acw_tc=xxx; session_id=yyy; ..."  # Paste full cookie string

When SAVE_LOGIN_STATE = True (line 53), the crawler persists cookies between runs to avoid repeated authentication.

Browser and CDP Settings

The framework uses Chrome DevTools Protocol (CDP) for anti-detection. Key settings from lines 59-86:


# config/base_config.py

ENABLE_CDP_MODE = True          # Use real browser instead of Playwright

CDP_DEBUG_PORT = 9222           # Chrome remote debugging port

HEADLESS = False                # Show browser UI (False) or hide it (True)

CDP_HEADLESS = False            # Headless inside CDP (may trigger detection)

CDP_CONNECT_EXISTING = True     # Attach to already-running Chrome

CUSTOM_BROWSER_PATH = ""        # Path to Chrome/Edge binary (empty = auto-detect)

BROWSER_LAUNCH_TIMEOUT = 60     # Seconds to wait for browser startup

AUTO_CLOSE_BROWSER = True       # Close browser when script ends

To connect to an existing Chrome instance, start it with remote debugging and match the port:

google-chrome --remote-debugging-port=9222

Proxy Configuration

Enable IP rotation via ENABLE_IP_PROXY (line 33) with three provider options:


# config/base_config.py

ENABLE_IP_PROXY = True
IP_PROXY_POOL_COUNT = 2
IP_PROXY_PROVIDER_NAME = "kuaidaili"  # Alternatives: "wandouhttp", "static"

# For static proxy:

# IP_PROXY_PROVIDER_NAME = "static"

# STATIC_PROXY_URL = "http://myproxy.example.com:3128"

Data Output Format

Control persistence via SAVE_DATA_OPTION (line 90):


# config/base_config.py

SAVE_DATA_OPTION = "jsonl"      # Options: csv, db, json, jsonl, sqlite, excel, postgres

SAVE_DATA_PATH = ""             # Empty defaults to data/ directory

The store_factory.py module (in store/) instantiates the appropriate storage class based on this constant.

Platform-Specific Configuration

Each supported platform has dedicated config files for URL lists and special parameters. These are imported alongside base_config.py at runtime.

Xiaohongshu (config/xhs_config.py)

Define specific notes or creators to crawl when CRAWLER_TYPE is detail or creator:


# config/xhs_config.py

XHS_SPECIFIED_NOTE_URL_LIST = [
    "https://www.xiaohongshu.com/explore/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
]

XHS_CREATOR_ID_LIST = [
    "https://www.xiaohongshu.com/user/profile/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
]

URLs must include the xsec_token query parameter for authentication validation (lines 26-37).

Zhihu (config/zhihu_config.py)

Configure user profiles, questions, answers, and videos:


# config/zhihu_config.py

ZHIHU_CREATOR_URL_LIST = [
    "https://www.zhihu.com/people/example-user",
]

ZHIHU_SPECIFIED_ID_LIST = [
    "https://www.zhihu.com/question/123456789",
    "https://www.zhihu.com/zvideo/987654321",
]

Other platforms follow identical patterns: dy_config.py (Douyin), ks_config.py (Kuaishou), bilibili_config.py, tieba_config.py, and weibo_config.py.

Command-Line Overrides

The cmd_arg/arg.py module defines enumerations that map directly to configuration constants, enabling runtime overrides without code changes:

python main.py \
    --platform zhihu \
    --login-type phone \
    --crawler-type creator \
    --save-data-option csv \
    --headless true

Available enum mappings (lines 40-68):

  • PlatformEnum → PLATFORM
  • LoginTypeEnum → LOGIN_TYPE
  • CrawlerTypeEnum → CRAWLER_TYPE
  • SaveDataOptionEnum → SAVE_DATA_OPTION

CLI arguments take precedence over file-based configuration.

Database Configuration (config/db_config.py)

Database connectivity for MySQL, PostgreSQL, SQLite, and MongoDB is declared in db_config.py. This file is imported by database/db_session.py and storage factories, so credential changes propagate across all persistence operations.

Environment Variables for Secrets

Sensitive values can be externalized to a .env file (optional, requires python-dotenv):


# .env in project root

COOKIES=your_session_cookie_string
STATIC_PROXY_URL=http://auth:pass@proxy.example.com:8080

Reference these in base_config.py:

import os
COOKIES = os.getenv("COOKIES", "")

Configuration Execution Flow

When you run python main.py, the following sequence executes (as documented in docs/项目架构文档.md):

  1. Entry point (main.py) parses CLI arguments
  2. Config import loads config/ package → base_config.py, db_config.py, platform configs
  3. CrawlerFactory (media_platform/) instantiates platform-specific crawler using PLATFORM
  4. Crawler execution uses authentication, proxy, and browser settings from config
  5. StoreFactory (store/) writes data according to SAVE_DATA_OPTION

No hard-coded values remain after import—all behavior is externally configurable.

Practical Configuration Examples


# config/base_config.py

PLATFORM = "dy"
LOGIN_TYPE = "cookie"
COOKIES = "sessionid=xxx; sid_guard=xxx"
CRAWLER_TYPE = "search"
KEYWORDS = "美食,旅游"

Enable SQLite Output with Custom Path


# config/base_config.py

SAVE_DATA_OPTION = "sqlite"
SAVE_DATA_PATH = "/var/media_crawler/db"

Crawl Specific Weibo Posts


# config/weibo_config.py

WEIBO_SEARCH_TYPE = "1"  # Comprehensive search

WEIBO_SPECIFIED_ID_LIST = ["4892012345678901", "4892012345678902"]

Connect to Remote Debug Chrome


# Terminal 1: Start Chrome

/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome \
    --remote-debugging-port=9222 \
    --user-data-dir=/tmp/chrome_debug_profile

# Terminal 2: Run crawler

python main.py --cdp-connect-existing true --headless false

Summary

  • All configuration lives in config/ — edit base_config.py for global settings and {platform}_config.py for URL lists
  • Runtime overrides via CLI flags defined in cmd_arg/arg.py take precedence over file values
  • Authentication supports QR code, phone, or cookie-based login with optional persistence
  • Proxy support includes Kuaidaili, Wandou, or static proxy configurations
  • Data output spans 7 formats: csv, json, jsonl, sqlite, excel, postgres, db
  • CDP mode uses real Chrome/Edge for anti-detection with configurable headless and connection options

Frequently Asked Questions

Can I configure MediaCrawler without editing Python files?

Yes. Use command-line arguments to override any base_config.py setting. For example: python main.py --platform dy --login-type cookie --save-data-option csv. CLI flags map directly to configuration constants via the enums in cmd_arg/arg.py.

Where do I set cookies for authentication?

Paste the complete cookie string into COOKIES in config/base_config.py (line 30), or use the COOKIES environment variable in a .env file. Then set LOGIN_TYPE = "cookie". The crawler will use this session for authenticated requests.

Why does my Xiaohongshu URL fail with authentication errors?

Xiaohongshu URLs require the xsec_token query parameter for validation. Ensure every URL in XHS_SPECIFIED_NOTE_URL_LIST or XHS_CREATOR_ID_LIST includes this token, which you can extract from your browser's URL bar when viewing the content.

How do I switch from JSONL to database storage?

Change SAVE_DATA_OPTION to your preferred format: sqlite, postgres, db (MySQL), or mongo. Update config/db_config.py with connection credentials if using a server database. The StoreFactory automatically instantiates the correct storage class.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →