How to Configure MediaCrawler Settings: Complete Guide to Python Config Files and CLI Options
Configure MediaCrawler by editing Python modules in the config/ directory—primarily base_config.py for global settings and platform-specific files like xhs_config.py for URL lists—without modifying any source code.
MediaCrawler is an open-source, multi-platform social media scraper that centralizes all behavior in configuration files located under config/. As implemented in NanmiCoder/MediaCrawler, the framework imports these modules at runtime through config/__init__.py, which re-exports every constant from base_config.py and db_config.py. This architecture lets you customize crawling behavior, authentication, proxies, and data output formats by changing values in a few Python files.
Core Global Configuration (config/base_config.py)
The config/base_config.py file controls platform selection, authentication, browser behavior, proxy settings, and data persistence. All constants are exposed via from .base_config import * in config/__init__.py, making them available project-wide.
Platform and Crawl Target Settings
| Setting | Description | Default Value |
|---|---|---|
PLATFORM |
Target platform: xhs, dy, ks, bili, wb, tieba, zhihu |
"xhs" |
XHS_INTERNATIONAL |
Use overseas Xiaohongshu (rednote.com) |
False |
KEYWORDS |
Comma-separated search terms | "编程副业,编程兼职" |
CRAWLER_TYPE |
Crawl mode: search, detail, or creator |
"search" |
In base_config.py line 21, PLATFORM defaults to "xhs" but can be switched to any supported platform identifier.
Authentication Configuration
MediaCrawler supports three login methods configured via LOGIN_TYPE (line 29):
qrcode– Scan QR code with mobile app (default)phone– SMS verification logincookie– Use existing session viaCOOKIESstring (line 30)
# config/base_config.py
LOGIN_TYPE = "cookie"
COOKIES = "acw_tc=xxx; session_id=yyy; ..." # Paste full cookie string
When SAVE_LOGIN_STATE = True (line 53), the crawler persists cookies between runs to avoid repeated authentication.
Browser and CDP Settings
The framework uses Chrome DevTools Protocol (CDP) for anti-detection. Key settings from lines 59-86:
# config/base_config.py
ENABLE_CDP_MODE = True # Use real browser instead of Playwright
CDP_DEBUG_PORT = 9222 # Chrome remote debugging port
HEADLESS = False # Show browser UI (False) or hide it (True)
CDP_HEADLESS = False # Headless inside CDP (may trigger detection)
CDP_CONNECT_EXISTING = True # Attach to already-running Chrome
CUSTOM_BROWSER_PATH = "" # Path to Chrome/Edge binary (empty = auto-detect)
BROWSER_LAUNCH_TIMEOUT = 60 # Seconds to wait for browser startup
AUTO_CLOSE_BROWSER = True # Close browser when script ends
To connect to an existing Chrome instance, start it with remote debugging and match the port:
google-chrome --remote-debugging-port=9222
Proxy Configuration
Enable IP rotation via ENABLE_IP_PROXY (line 33) with three provider options:
# config/base_config.py
ENABLE_IP_PROXY = True
IP_PROXY_POOL_COUNT = 2
IP_PROXY_PROVIDER_NAME = "kuaidaili" # Alternatives: "wandouhttp", "static"
# For static proxy:
# IP_PROXY_PROVIDER_NAME = "static"
# STATIC_PROXY_URL = "http://myproxy.example.com:3128"
Data Output Format
Control persistence via SAVE_DATA_OPTION (line 90):
# config/base_config.py
SAVE_DATA_OPTION = "jsonl" # Options: csv, db, json, jsonl, sqlite, excel, postgres
SAVE_DATA_PATH = "" # Empty defaults to data/ directory
The store_factory.py module (in store/) instantiates the appropriate storage class based on this constant.
Platform-Specific Configuration
Each supported platform has dedicated config files for URL lists and special parameters. These are imported alongside base_config.py at runtime.
Xiaohongshu (config/xhs_config.py)
Define specific notes or creators to crawl when CRAWLER_TYPE is detail or creator:
# config/xhs_config.py
XHS_SPECIFIED_NOTE_URL_LIST = [
"https://www.xiaohongshu.com/explore/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
]
XHS_CREATOR_ID_LIST = [
"https://www.xiaohongshu.com/user/profile/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
]
URLs must include the xsec_token query parameter for authentication validation (lines 26-37).
Zhihu (config/zhihu_config.py)
Configure user profiles, questions, answers, and videos:
# config/zhihu_config.py
ZHIHU_CREATOR_URL_LIST = [
"https://www.zhihu.com/people/example-user",
]
ZHIHU_SPECIFIED_ID_LIST = [
"https://www.zhihu.com/question/123456789",
"https://www.zhihu.com/zvideo/987654321",
]
Other platforms follow identical patterns: dy_config.py (Douyin), ks_config.py (Kuaishou), bilibili_config.py, tieba_config.py, and weibo_config.py.
Command-Line Overrides
The cmd_arg/arg.py module defines enumerations that map directly to configuration constants, enabling runtime overrides without code changes:
python main.py \
--platform zhihu \
--login-type phone \
--crawler-type creator \
--save-data-option csv \
--headless true
Available enum mappings (lines 40-68):
PlatformEnum→PLATFORMLoginTypeEnum→LOGIN_TYPECrawlerTypeEnum→CRAWLER_TYPESaveDataOptionEnum→SAVE_DATA_OPTION
CLI arguments take precedence over file-based configuration.
Database Configuration (config/db_config.py)
Database connectivity for MySQL, PostgreSQL, SQLite, and MongoDB is declared in db_config.py. This file is imported by database/db_session.py and storage factories, so credential changes propagate across all persistence operations.
Environment Variables for Secrets
Sensitive values can be externalized to a .env file (optional, requires python-dotenv):
# .env in project root
COOKIES=your_session_cookie_string
STATIC_PROXY_URL=http://auth:pass@proxy.example.com:8080
Reference these in base_config.py:
import os
COOKIES = os.getenv("COOKIES", "")
Configuration Execution Flow
When you run python main.py, the following sequence executes (as documented in docs/项目架构文档.md):
- Entry point (
main.py) parses CLI arguments - Config import loads
config/package →base_config.py,db_config.py, platform configs - CrawlerFactory (
media_platform/) instantiates platform-specific crawler usingPLATFORM - Crawler execution uses authentication, proxy, and browser settings from config
- StoreFactory (
store/) writes data according toSAVE_DATA_OPTION
No hard-coded values remain after import—all behavior is externally configurable.
Practical Configuration Examples
Switch to Douyin with Cookie Login
# config/base_config.py
PLATFORM = "dy"
LOGIN_TYPE = "cookie"
COOKIES = "sessionid=xxx; sid_guard=xxx"
CRAWLER_TYPE = "search"
KEYWORDS = "美食,旅游"
Enable SQLite Output with Custom Path
# config/base_config.py
SAVE_DATA_OPTION = "sqlite"
SAVE_DATA_PATH = "/var/media_crawler/db"
Crawl Specific Weibo Posts
# config/weibo_config.py
WEIBO_SEARCH_TYPE = "1" # Comprehensive search
WEIBO_SPECIFIED_ID_LIST = ["4892012345678901", "4892012345678902"]
Connect to Remote Debug Chrome
# Terminal 1: Start Chrome
/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome \
--remote-debugging-port=9222 \
--user-data-dir=/tmp/chrome_debug_profile
# Terminal 2: Run crawler
python main.py --cdp-connect-existing true --headless false
Summary
- All configuration lives in
config/— editbase_config.pyfor global settings and{platform}_config.pyfor URL lists - Runtime overrides via CLI flags defined in
cmd_arg/arg.pytake precedence over file values - Authentication supports QR code, phone, or cookie-based login with optional persistence
- Proxy support includes Kuaidaili, Wandou, or static proxy configurations
- Data output spans 7 formats:
csv,json,jsonl,sqlite,excel,postgres,db - CDP mode uses real Chrome/Edge for anti-detection with configurable headless and connection options
Frequently Asked Questions
Can I configure MediaCrawler without editing Python files?
Yes. Use command-line arguments to override any base_config.py setting. For example: python main.py --platform dy --login-type cookie --save-data-option csv. CLI flags map directly to configuration constants via the enums in cmd_arg/arg.py.
Where do I set cookies for authentication?
Paste the complete cookie string into COOKIES in config/base_config.py (line 30), or use the COOKIES environment variable in a .env file. Then set LOGIN_TYPE = "cookie". The crawler will use this session for authenticated requests.
Why does my Xiaohongshu URL fail with authentication errors?
Xiaohongshu URLs require the xsec_token query parameter for validation. Ensure every URL in XHS_SPECIFIED_NOTE_URL_LIST or XHS_CREATOR_ID_LIST includes this token, which you can extract from your browser's URL bar when viewing the content.
How do I switch from JSONL to database storage?
Change SAVE_DATA_OPTION to your preferred format: sqlite, postgres, db (MySQL), or mongo. Update config/db_config.py with connection credentials if using a server database. The StoreFactory automatically instantiates the correct storage class.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →