How to Configure MediaCrawler: A Complete Guide to Base, Platform-Specific, and Database Settings
MediaCrawler behavior is controlled entirely through Python configuration modules in the config/ package, allowing you to customize platforms, login methods, proxies, and storage formats without modifying source code.
The NanmiCoder/MediaCrawler repository organizes all tunable parameters into modular config files. At runtime, config/__init__.py re-exports constants from base_config.py and db_config.py, while platform-specific modules like xhs_config.py handle unique URL lists and switches. This architecture decouples crawler logic from deployment settings.
Core Configuration (config/base_config.py)
The config/base_config.py file contains global constants that affect every platform. These settings dictate target platforms, authentication strategies, network proxies, browser behavior, and output formats.
Platform Selection and Login Methods
The PLATFORM constant (line 21) selects your target: "xhs" (Xiaohongshu/Rednote), "dy" (Douyin), "ks" (Kuaishou), "bili" (Bilibili), "wb" (Weibo), "tieba", or "zhihu".
Authentication is controlled by LOGIN_TYPE (line 29), supporting "qrcode", "phone", or "cookie". When using cookie-based login, populate the COOKIES string (line 30). To persist sessions between runs, enable SAVE_LOGIN_STATE (line 53, default True).
For Xiaohongshu specifically, set XHS_INTERNATIONAL (line 25) to True to crawl the overseas version (rednote.com) instead of the domestic site.
Define search queries in KEYWORDS (line 27) as a comma-separated string:
# config/base_config.py
PLATFORM = "xhs"
KEYWORDS = "编程副业,编程兼职"
LOGIN_TYPE = "qrcode"
Proxy and Network Settings
Enable proxy pools by setting ENABLE_IP_PROXY (line 33) to True. Configure the provider via IP_PROXY_PROVIDER_NAME (line 40), choosing from "kuaidaili", "wandouhttp", or "static".
For static proxies, specify the URL in STATIC_PROXY_URL (line 44). Parallel proxy pools are controlled by IP_PROXY_POOL_COUNT (line 37, default 2).
# config/base_config.py
ENABLE_IP_PROXY = True
IP_PROXY_PROVIDER_NAME = "static"
STATIC_PROXY_URL = "http://myproxy.example.com:3128"
Browser and CDP Configuration
MediaCrawler uses Chrome DevTools Protocol (CDP) for anti-detection. Key settings include:
ENABLE_CDP_MODE(line 59): Set toTrueto use real browser automation instead of Playwright (defaultTrue).CDP_DEBUG_PORT(line 63): CDP communication port (default9222).CDP_CONNECT_EXISTING(line 81): Attach to an already-running Chrome instance (defaultTrue).CUSTOM_BROWSER_PATH(line 69): Absolute path to Chrome/Edge binary; leave empty for auto-detection.HEADLESS(line 50): Run browser without UI (defaultFalse).BROWSER_LAUNCH_TIMEOUT(line 76): Seconds to wait for browser startup (default60).
To connect to an existing Chrome instance:
# Start Chrome with remote debugging
google-chrome --remote-debugging-port=9222
# Then run the crawler
python main.py --headless false --cdp-connect-existing true
Data Persistence Options
The SAVE_DATA_OPTION constant (line 90) determines output format: "csv", "db", "json", "jsonl", "sqlite", "excel", or "postgres". The export directory is set via SAVE_DATA_PATH (line 93); if empty, files write to data/.
# config/base_config.py
SAVE_DATA_OPTION = "jsonl"
SAVE_DATA_PATH = "/var/media_crawler/output"
Platform-Specific Configuration
Each platform has dedicated config files under config/ that define URL lists and specialized toggles.
Xiaohongshu (XHS) Settings
In config/xhs_config.py (lines 26-37), define specific content to crawl:
XHS_SPECIFIED_NOTE_URL_LIST: List of note URLs requiring thexsec_tokenquery parameter.XHS_CREATOR_ID_LIST: Creator profile URLs for user-centric crawling.
# config/xhs_config.py
XHS_SPECIFIED_NOTE_URL_LIST = [
"https://www.xiaohongshu.com/explore/64b95d01000000000c034587?xsec_token=YOUR_TOKEN",
"https://www.xiaohongshu.com/explore/64b962a2000000000c0345a9?xsec_token=YOUR_TOKEN",
]
Zhihu, Weibo, and Other Platforms
Platform modules follow identical patterns:
- Zhihu (
config/zhihu_config.py, lines 24-34): UseZHIHU_CREATOR_URL_LISTandZHIHU_SPECIFIED_ID_LISTfor user pages and Q&A content. - Weibo (
config/weibo_config.py, lines 23-28): SetWEIBO_SEARCH_TYPEfor strategy selection andWEIBO_SPECIFIED_ID_LISTfor concrete post IDs. - Others (
dy_config.py,ks_config.py,bilibili_config.py,tieba_config.py): Each exports analogous URL lists and platform-specific constants.
Database Configuration (config/db_config.py)
Database connectivity for MySQL, PostgreSQL, MongoDB, and SQLite is declared in config/db_config.py. This file is imported by database/db_session.py and storage factories, ensuring credential changes propagate across the entire persistence layer. Update connection strings and pool settings here to switch between local SQLite and production PostgreSQL instances.
Command-Line Overrides
The cmd_arg/arg.py module defines enumerations that map directly to configuration constants, allowing runtime overrides without file edits:
PlatformEnum(line 40): OverridesPLATFORM.LoginTypeEnum(line 52): OverridesLOGIN_TYPE.CrawlerTypeEnum(line 60): OverridesCRAWLER_TYPE("search","detail", or"creator").SaveDataOptionEnum(line 68): OverridesSAVE_DATA_OPTION.
Example CLI invocation:
python main.py \
--platform zhihu \
--login-type phone \
--crawler-type creator \
--save-data-option csv \
--headless true
Practical Configuration Examples
Switching to SQLite Output
# config/base_config.py
SAVE_DATA_OPTION = "sqlite"
SAVE_DATA_PATH = "/var/media_crawler/sqlite"
Enabling Douyin with Phone Login
# config/base_config.py
PLATFORM = "dy"
LOGIN_TYPE = "phone"
CRAWLER_TYPE = "search"
Using Environment Variables for Secrets
Create a .env file in the project root (requires python-dotenv):
COOKIES=your_cookie_string_here
STATIC_PROXY_URL=http://myproxy.example.com:8080
Reference these in base_config.py via os.getenv to keep credentials out of version control.
Summary
- Global settings reside in
config/base_config.pyand control platforms, login methods, proxies, CDP behavior, and storage formats. - Platform-specific modules (
xhs_config.py,zhihu_config.py, etc.) manage URL lists and unique parameters. - Database credentials are centralized in
config/db_config.pyand consumed bydatabase/db_session.py. - CLI arguments defined in
cmd_arg/arg.pyallow temporary overrides without code changes. - All constants are re-exported through
config/__init__.py, ensuring the runtime remains decoupled from configuration files.
Frequently Asked Questions
How do I switch from Xiaohongshu to Douyin without editing code?
Use the command-line interface: python main.py --platform dy. The PlatformEnum in cmd_arg/arg.py (line 40) maps the CLI value to the PLATFORM constant in base_config.py at runtime.
What is the difference between CRAWLER_TYPE options?
The CRAWLER_TYPE setting (line 31) accepts three values: "search" for keyword-based crawling, "detail" for specific post/note URLs, and "creator" for user profile scraping. Each type triggers different logic in the platform-specific crawler implementations.
Why should I use CDP mode instead of standard Playwright?
ENABLE_CDP_MODE (line 59, default True) uses Chrome DevTools Protocol with a real browser instance, significantly reducing detection rates by anti-bot systems. Standard Playwright operates in a more detectable automation environment.
Where do I store sensitive data like cookies and proxy URLs?
Place sensitive values in a .env file at the project root and reference them via os.getenv() in config/base_config.py. Alternatively, pass cookies via the COOKIES constant (line 30) if your deployment environment supports secure configuration injection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →