How to Customize MediaCrawler's Settings Using the Configuration System
MediaCrawler reads runtime settings from Python module-level variables in the config package, allowing customization through direct file edits, command-line overrides, or programmatic modification before execution.
The NanmiCoder/MediaCrawler repository organizes all tunable parameters in a centralized configuration system located under main/config. This architecture uses plain Python variables rather than external formats like YAML or JSON, making it straightforward to customize MediaCrawler's settings using the configuration system by editing source files or injecting values via the CLI.
Understanding the Configuration Architecture
The configuration system consists of a base configuration file that defines global defaults and platform-specific modules that extend these settings with URL lists and identifiers.
Base Configuration File
The config/base_config.py file serves as the single source of truth for all crawler behavior. It defines module-level variables such as PLATFORM, HEADLESS, and CRAWLER_MAX_NOTES_COUNT that control core functionality. At the bottom of this file, the system imports platform-specific configurations using wildcard imports:
# config/base_config.py (excerpt)
from .bilibili_config import *
from .xhs_config import *
from .dy_config import *
from .ks_config import *
from .weibo_config import *
from .tieba_config import *
from .zhihu_config import *
Platform-Specific Configuration Files
Each supported platform maintains its own configuration file (e.g., config/zhihu_config.py, config/xhs_config.py, config/ks_config.py). These files contain lists such as XHS_SPECIFIED_NOTE_URL_LIST or ZHIHU_SPECIFIED_ID_LIST that specify which content to fetch when operating in detail mode.
Key Configuration Variables
The following variables in config/base_config.py control essential crawler behavior:
PLATFORM– Target platform identifier ("xhs","dy","ks","bili","wb","tieba","zhihu")HEADLESS– Boolean flag to run browsers in headless mode (default:False)ENABLE_IP_PROXY– Boolean to enable HTTP proxy handling for all requests (default:False)STATIC_PROXY_URL– Proxy URL string when using static proxy configurationCRAWLER_MAX_NOTES_COUNT– Integer limiting the maximum number of posts or videos to fetch (default:15)SAVE_DATA_OPTION– Output format specification ("csv","db","json","jsonl","sqlite","excel","postgres")
How the CLI Overrides Configuration Values
The command-line interface in cmd_arg/arg.py implements a two-way binding with the configuration module. When you execute the crawler, the parse_cmd() function (around lines 48-68) performs the following:
- Reads the current value from
configto display as defaults in help text - Writes the parsed flag value directly back to the same module variable
# cmd_arg/arg.py (excerpt)
config.PLATFORM = platform.value
config.LOGIN_TYPE = lt.value
config.CRAWLER_TYPE = crawler_type.value
config.ENABLE_IP_PROXY = enable_ip_proxy_value
Because the CLI writes directly into the config module namespace, any subsequent code imports automatically receive the overridden values without requiring additional parsing steps.
Three Methods to Customize MediaCrawler Settings
Method 1: Edit Configuration Files Directly
Modify the Python variables in config/base_config.py or platform-specific files for persistent changes:
# main/config/base_config.py
HEADLESS = True # Run Chrome/Edge without UI
ENABLE_IP_PROXY = True # Enable proxy handling
STATIC_PROXY_URL = "http://proxy.example.com:3128"
CRAWLER_MAX_NOTES_COUNT = 30 # Fetch up to 30 items per run
SAVE_DATA_OPTION = "csv" # Export results as CSV files
Method 2: Override via Command-Line Arguments
Apply temporary changes without touching source files by passing flags to the main module:
python -m MediaCrawler.main \
--platform ks \
--headless true \
--enable_ip_proxy true \
--static_proxy_url http://proxy.example.com:3128 \
--crawler_max_notes_count 50 \
--save_data_option jsonl
Method 3: Programmatic Configuration in Python
Import the config module from external scripts to modify settings at runtime before invoking the crawler:
import config
from MediaCrawler.main import main
# Change settings at runtime
config.PLATFORM = "zhihu"
config.ZHIHU_SPECIFIED_ID_LIST = [
"https://www.zhihu.com/question/123456789/answer/987654321"
]
config.ENABLE_GET_COMMENTS = False
# Run the crawler with the new configuration
main()
Platform-Specific Settings
For targeted crawling, populate the list variables in platform-specific configuration files:
XHS_SPECIFIED_NOTE_URL_LISTinconfig/xhs_config.py– XiaoHongShu note URLs for detail modeZHIHU_SPECIFIED_ID_LISTinconfig/zhihu_config.py– Zhihu answer or article URLsKS_CREATOR_ID_LISTinconfig/ks_config.py– Kuaishou creator identifiers (full URLs or plain IDs)
Summary
- MediaCrawler stores settings as module-level Python variables in
main/config, withbase_config.pydefining globals and platform-specific files extending them. - The CLI in
cmd_arg/arg.pyreads from and writes directly to theconfigmodule, ensuring command-line flags override file-based defaults. - You can customize settings through three approaches: editing Python source files, supplying CLI arguments, or programmatically modifying the
configobject before execution. - Platform-specific lists (URLs, IDs) reside in dedicated configuration files like
xhs_config.pyandzhihu_config.py.
Frequently Asked Questions
Where are MediaCrawler's configuration files located?
All configuration files reside in the main/config directory. The primary file is base_config.py, which contains global defaults, while platform-specific settings live in files like xhs_config.py, zhihu_config.py, and ks_config.py. These files use standard Python syntax, allowing you to edit variables directly without parsing external formats.
How do I change the target platform in MediaCrawler?
Set the PLATFORM variable in config/base_config.py to your desired platform code ("xhs", "dy", "ks", "bili", "wb", "tieba", or "zhihu"). Alternatively, pass the --platform flag when running from the command line, which updates config.PLATFORM via the CLI parser in cmd_arg/arg.py before the crawler initializes.
Can I use environment variables instead of editing config files?
The standard configuration system does not automatically read environment variables. However, because the settings are plain Python variables, you can modify base_config.py to read os.environ values, or import the config module in a wrapper script that sets variables based on your environment before calling the main crawler function.
How do I configure proxy settings in MediaCrawler?
Set ENABLE_IP_PROXY = True in base_config.py and provide the proxy URL in STATIC_PROXY_URL (e.g., "http://user:pwd@host:port"). For dynamic proxy providers, configure the IP_PROXY_PROVIDER_NAME variable. These settings apply to both Playwright and CDP browser instances, as well as direct HTTP requests made by the crawler.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →