How MediaCrawler Manages Configuration for Different Platform Scrapers
MediaCrawler uses a centralized config package where base_config.py defines global defaults and imports platform-specific modules (like zhihu_config.py or weibo_config.py) into a unified namespace, with the active platform selected via the PLATFORM constant that can be overridden by command-line arguments.
MediaCrawler implements a modular configuration system that separates global scraping settings from platform-specific parameters. The repository organizes all runtime settings within a central config package, allowing developers to switch between scrapers for Zhihu, Weibo, Douyin, and other platforms without modifying core code. This architecture ensures that shared settings like proxy handling and headless mode remain consistent while platform-specific variables remain isolated in dedicated modules.
Centralized Configuration Architecture
The configuration system relies on a two-layer hierarchy that merges global and platform-specific settings into a single namespace.
Global defaults reside in config/base_config.py. This file contains universal parameters such as request limits, proxy handling, CDP (Chrome DevTools Protocol) settings, and the crucial PLATFORM selector that determines which scraper to activate.
Platform-specific modules extend these defaults. Each supported platform—Zhihu, Weibo, Tieba, Douyin, Kuaishou, Bilibili, and Xiaohongshu—maintains its own configuration file (e.g., config/zhihu_config.py, config/weibo_config.py). These files define unique variables such as API endpoints, creator URL lists, and platform-specific flags.
At the bottom of base_config.py, the system executes wildcard imports to merge all platform configurations:
# config/base_config.py (excerpt)
from .zhihu_config import *
from .weibo_config import *
from .tieba_config import *
# ... additional platform imports
This pattern places all configuration variables into a unified namespace, allowing any part of the codebase to access config.SOME_SETTING regardless of which platform is currently active.
Key Configuration Files and Their Roles
Understanding the file structure is essential for navigating the configuration system:
| File | Purpose |
|---|---|
config/base_config.py |
Defines global defaults, imports all platform modules, and declares the PLATFORM selector constant |
config/zhihu_config.py |
Platform-specific URLs, creator ID lists, and Zhihu-specific parameters |
config/weibo_config.py |
Weibo-specific settings including API endpoints and search configurations |
cmd_arg/arg.py |
Parses --platform CLI arguments and updates config.PLATFORM at runtime |
main/main.py |
Application entry point that builds crawlers based on config.PLATFORM |
base/base_crawler.py |
Abstract base class that all platform crawlers inherit from |
media_platform/<platform>/core.py |
Concrete implementations that reference the unified config namespace |
Runtime Configuration Flow
The configuration system operates through a specific initialization sequence when the scraper starts:
-
Import Phase: When
main.pyimportsconfig.base_config, Python executes the module top-to-bottom. -
Module Aggregation:
base_config.pyfirst sets generic defaults, then executesfrom .zhihu_config import *(and similar imports for other platforms), merging all symbols into theconfignamespace. -
Platform Selection: The
PLATFORMconstant (defaulting to"xhs"for Xiaohongshu) determines which crawler implementation to instantiate. -
Factory Instantiation: The CrawlerFactory creates the appropriate crawler via
CrawlerFactory.create_crawler(platform=config.PLATFORM). -
Configuration Consumption: Inside each platform's
core.pyfile (e.g.,media_platform/zhihu/core.py), the scraper reads values such asconfig.KEYWORDS,config.HEADLESS, or platform-specific lists likeconfig.ZHIHU_CREATOR_URL_LIST. -
CLI Override: If the user supplies
--platform weiboon the command line,cmd_arg/arg.pyupdatesconfig.PLATFORMbefore the factory is called, redirecting execution to the Weibo implementation while preserving shared global settings.
Switching Platforms via Command Line
The configuration system supports dynamic platform selection without code changes. The cmd_arg/arg.py module parses command-line arguments and writes directly back to the configuration namespace:
# Simplified runtime illustration
from config import base_config as cfg
from cmd_arg.arg import parse_args
from main import CrawlerFactory
if __name__ == "__main__":
# Parse CLI arguments; this may change cfg.PLATFORM
args = parse_args()
# Build the appropriate crawler based on active configuration
crawler = CrawlerFactory.create_crawler(platform=cfg.PLATFORM)
crawler.run()
Command-line usage examples:
# Use the default platform (Xiaohongshu)
python run.py
# Switch to Zhihu scraper
python run.py --platform zhihu
# Switch to Weibo scraper with headless mode enabled
python run.py --platform weibo --headless
All commands rely on the same configuration objects; only the PLATFORM variable changes which platform module's settings are consulted at runtime.
Extending the System for New Platforms
Adding a new scraper (e.g., for Twitter) requires minimal changes to the existing configuration infrastructure:
-
Create
config/twitter_config.pycontaining platform-specific constants likeTWITTER_API_KEYandTWITTER_CREATOR_URL_LIST. -
Add
from .twitter_config import *at the end ofconfig/base_config.pyto merge the new symbols into the global namespace. -
Implement
media_platform/twitter/core.pythat reads from the unifiedconfigobject (e.g.,config.TWITTER_API_KEY). -
Register the new platform in
cmd_arg/arg.pyand extendCrawlerFactoryto return the new crawler class.
Because all settings share a single namespace through the wildcard import pattern, the rest of the codebase remains untouched when adding new platforms.
Summary
- MediaCrawler uses a unified configuration namespace created by
base_config.pyimporting platform-specific modules via wildcard imports. - The
PLATFORMconstant inbase_config.pydetermines which scraper implementation runs, defaulting to"xhs"(Xiaohongshu) but overridable via CLI. - Platform-specific settings live in dedicated files like
zhihu_config.pyandweibo_config.py, while global settings (proxies, CDP, limits) remain inbase_config.py. - Command-line arguments parsed in
cmd_arg/arg.pycan modifyconfig.PLATFORMat runtime, enabling the same binary to run any supported scraper. - CrawlerFactory instantiates the correct implementation based on the active
PLATFORMvalue, with all crawlers reading from the sameconfignamespace.
Frequently Asked Questions
How does MediaCrawler determine which platform to scrape?
MediaCrawler checks the PLATFORM constant defined in config/base_config.py. This value defaults to "xhs" for Xiaohongshu, but the command-line parser in cmd_arg/arg.py can override it when users pass the --platform argument. The CrawlerFactory then instantiates the appropriate crawler class based on this configuration value.
Where are platform-specific URLs and creator lists defined?
Platform-specific variables reside in dedicated configuration files within the config/ directory. For example, Zhihu-specific settings live in config/zhihu_config.py, while Weibo settings are in config/weibo_config.py. These files contain unique constants like ZHIHU_CREATOR_URL_LIST or WEIBO_SEARCH_KEYWORDS, which are merged into the global config namespace via wildcard imports in base_config.py.
Can I override configuration settings via command line?
Yes. The cmd_arg/arg.py module parses CLI arguments and updates the configuration namespace before the crawler initializes. Users can specify --platform to switch scrapers, --headless to toggle browser visibility, and other flags that modify the runtime configuration without editing Python files.
How do I add a new platform scraper to the configuration system?
Create a new configuration file in config/ (e.g., twitter_config.py) with your platform-specific constants, then add from .twitter_config import * to the bottom of config/base_config.py. Implement the crawler logic in media_platform/twitter/core.py, register the platform in cmd_arg/arg.py, and extend CrawlerFactory to handle the new platform type. The existing wildcard import system automatically incorporates your new settings into the global namespace.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →