How MediaCrawler Manages Configuration for Different Platform Scrapers

MediaCrawler uses a centralized config package where base_config.py defines global defaults and imports platform-specific modules (like zhihu_config.py or weibo_config.py) into a unified namespace, with the active platform selected via the PLATFORM constant that can be overridden by command-line arguments.

MediaCrawler implements a modular configuration system that separates global scraping settings from platform-specific parameters. The repository organizes all runtime settings within a central config package, allowing developers to switch between scrapers for Zhihu, Weibo, Douyin, and other platforms without modifying core code. This architecture ensures that shared settings like proxy handling and headless mode remain consistent while platform-specific variables remain isolated in dedicated modules.

Centralized Configuration Architecture

The configuration system relies on a two-layer hierarchy that merges global and platform-specific settings into a single namespace.

Global defaults reside in config/base_config.py. This file contains universal parameters such as request limits, proxy handling, CDP (Chrome DevTools Protocol) settings, and the crucial PLATFORM selector that determines which scraper to activate.

Platform-specific modules extend these defaults. Each supported platform—Zhihu, Weibo, Tieba, Douyin, Kuaishou, Bilibili, and Xiaohongshu—maintains its own configuration file (e.g., config/zhihu_config.py, config/weibo_config.py). These files define unique variables such as API endpoints, creator URL lists, and platform-specific flags.

At the bottom of base_config.py, the system executes wildcard imports to merge all platform configurations:


# config/base_config.py (excerpt)

from .zhihu_config import *
from .weibo_config import *
from .tieba_config import *

# ... additional platform imports

This pattern places all configuration variables into a unified namespace, allowing any part of the codebase to access config.SOME_SETTING regardless of which platform is currently active.

Key Configuration Files and Their Roles

Understanding the file structure is essential for navigating the configuration system:

File Purpose
config/base_config.py Defines global defaults, imports all platform modules, and declares the PLATFORM selector constant
config/zhihu_config.py Platform-specific URLs, creator ID lists, and Zhihu-specific parameters
config/weibo_config.py Weibo-specific settings including API endpoints and search configurations
cmd_arg/arg.py Parses --platform CLI arguments and updates config.PLATFORM at runtime
main/main.py Application entry point that builds crawlers based on config.PLATFORM
base/base_crawler.py Abstract base class that all platform crawlers inherit from
media_platform/<platform>/core.py Concrete implementations that reference the unified config namespace

Runtime Configuration Flow

The configuration system operates through a specific initialization sequence when the scraper starts:

  1. Import Phase: When main.py imports config.base_config, Python executes the module top-to-bottom.

  2. Module Aggregation: base_config.py first sets generic defaults, then executes from .zhihu_config import * (and similar imports for other platforms), merging all symbols into the config namespace.

  3. Platform Selection: The PLATFORM constant (defaulting to "xhs" for Xiaohongshu) determines which crawler implementation to instantiate.

  4. Factory Instantiation: The CrawlerFactory creates the appropriate crawler via CrawlerFactory.create_crawler(platform=config.PLATFORM).

  5. Configuration Consumption: Inside each platform's core.py file (e.g., media_platform/zhihu/core.py), the scraper reads values such as config.KEYWORDS, config.HEADLESS, or platform-specific lists like config.ZHIHU_CREATOR_URL_LIST.

  6. CLI Override: If the user supplies --platform weibo on the command line, cmd_arg/arg.py updates config.PLATFORM before the factory is called, redirecting execution to the Weibo implementation while preserving shared global settings.

Switching Platforms via Command Line

The configuration system supports dynamic platform selection without code changes. The cmd_arg/arg.py module parses command-line arguments and writes directly back to the configuration namespace:


# Simplified runtime illustration

from config import base_config as cfg
from cmd_arg.arg import parse_args
from main import CrawlerFactory

if __name__ == "__main__":
    # Parse CLI arguments; this may change cfg.PLATFORM

    args = parse_args()
    
    # Build the appropriate crawler based on active configuration

    crawler = CrawlerFactory.create_crawler(platform=cfg.PLATFORM)
    crawler.run()

Command-line usage examples:


# Use the default platform (Xiaohongshu)

python run.py

# Switch to Zhihu scraper

python run.py --platform zhihu

# Switch to Weibo scraper with headless mode enabled

python run.py --platform weibo --headless

All commands rely on the same configuration objects; only the PLATFORM variable changes which platform module's settings are consulted at runtime.

Extending the System for New Platforms

Adding a new scraper (e.g., for Twitter) requires minimal changes to the existing configuration infrastructure:

  1. Create config/twitter_config.py containing platform-specific constants like TWITTER_API_KEY and TWITTER_CREATOR_URL_LIST.

  2. Add from .twitter_config import * at the end of config/base_config.py to merge the new symbols into the global namespace.

  3. Implement media_platform/twitter/core.py that reads from the unified config object (e.g., config.TWITTER_API_KEY).

  4. Register the new platform in cmd_arg/arg.py and extend CrawlerFactory to return the new crawler class.

Because all settings share a single namespace through the wildcard import pattern, the rest of the codebase remains untouched when adding new platforms.

Summary

  • MediaCrawler uses a unified configuration namespace created by base_config.py importing platform-specific modules via wildcard imports.
  • The PLATFORM constant in base_config.py determines which scraper implementation runs, defaulting to "xhs" (Xiaohongshu) but overridable via CLI.
  • Platform-specific settings live in dedicated files like zhihu_config.py and weibo_config.py, while global settings (proxies, CDP, limits) remain in base_config.py.
  • Command-line arguments parsed in cmd_arg/arg.py can modify config.PLATFORM at runtime, enabling the same binary to run any supported scraper.
  • CrawlerFactory instantiates the correct implementation based on the active PLATFORM value, with all crawlers reading from the same config namespace.

Frequently Asked Questions

How does MediaCrawler determine which platform to scrape?

MediaCrawler checks the PLATFORM constant defined in config/base_config.py. This value defaults to "xhs" for Xiaohongshu, but the command-line parser in cmd_arg/arg.py can override it when users pass the --platform argument. The CrawlerFactory then instantiates the appropriate crawler class based on this configuration value.

Where are platform-specific URLs and creator lists defined?

Platform-specific variables reside in dedicated configuration files within the config/ directory. For example, Zhihu-specific settings live in config/zhihu_config.py, while Weibo settings are in config/weibo_config.py. These files contain unique constants like ZHIHU_CREATOR_URL_LIST or WEIBO_SEARCH_KEYWORDS, which are merged into the global config namespace via wildcard imports in base_config.py.

Can I override configuration settings via command line?

Yes. The cmd_arg/arg.py module parses CLI arguments and updates the configuration namespace before the crawler initializes. Users can specify --platform to switch scrapers, --headless to toggle browser visibility, and other flags that modify the runtime configuration without editing Python files.

How do I add a new platform scraper to the configuration system?

Create a new configuration file in config/ (e.g., twitter_config.py) with your platform-specific constants, then add from .twitter_config import * to the bottom of config/base_config.py. Implement the crawler logic in media_platform/twitter/core.py, register the platform in cmd_arg/arg.py, and extend CrawlerFactory to handle the new platform type. The existing wildcard import system automatically incorporates your new settings into the global namespace.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →