Where Are the Configuration Files Located in MediaCrawler?

MediaCrawler stores all runtime configuration in the main/config/ directory, with platform-specific modules for each social network and a shared base_config.py for common settings.

MediaCrawler, the open-source social media crawling framework maintained by NanmiCoder, organizes its configuration files within a dedicated Python package under the main/ source tree. Understanding the exact location and structure of these files is essential for customizing API endpoints, authentication credentials, and database connections across supported platforms.

Configuration Directory Structure

All configuration modules reside under the main/config/ path in the repository. This location follows a modular architecture that separates shared defaults from platform-specific parameters, allowing each crawler to operate with isolated settings.

The directory structure follows this pattern:


MediaCrawler/
└── main/
    └── config/
        ├── base_config.py
        ├── zhihu_config.py
        ├── xhs_config.py
        ├── weibo_config.py
        ├── tieba_config.py
        ├── ks_config.py
        ├── dy_config.py
        ├── bilibili_config.py
        └── db_config.py

Core Configuration Files Explained

base_config.py

The main/config/base_config.py module defines shared defaults and utility functions used across all crawling operations. This file contains common HTTP headers, retry logic parameters, and helper methods for loading environment-specific values. When adjusting global behavior like request timeouts or user-agent rotation, modify this central file to affect all platform crawlers.

Platform-Specific Configurations

Each supported social media platform maintains its own dedicated configuration module in the same directory:

  • zhihu_config.py – Stores Zhihu API endpoints, authentication keys, and crawling parameters
  • xhs_config.py – Contains XiaoHongShu configuration including API tokens and pagination controls
  • weibo_config.py – Houses Weibo authentication cookies, API URLs, and rate-limit settings
  • tieba_config.py – Defines Baidu Tieba forum URLs, thread limits, and scraping intervals
  • ks_config.py – Configures Kuaishou video fetch parameters and credential handling
  • dy_config.py – Stores Douyin-specific API keys and endpoint paths
  • bilibili_config.py – Contains Bilibili API token placeholders and request configurations

These modules enable platform-specific customization without modifying the core BaseCrawler logic.

db_config.py

The main/config/db_config.py file centralizes database connection management. It defines connection strings, cache settings, and SQLAlchemy ORM options used throughout the application. When deploying MediaCrawler with persistent storage, this file controls how the system connects to your backend database.

How to Access Configuration Values in Code

Configuration classes are imported directly by crawler implementations at runtime. The system uses simple Python imports to retrieve settings, ensuring type-safe access to constants and methods.

Loading Platform Configuration

To access Zhihu-specific settings in your crawler:

from config.zhihu_config import ZhihuConfig

# Access a specific setting

api_endpoint = ZhihuConfig.API_URL

Using Base Configuration

For common headers shared across all platforms:

from config.base_config import BaseConfig

# Retrieve default headers used for all HTTP requests

default_headers = BaseConfig.DEFAULT_HEADERS

Database Configuration

To initialize database connections using configured values:

from config.db_config import DBConfig

# Create engine using configuration values

engine = DBConfig.create_engine()

Summary

  • MediaCrawler configuration files are located in the main/config/ directory of the repository
  • base_config.py provides shared defaults and utility functions for all crawling operations
  • Each platform (Zhihu, XiaoHongShu, Weibo, etc.) has its own dedicated config module for specific endpoints and credentials
  • db_config.py manages database connection strings and ORM settings
  • Configuration values are accessed via direct Python imports from the config package

Frequently Asked Questions

Where is the main configuration directory in MediaCrawler?

The main configuration directory is located at main/config/ in the repository root. This package contains all platform-specific modules and the shared base configuration. According to the NanmiCoder/MediaCrawler source code, core classes like BaseCrawler import settings from this path to retrieve runtime parameters.

How do I add custom settings to MediaCrawler?

You can add custom settings by extending the appropriate configuration module in main/config/. For platform-specific values, edit the corresponding file (e.g., xhs_config.py for XiaoHongShu). For global settings that apply to all crawlers, define new constants in base_config.py where utility functions for environment-specific loading are also provided.

Which file should I edit for database connection settings?

Database configuration is handled exclusively in main/config/db_config.py. This file contains connection strings, pool settings, and SQLAlchemy ORM configurations. Modify this file when changing your database backend or adjusting connection parameters like timeouts, SSL settings, or cache options.

Can I use environment variables with MediaCrawler configs?

Yes, the configuration system supports environment variables through utility functions defined in main/config/base_config.py. You can leverage these helpers to load sensitive credentials like API keys and database URLs from environment variables, keeping secrets out of version control while maintaining the structured configuration approach.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →