How to Contribute to the MediaCrawler Project: A Developer's Guide
MediaCrawler is a modular, multi-platform web-scraping framework built on Playwright that accepts contributions through GitHub pull requests following a standardized workflow based on the AbstractCrawler base class.
The NanmiCoder/MediaCrawler repository provides a pluggable architecture for scraping Chinese social media platforms like XiaoHongShu, Douyin, and Zhihu. Whether you want to add support for a new platform, fix a data storage bug, or improve the CDP mode documentation, understanding the project structure is essential for a successful contribution.
Understanding MediaCrawler's Architecture
MediaCrawler organizes its codebase into distinct layers that separate concerns between command-line entry, configuration, crawling logic, and data persistence.
CLI and Entry Points
The main entry point resides in main.py, which handles command-line argument parsing and instantiates the correct crawler via the CrawlerFactory mapping (lines 50-66). This factory pattern maps platform identifiers (e.g., "xhs", "zhihu") to concrete crawler classes, ensuring that adding a new platform only requires registration in one location.
Configuration Layer
Centralized settings live in config/base_config.py, defining feature toggles such as ENABLE_CDP_MODE, SAVE_DATA_OPTION, and proxy configurations. When contributing new features that require user configuration, you must add the corresponding flags here to maintain consistency with the existing schema.
Abstract Crawler Pattern
All platform-specific implementations inherit from AbstractCrawler defined in base/base_crawler.py. This base class standardizes the lifecycle methods (start, stop) and optional CDP manager handling. Your contribution must extend this abstraction rather than implementing standalone scraping logic from scratch.
Platform Implementations
Concrete crawlers for each media site reside under media_platform/ (e.g., media_platform/xhs.py, media_platform/zhihu.py). These modules contain the site-specific login flows, API wrappers, and parsing logic that extract data from target platforms.
Data Storage Layer
The framework supports multiple output formats through store/excel_store_base.py and platform-specific implementations like store/zhihu/_store_impl.py or store/bilibili/_store_impl.py. When modifying how data persists, ensure you update both the base logic and any platform-specific storage subclasses.
Utilities and Testing
Helper utilities such as the async file writer (tools/async_file_writer.py) support word-cloud generation and cleanup tasks. The tests/ directory contains unit and integration tests that enforce behavior across platforms—essential for validating that your changes do not break existing functionality.
Setting Up Your Development Environment
MediaCrawler uses uv for dependency management and requires Python 3.8+. You must install Playwright browsers after syncing dependencies.
# Clone your fork
git clone https://github.com/<your-username>/MediaCrawler.git
cd MediaCrawler
# Install uv if you haven't already
curl -LsSf https://astral.sh/uv/install.sh | sh
# Sync the Python environment (includes Playwright)
uv sync
Contribution Workflow
Follow this standardized process to ensure your pull request meets the project's quality standards.
1. Fork and Clone
Create a personal fork on GitHub, then clone it locally using the commands above.
2. Run Baseline Tests
Verify the existing codebase passes before making changes:
uv run pytest
3. Create a Feature Branch
Use descriptive branch names that indicate the contribution type:
git checkout -b feature/add-tiktok-support
4. Implement Your Change
The implementation strategy depends on your contribution type:
- New Platform: Create a subclass in
media_platform/that inherits fromAbstractCrawler, then register it inmain.py(lines 50-66) within theCrawlerFactory.CRAWLERSdictionary. - Bug Fix: Locate the failing test in
tests/(e.g.,test_tieba_extractor.py), modify the implementation in the relevantstore/*_impl.pyfile, and ensure the test passes. - Documentation: Update
README.mdor add guides underdocs/. The CDP mode guide (docs/CDP模式使用指南.md, lines 80-84) already contains a "贡献" section welcoming improvements to CDP documentation.
5. Validate Your Changes
Run the full test suite and linting tools:
uv run pytest
uv run ruff .
6. Submit Your Pull Request
Commit with clear, conventional commit messages and push to your fork:
git add .
git commit -m "feat: add TikTok crawler with QR code login support"
git push origin feature/add-tiktok-support
Open a Pull Request against the main repository. The maintainers may request additional test coverage, linting fixes, or documentation updates before merging.
Code Standards and Best Practices
Adhering to these guidelines ensures your contribution aligns with the existing codebase quality.
Code Style: Follow PEP 8 conventions. The repository uses ruff and black via uv lock; always run uv run ruff . before committing.
Type Hints: Maintain type consistency across all public APIs. The AbstractCrawler class and its implementations use explicit type annotations that you must preserve.
Testing: Add unit tests covering new code paths. Use the tmp_path fixture for filesystem-related tests to ensure isolation.
Documentation: When adding platforms, document required configuration keys in config/<platform>_config.py and update the main README.md with setup instructions.
License Compliance: All contributions fall under the project's NON-COMMERCIAL LEARNING LICENSE 1.1. Ensure you have read the LICENSE file before submitting code.
CI Pipeline: GitHub Actions automatically runs the test suite on each PR. Do not merge if the pipeline fails.
Example: Adding a New Platform
Here is a minimal implementation pattern for adding a crawler called "example":
# File: media_platform/example.py
from base.base_crawler import AbstractCrawler
class ExampleCrawler(AbstractCrawler):
async def start(self) -> None:
print("[Example] Crawling started")
await self.fetch_some_data()
print("[Example] Crawling finished")
async def fetch_some_data(self):
# Implementation details here
pass
Register the new crawler in the factory:
# Edit main.py (lines 50-66)
from media_platform.example import ExampleCrawler
CrawlerFactory.CRAWLERS["example"] = ExampleCrawler
Test your implementation locally:
uv run main.py --platform example --lt qrcode --type search
Summary
- MediaCrawler uses a factory pattern in
main.pyto manage platform-specific crawlers inheriting fromAbstractCrawlerinbase/base_crawler.py. - Development setup requires
uvfor dependency management andpytestfor validation. - New platforms require a class in
media_platform/, registration inCrawlerFactory, and corresponding configuration inconfig/base_config.py. - Code quality depends on
rufflinting, type hints, and comprehensive testing undertests/. - Documentation contributions are welcomed, particularly for CDP mode improvements referenced in
docs/CDP模式使用指南.md.
Frequently Asked Questions
What license covers contributions to MediaCrawler?
All contributions are governed by the NON-COMMERCIAL LEARNING LICENSE 1.1. This license restricts commercial usage while permitting educational and research purposes. Review the LICENSE file in the repository root before submitting your pull request.
How do I add support for a new social media platform?
Create a Python module in media_platform/ containing a class that inherits from AbstractCrawler. Implement the required start() and stop() methods, then register your class in main.py within the CrawlerFactory.CRAWLERS dictionary (lines 50-66). Add corresponding configuration defaults in config/base_config.py and create unit tests under tests/.
What testing framework does MediaCrawler use?
The project uses pytest for unit and integration testing. Run uv run pytest to execute the full suite. When contributing bug fixes or new features, add tests that cover your specific code paths, using fixtures like tmp_path for filesystem operations to ensure test isolation.
Where should I document my changes?
Update README.md for user-facing features or platform additions. For technical details regarding Chrome DevTools Protocol (CDP) usage, contribute to docs/CDP模式使用指南.md, which already contains a dedicated "贡献" section (lines 80-84) welcoming improvements to the CDP implementation guide.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →