# How to Contribute to the MediaCrawler Project: A Developer's Guide

> Learn how to contribute to the MediaCrawler project a modular web-scraping framework built on Playwright. Follow our GitHub pull request guide to get started.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**MediaCrawler is a modular, multi-platform web-scraping framework built on Playwright that accepts contributions through GitHub pull requests following a standardized workflow based on the `AbstractCrawler` base class.**

The [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) repository provides a pluggable architecture for scraping Chinese social media platforms like XiaoHongShu, Douyin, and Zhihu. Whether you want to add support for a new platform, fix a data storage bug, or improve the CDP mode documentation, understanding the project structure is essential for a successful contribution.

## Understanding MediaCrawler's Architecture

MediaCrawler organizes its codebase into distinct layers that separate concerns between command-line entry, configuration, crawling logic, and data persistence.

### CLI and Entry Points

The main entry point resides in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), which handles command-line argument parsing and instantiates the correct crawler via the **`CrawlerFactory`** mapping (lines 50-66). This factory pattern maps platform identifiers (e.g., `"xhs"`, `"zhihu"`) to concrete crawler classes, ensuring that adding a new platform only requires registration in one location.

### Configuration Layer

Centralized settings live in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), defining feature toggles such as `ENABLE_CDP_MODE`, `SAVE_DATA_OPTION`, and proxy configurations. When contributing new features that require user configuration, you must add the corresponding flags here to maintain consistency with the existing schema.

### Abstract Crawler Pattern

All platform-specific implementations inherit from **`AbstractCrawler`** defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This base class standardizes the lifecycle methods (`start`, `stop`) and optional CDP manager handling. Your contribution must extend this abstraction rather than implementing standalone scraping logic from scratch.

### Platform Implementations

Concrete crawlers for each media site reside under `media_platform/` (e.g., [`media_platform/xhs.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs.py), [`media_platform/zhihu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu.py)). These modules contain the site-specific login flows, API wrappers, and parsing logic that extract data from target platforms.

### Data Storage Layer

The framework supports multiple output formats through [`store/excel_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/excel_store_base.py) and platform-specific implementations like [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) or [`store/bilibili/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/bilibili/_store_impl.py). When modifying how data persists, ensure you update both the base logic and any platform-specific storage subclasses.

### Utilities and Testing

Helper utilities such as the async file writer ([`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py)) support word-cloud generation and cleanup tasks. The `tests/` directory contains unit and integration tests that enforce behavior across platforms—essential for validating that your changes do not break existing functionality.

## Setting Up Your Development Environment

MediaCrawler uses **`uv`** for dependency management and requires Python 3.8+. You must install Playwright browsers after syncing dependencies.

```bash

# Clone your fork

git clone https://github.com/<your-username>/MediaCrawler.git
cd MediaCrawler

# Install uv if you haven't already

curl -LsSf https://astral.sh/uv/install.sh | sh

# Sync the Python environment (includes Playwright)

uv sync

```

## Contribution Workflow

Follow this standardized process to ensure your pull request meets the project's quality standards.

### 1. Fork and Clone

Create a personal fork on GitHub, then clone it locally using the commands above.

### 2. Run Baseline Tests

Verify the existing codebase passes before making changes:

```bash
uv run pytest

```

### 3. Create a Feature Branch

Use descriptive branch names that indicate the contribution type:

```bash
git checkout -b feature/add-tiktok-support

```

### 4. Implement Your Change

The implementation strategy depends on your contribution type:

- **New Platform**: Create a subclass in `media_platform/` that inherits from `AbstractCrawler`, then register it in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) (lines 50-66) within the `CrawlerFactory.CRAWLERS` dictionary.
- **Bug Fix**: Locate the failing test in `tests/` (e.g., [`test_tieba_extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/test_tieba_extractor.py)), modify the implementation in the relevant `store/*_impl.py` file, and ensure the test passes.
- **Documentation**: Update [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md) or add guides under `docs/`. The CDP mode guide (`docs/CDP模式使用指南.md`, lines 80-84) already contains a "贡献" section welcoming improvements to CDP documentation.

### 5. Validate Your Changes

Run the full test suite and linting tools:

```bash
uv run pytest
uv run ruff .

```

### 6. Submit Your Pull Request

Commit with clear, conventional commit messages and push to your fork:

```bash
git add .
git commit -m "feat: add TikTok crawler with QR code login support"
git push origin feature/add-tiktok-support

```

Open a Pull Request against the main repository. The maintainers may request additional test coverage, linting fixes, or documentation updates before merging.

## Code Standards and Best Practices

Adhering to these guidelines ensures your contribution aligns with the existing codebase quality.

**Code Style**: Follow PEP 8 conventions. The repository uses `ruff` and `black` via `uv lock`; always run `uv run ruff .` before committing.

**Type Hints**: Maintain type consistency across all public APIs. The `AbstractCrawler` class and its implementations use explicit type annotations that you must preserve.

**Testing**: Add unit tests covering new code paths. Use the `tmp_path` fixture for filesystem-related tests to ensure isolation.

**Documentation**: When adding platforms, document required configuration keys in `config/<platform>_config.py` and update the main [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md) with setup instructions.

**License Compliance**: All contributions fall under the project's **NON-COMMERCIAL LEARNING LICENSE 1.1**. Ensure you have read the LICENSE file before submitting code.

**CI Pipeline**: GitHub Actions automatically runs the test suite on each PR. Do not merge if the pipeline fails.

## Example: Adding a New Platform

Here is a minimal implementation pattern for adding a crawler called "example":

```python

# File: media_platform/example.py

from base.base_crawler import AbstractCrawler

class ExampleCrawler(AbstractCrawler):
    async def start(self) -> None:
        print("[Example] Crawling started")
        await self.fetch_some_data()
        print("[Example] Crawling finished")
    
    async def fetch_some_data(self):
        # Implementation details here

        pass

```

Register the new crawler in the factory:

```python

# Edit main.py (lines 50-66)

from media_platform.example import ExampleCrawler

CrawlerFactory.CRAWLERS["example"] = ExampleCrawler

```

Test your implementation locally:

```bash
uv run main.py --platform example --lt qrcode --type search

```

## Summary

- **MediaCrawler** uses a factory pattern in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to manage platform-specific crawlers inheriting from `AbstractCrawler` in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py).
- **Development setup** requires `uv` for dependency management and `pytest` for validation.
- **New platforms** require a class in `media_platform/`, registration in `CrawlerFactory`, and corresponding configuration in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).
- **Code quality** depends on `ruff` linting, type hints, and comprehensive testing under `tests/`.
- **Documentation** contributions are welcomed, particularly for CDP mode improvements referenced in `docs/CDP模式使用指南.md`.

## Frequently Asked Questions

### What license covers contributions to MediaCrawler?

All contributions are governed by the **NON-COMMERCIAL LEARNING LICENSE 1.1**. This license restricts commercial usage while permitting educational and research purposes. Review the LICENSE file in the repository root before submitting your pull request.

### How do I add support for a new social media platform?

Create a Python module in `media_platform/` containing a class that inherits from `AbstractCrawler`. Implement the required `start()` and `stop()` methods, then register your class in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) within the `CrawlerFactory.CRAWLERS` dictionary (lines 50-66). Add corresponding configuration defaults in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and create unit tests under `tests/`.

### What testing framework does MediaCrawler use?

The project uses **pytest** for unit and integration testing. Run `uv run pytest` to execute the full suite. When contributing bug fixes or new features, add tests that cover your specific code paths, using fixtures like `tmp_path` for filesystem operations to ensure test isolation.

### Where should I document my changes?

Update [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md) for user-facing features or platform additions. For technical details regarding Chrome DevTools Protocol (CDP) usage, contribute to `docs/CDP模式使用指南.md`, which already contains a dedicated "贡献" section (lines 80-84) welcoming improvements to the CDP implementation guide.