# How to Contribute to NanmiCoder/MediaCrawler: A Complete Guide for Developers

> Learn how to contribute to NanmiCoderMediaCrawler. Fork the repo, set up your environment with uv, run tests, implement changes, and submit a pull request to join the development.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-12

---

**To contribute to NanmiCoder/MediaCrawler, fork the repository, set up the development environment with `uv`, run the test suite, implement your changes following the modular architecture, and submit a pull request.**

The NanmiCoder/MediaCrawler project is a Python-based, multi-platform web scraper designed to extract data from popular Chinese social media platforms like Xiaohongshu, Weibo, Douyin, and Zhihu. Whether you want to add support for a new platform, fix bugs, or improve existing crawlers, understanding how to contribute to NanmiCoder/MediaCrawler requires familiarity with its layered architecture and established workflow.

## Understanding the MediaCrawler Architecture

Before writing code, you must understand how the repository is organized. The project follows a clean separation of concerns across eight distinct layers.

### Core Components Overview

| Layer | Directory | Responsibility |
|-------|-----------|----------------|
| **Core Engine** | [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py), [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) | Defines the `BaseCrawler` abstract class and CLI entry point |
| **Configuration** | [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py), etc. | Platform-specific settings and global defaults |
| **Platform Implementations** | `media_platform/xhs/`, `media_platform/weibo/`, etc. | Concrete crawler implementations for each platform |
| **Data Models** | [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py), [`model/m_weibo.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_weibo.py), etc. | Pydantic-style data classes for extracted content |
| **Storage Layer** | `store/xhs/`, `store/weibo/`, etc. | Persistence to CSV, JSON, Excel, SQLite, or MySQL |
| **Cache Layer** | [`cache/abs_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/abs_cache.py), [`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py), etc. | Session storage and rate-limit management |
| **Proxy & IP Pool** | [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py), `proxy/providers/` | Rotating proxy handling for IP rotation |
| **Utility Tools** | [`tools/browser_launcher.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/browser_launcher.py), [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) | Browser automation helpers |

### How Components Interact

The execution flow in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) demonstrates how to contribute to NanmiCoder/MediaCrawler effectively:

1. CLI arguments parsed in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) determine platform and crawl mode
2. Configuration loaded from files like [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py)
3. Platform-specific crawler instantiated (subclass of `BaseCrawler` from [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py))
4. Login session retrieved from cache ([`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py)) or fresh login via browser launcher
5. Data extracted, mapped to models ([`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py)), and persisted via store layer
6. Proxy handling and rate-limit checks applied throughout

## Setting Up Your Development Environment

### Prerequisites and Installation

The project **strongly recommends** using `uv` for dependency management. This ensures reproducible builds across contributor environments.

```bash

# Fork the repository on GitHub, then clone your fork

git clone https://github.com/YOUR_USERNAME/MediaCrawler.git
cd MediaCrawler

# Install dependencies with uv

uv sync

# Install Playwright browsers if needed for your contribution

uv run playwright install

```

### Running the Test Suite

Always verify the baseline before making changes. The test suite lives in `tests/` and `test/` directories.

```bash

# Run all tests

uv run pytest

# Or with standard Python

python -m pytest

```

## Contribution Workflow: Step by Step

Follow this proven workflow when you contribute to NanmiCoder/MediaCrawler:

1. **Fork and clone** the repository to your GitHub account
2. **Create a feature branch**: `git checkout -b feature/your-contribution`
3. **Implement your changes** following architectural patterns
4. **Add or update tests** in the appropriate `tests/` subdirectory
5. **Run the full test suite** to catch regressions: `uv run pytest`
6. **Commit with descriptive messages** referencing any related issues
7. **Push to your fork** and open a Pull Request against `NanmiCoder/MediaCrawler:main`
8. **Ensure CI passes** before requesting review

## Common Contribution Patterns

### Adding a New Platform Crawler

The most impactful way to contribute to NanmiCoder/MediaCrawler is adding support for new social media platforms. Here's the complete implementation pattern.

Create the crawler class in a new directory:

```python

# media_platform/foobar/foobar_crawler.py

from base.base_crawler import BaseCrawler
from config.foobar_config import FoobarConfig
from model.m_foobar import FoobarPost
from store.foobar import FoobarStore
from tools.browser_launcher import launch_browser


class FoobarCrawler(BaseCrawler):
    def __init__(self, cfg: FoobarConfig):
        super().__init__(cfg)
        self.store = FoobarStore()

    async def login(self):
        """Authenticate using CDP or Playwright."""
        self.browser = await launch_browser(self.cfg)
        await self.browser.goto(self.cfg.LOGIN_URL)
        # Implement QR-code login or cookie injection here

        pass

    async def fetch_posts(self, keyword: str):
        """Paginate through search results and persist posts."""
        async for page in self.cdp_browser.paginate(
            self.cfg.SEARCH_URL, 
            params={"q": keyword}
        ):
            for raw in page.json()["data"]:
                post = FoobarPost.from_raw(raw)
                await self.store.save(post)

```

Required supporting files for a complete platform addition:

- [`config/foobar_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/foobar_config.py) — platform-specific constants and credentials
- [`model/m_foobar.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_foobar.py) — Pydantic data model with fields like `post_id`, `author`, `content`, `timestamp`
- `store/foobar/` — persistence handlers for your chosen output formats
- Update [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to register the new crawler in the CLI dispatcher

### Extending the Cache Layer

Contributions to infrastructure components follow similar patterns. To add a custom LRU cache:

```python

# cache/custom_lru_cache.py

from cache.abs_cache import AbsCache
from collections import OrderedDict
from typing import Any


class CustomLRUCache(AbsCache):
    def __init__(self, maxsize: int = 1024):
        self.store = OrderedDict()
        self.maxsize = maxsize

    def get(self, key: str) -> Any:
        value = self.store.pop(key, None)
        if value is not None:
            # Re-insert to mark as most-recently used

            self.store[key] = value
        return value

    def set(self, key: str, value: Any, ttl: int = 0) -> None:
        self.store[key] = value
        if len(self.store) > self.maxsize:
            # Remove least-recently used entry

            self.store.popitem(last=False)

```

Register your implementation in [`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py) to make it selectable via configuration.

## Key Source Files Reference

| File | Purpose | Direct Link |
|------|---------|-------------|
| [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) | Abstract base class all crawlers inherit | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) |
| [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) | CLI entry point and crawler dispatcher | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) |
| [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) | Global configuration defaults | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) |
| [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) | Example platform-specific config | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) |
| [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py) | Data model for Xiaohongshu posts | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py) |
| [`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py) | Redis-backed cache implementation | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py) |
| [`tools/browser_launcher.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/browser_launcher.py) | Browser automation helper | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/browser_launcher.py) |
| `docs/项目代码结构.md` | Visual architecture overview (Chinese) | [View source](https://github.com/NanmiCoder/MediaCrawler/blob/main/docs/项目代码结构.md) |

## Summary

To successfully contribute to NanmiCoder/MediaCrawler, remember these essentials:

- **Use `uv`** for dependency management and `uv run pytest` for testing
- **Inherit from `BaseCrawler`** in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) when adding new platforms
- **Follow the eight-layer architecture**: core → config → platform → model → store → cache → proxy → tools
- **Always run tests** before and after your changes
- **Register new crawlers** in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to expose them via CLI

## Frequently Asked Questions

### Do I need to understand Chinese to contribute to NanmiCoder/MediaCrawler?

No. While the repository includes Chinese documentation like `docs/项目代码结构.md` and targets Chinese social media platforms, all source code, variable names, and docstrings use English. Platform-specific implementations handle Chinese content extraction automatically. English-only contributors can confidently submit bug fixes, infrastructure improvements, and new platform support.

### What Python version does MediaCrawler require?

The project uses modern Python features. Check [`pyproject.toml`](https://github.com/NanmiCoder/MediaCrawler/blob/main/pyproject.toml) in the repository root for the exact version specification. As of the latest commit, Python 3.10+ is typically required to support `async`/`await` patterns and type hint syntax used throughout [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) and platform implementations.

### How do I test my crawler without hitting rate limits?

The `cache/` layer provides mechanisms to store and reuse login sessions. For development, configure `ENABLE_CDP_MODE` in your platform config to use local browser automation rather than direct API calls. Additionally, the [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py) system can route requests through rotating proxies during integration testing. Use small `LIMIT` values in config files to minimize requests while validating functionality.

### Can I contribute documentation or translations instead of code?

Yes. Documentation contributions follow the same workflow: fork, branch, commit, and pull request. The [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md) and `docs/` directory welcome improvements. For translations, maintain parallel structure to preserve deep links. Code comments in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) and [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) also benefit from clarification—submit these as dedicated "docs:" commits for easy review.