How to Contribute to NanmiCoder/MediaCrawler: A Complete Guide for Developers

To contribute to NanmiCoder/MediaCrawler, fork the repository, set up the development environment with uv, run the test suite, implement your changes following the modular architecture, and submit a pull request.

The NanmiCoder/MediaCrawler project is a Python-based, multi-platform web scraper designed to extract data from popular Chinese social media platforms like Xiaohongshu, Weibo, Douyin, and Zhihu. Whether you want to add support for a new platform, fix bugs, or improve existing crawlers, understanding how to contribute to NanmiCoder/MediaCrawler requires familiarity with its layered architecture and established workflow.

Understanding the MediaCrawler Architecture

Before writing code, you must understand how the repository is organized. The project follows a clean separation of concerns across eight distinct layers.

Core Components Overview

Layer Directory Responsibility
Core Engine base/base_crawler.py, main.py Defines the BaseCrawler abstract class and CLI entry point
Configuration config/base_config.py, config/xhs_config.py, etc. Platform-specific settings and global defaults
Platform Implementations media_platform/xhs/, media_platform/weibo/, etc. Concrete crawler implementations for each platform
Data Models model/m_xiaohongshu.py, model/m_weibo.py, etc. Pydantic-style data classes for extracted content
Storage Layer store/xhs/, store/weibo/, etc. Persistence to CSV, JSON, Excel, SQLite, or MySQL
Cache Layer cache/abs_cache.py, cache/redis_cache.py, etc. Session storage and rate-limit management
Proxy & IP Pool proxy/proxy_ip_pool.py, proxy/providers/ Rotating proxy handling for IP rotation
Utility Tools tools/browser_launcher.py, tools/cdp_browser.py Browser automation helpers

How Components Interact

The execution flow in main.py demonstrates how to contribute to NanmiCoder/MediaCrawler effectively:

  1. CLI arguments parsed in main.py determine platform and crawl mode
  2. Configuration loaded from files like config/xhs_config.py
  3. Platform-specific crawler instantiated (subclass of BaseCrawler from base/base_crawler.py)
  4. Login session retrieved from cache (cache/redis_cache.py) or fresh login via browser launcher
  5. Data extracted, mapped to models (model/m_xiaohongshu.py), and persisted via store layer
  6. Proxy handling and rate-limit checks applied throughout

Setting Up Your Development Environment

Prerequisites and Installation

The project strongly recommends using uv for dependency management. This ensures reproducible builds across contributor environments.


# Fork the repository on GitHub, then clone your fork

git clone https://github.com/YOUR_USERNAME/MediaCrawler.git
cd MediaCrawler

# Install dependencies with uv

uv sync

# Install Playwright browsers if needed for your contribution

uv run playwright install

Running the Test Suite

Always verify the baseline before making changes. The test suite lives in tests/ and test/ directories.


# Run all tests

uv run pytest

# Or with standard Python

python -m pytest

Contribution Workflow: Step by Step

Follow this proven workflow when you contribute to NanmiCoder/MediaCrawler:

  1. Fork and clone the repository to your GitHub account
  2. Create a feature branch: git checkout -b feature/your-contribution
  3. Implement your changes following architectural patterns
  4. Add or update tests in the appropriate tests/ subdirectory
  5. Run the full test suite to catch regressions: uv run pytest
  6. Commit with descriptive messages referencing any related issues
  7. Push to your fork and open a Pull Request against NanmiCoder/MediaCrawler:main
  8. Ensure CI passes before requesting review

Common Contribution Patterns

Adding a New Platform Crawler

The most impactful way to contribute to NanmiCoder/MediaCrawler is adding support for new social media platforms. Here's the complete implementation pattern.

Create the crawler class in a new directory:


# media_platform/foobar/foobar_crawler.py

from base.base_crawler import BaseCrawler
from config.foobar_config import FoobarConfig
from model.m_foobar import FoobarPost
from store.foobar import FoobarStore
from tools.browser_launcher import launch_browser


class FoobarCrawler(BaseCrawler):
    def __init__(self, cfg: FoobarConfig):
        super().__init__(cfg)
        self.store = FoobarStore()

    async def login(self):
        """Authenticate using CDP or Playwright."""
        self.browser = await launch_browser(self.cfg)
        await self.browser.goto(self.cfg.LOGIN_URL)
        # Implement QR-code login or cookie injection here

        pass

    async def fetch_posts(self, keyword: str):
        """Paginate through search results and persist posts."""
        async for page in self.cdp_browser.paginate(
            self.cfg.SEARCH_URL, 
            params={"q": keyword}
        ):
            for raw in page.json()["data"]:
                post = FoobarPost.from_raw(raw)
                await self.store.save(post)

Required supporting files for a complete platform addition:

  • config/foobar_config.py — platform-specific constants and credentials
  • model/m_foobar.py — Pydantic data model with fields like post_id, author, content, timestamp
  • store/foobar/ — persistence handlers for your chosen output formats
  • Update main.py to register the new crawler in the CLI dispatcher

Extending the Cache Layer

Contributions to infrastructure components follow similar patterns. To add a custom LRU cache:


# cache/custom_lru_cache.py

from cache.abs_cache import AbsCache
from collections import OrderedDict
from typing import Any


class CustomLRUCache(AbsCache):
    def __init__(self, maxsize: int = 1024):
        self.store = OrderedDict()
        self.maxsize = maxsize

    def get(self, key: str) -> Any:
        value = self.store.pop(key, None)
        if value is not None:
            # Re-insert to mark as most-recently used

            self.store[key] = value
        return value

    def set(self, key: str, value: Any, ttl: int = 0) -> None:
        self.store[key] = value
        if len(self.store) > self.maxsize:
            # Remove least-recently used entry

            self.store.popitem(last=False)

Register your implementation in cache/cache_factory.py to make it selectable via configuration.

Key Source Files Reference

File Purpose Direct Link
base/base_crawler.py Abstract base class all crawlers inherit View source
main.py CLI entry point and crawler dispatcher View source
config/base_config.py Global configuration defaults View source
config/xhs_config.py Example platform-specific config View source
model/m_xiaohongshu.py Data model for Xiaohongshu posts View source
cache/redis_cache.py Redis-backed cache implementation View source
tools/browser_launcher.py Browser automation helper View source
docs/项目代码结构.md Visual architecture overview (Chinese) View source

Summary

To successfully contribute to NanmiCoder/MediaCrawler, remember these essentials:

  • Use uv for dependency management and uv run pytest for testing
  • Inherit from BaseCrawler in base/base_crawler.py when adding new platforms
  • Follow the eight-layer architecture: core → config → platform → model → store → cache → proxy → tools
  • Always run tests before and after your changes
  • Register new crawlers in main.py to expose them via CLI

Frequently Asked Questions

Do I need to understand Chinese to contribute to NanmiCoder/MediaCrawler?

No. While the repository includes Chinese documentation like docs/项目代码结构.md and targets Chinese social media platforms, all source code, variable names, and docstrings use English. Platform-specific implementations handle Chinese content extraction automatically. English-only contributors can confidently submit bug fixes, infrastructure improvements, and new platform support.

What Python version does MediaCrawler require?

The project uses modern Python features. Check pyproject.toml in the repository root for the exact version specification. As of the latest commit, Python 3.10+ is typically required to support async/await patterns and type hint syntax used throughout base/base_crawler.py and platform implementations.

How do I test my crawler without hitting rate limits?

The cache/ layer provides mechanisms to store and reuse login sessions. For development, configure ENABLE_CDP_MODE in your platform config to use local browser automation rather than direct API calls. Additionally, the proxy/proxy_ip_pool.py system can route requests through rotating proxies during integration testing. Use small LIMIT values in config files to minimize requests while validating functionality.

Can I contribute documentation or translations instead of code?

Yes. Documentation contributions follow the same workflow: fork, branch, commit, and pull request. The README.md and docs/ directory welcome improvements. For translations, maintain parallel structure to preserve deep links. Code comments in main.py and base/base_crawler.py also benefit from clarification—submit these as dedicated "docs:" commits for easy review.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →