How to Contribute to NanmiCoder/MediaCrawler: A Complete Guide for Developers
To contribute to NanmiCoder/MediaCrawler, fork the repository, set up the development environment with uv, run the test suite, implement your changes following the modular architecture, and submit a pull request.
The NanmiCoder/MediaCrawler project is a Python-based, multi-platform web scraper designed to extract data from popular Chinese social media platforms like Xiaohongshu, Weibo, Douyin, and Zhihu. Whether you want to add support for a new platform, fix bugs, or improve existing crawlers, understanding how to contribute to NanmiCoder/MediaCrawler requires familiarity with its layered architecture and established workflow.
Understanding the MediaCrawler Architecture
Before writing code, you must understand how the repository is organized. The project follows a clean separation of concerns across eight distinct layers.
Core Components Overview
| Layer | Directory | Responsibility |
|---|---|---|
| Core Engine | base/base_crawler.py, main.py |
Defines the BaseCrawler abstract class and CLI entry point |
| Configuration | config/base_config.py, config/xhs_config.py, etc. |
Platform-specific settings and global defaults |
| Platform Implementations | media_platform/xhs/, media_platform/weibo/, etc. |
Concrete crawler implementations for each platform |
| Data Models | model/m_xiaohongshu.py, model/m_weibo.py, etc. |
Pydantic-style data classes for extracted content |
| Storage Layer | store/xhs/, store/weibo/, etc. |
Persistence to CSV, JSON, Excel, SQLite, or MySQL |
| Cache Layer | cache/abs_cache.py, cache/redis_cache.py, etc. |
Session storage and rate-limit management |
| Proxy & IP Pool | proxy/proxy_ip_pool.py, proxy/providers/ |
Rotating proxy handling for IP rotation |
| Utility Tools | tools/browser_launcher.py, tools/cdp_browser.py |
Browser automation helpers |
How Components Interact
The execution flow in main.py demonstrates how to contribute to NanmiCoder/MediaCrawler effectively:
- CLI arguments parsed in
main.pydetermine platform and crawl mode - Configuration loaded from files like
config/xhs_config.py - Platform-specific crawler instantiated (subclass of
BaseCrawlerfrombase/base_crawler.py) - Login session retrieved from cache (
cache/redis_cache.py) or fresh login via browser launcher - Data extracted, mapped to models (
model/m_xiaohongshu.py), and persisted via store layer - Proxy handling and rate-limit checks applied throughout
Setting Up Your Development Environment
Prerequisites and Installation
The project strongly recommends using uv for dependency management. This ensures reproducible builds across contributor environments.
# Fork the repository on GitHub, then clone your fork
git clone https://github.com/YOUR_USERNAME/MediaCrawler.git
cd MediaCrawler
# Install dependencies with uv
uv sync
# Install Playwright browsers if needed for your contribution
uv run playwright install
Running the Test Suite
Always verify the baseline before making changes. The test suite lives in tests/ and test/ directories.
# Run all tests
uv run pytest
# Or with standard Python
python -m pytest
Contribution Workflow: Step by Step
Follow this proven workflow when you contribute to NanmiCoder/MediaCrawler:
- Fork and clone the repository to your GitHub account
- Create a feature branch:
git checkout -b feature/your-contribution - Implement your changes following architectural patterns
- Add or update tests in the appropriate
tests/subdirectory - Run the full test suite to catch regressions:
uv run pytest - Commit with descriptive messages referencing any related issues
- Push to your fork and open a Pull Request against
NanmiCoder/MediaCrawler:main - Ensure CI passes before requesting review
Common Contribution Patterns
Adding a New Platform Crawler
The most impactful way to contribute to NanmiCoder/MediaCrawler is adding support for new social media platforms. Here's the complete implementation pattern.
Create the crawler class in a new directory:
# media_platform/foobar/foobar_crawler.py
from base.base_crawler import BaseCrawler
from config.foobar_config import FoobarConfig
from model.m_foobar import FoobarPost
from store.foobar import FoobarStore
from tools.browser_launcher import launch_browser
class FoobarCrawler(BaseCrawler):
def __init__(self, cfg: FoobarConfig):
super().__init__(cfg)
self.store = FoobarStore()
async def login(self):
"""Authenticate using CDP or Playwright."""
self.browser = await launch_browser(self.cfg)
await self.browser.goto(self.cfg.LOGIN_URL)
# Implement QR-code login or cookie injection here
pass
async def fetch_posts(self, keyword: str):
"""Paginate through search results and persist posts."""
async for page in self.cdp_browser.paginate(
self.cfg.SEARCH_URL,
params={"q": keyword}
):
for raw in page.json()["data"]:
post = FoobarPost.from_raw(raw)
await self.store.save(post)
Required supporting files for a complete platform addition:
config/foobar_config.py— platform-specific constants and credentialsmodel/m_foobar.py— Pydantic data model with fields likepost_id,author,content,timestampstore/foobar/— persistence handlers for your chosen output formats- Update
main.pyto register the new crawler in the CLI dispatcher
Extending the Cache Layer
Contributions to infrastructure components follow similar patterns. To add a custom LRU cache:
# cache/custom_lru_cache.py
from cache.abs_cache import AbsCache
from collections import OrderedDict
from typing import Any
class CustomLRUCache(AbsCache):
def __init__(self, maxsize: int = 1024):
self.store = OrderedDict()
self.maxsize = maxsize
def get(self, key: str) -> Any:
value = self.store.pop(key, None)
if value is not None:
# Re-insert to mark as most-recently used
self.store[key] = value
return value
def set(self, key: str, value: Any, ttl: int = 0) -> None:
self.store[key] = value
if len(self.store) > self.maxsize:
# Remove least-recently used entry
self.store.popitem(last=False)
Register your implementation in cache/cache_factory.py to make it selectable via configuration.
Key Source Files Reference
| File | Purpose | Direct Link |
|---|---|---|
base/base_crawler.py |
Abstract base class all crawlers inherit | View source |
main.py |
CLI entry point and crawler dispatcher | View source |
config/base_config.py |
Global configuration defaults | View source |
config/xhs_config.py |
Example platform-specific config | View source |
model/m_xiaohongshu.py |
Data model for Xiaohongshu posts | View source |
cache/redis_cache.py |
Redis-backed cache implementation | View source |
tools/browser_launcher.py |
Browser automation helper | View source |
docs/项目代码结构.md |
Visual architecture overview (Chinese) | View source |
Summary
To successfully contribute to NanmiCoder/MediaCrawler, remember these essentials:
- Use
uvfor dependency management anduv run pytestfor testing - Inherit from
BaseCrawlerinbase/base_crawler.pywhen adding new platforms - Follow the eight-layer architecture: core → config → platform → model → store → cache → proxy → tools
- Always run tests before and after your changes
- Register new crawlers in
main.pyto expose them via CLI
Frequently Asked Questions
Do I need to understand Chinese to contribute to NanmiCoder/MediaCrawler?
No. While the repository includes Chinese documentation like docs/项目代码结构.md and targets Chinese social media platforms, all source code, variable names, and docstrings use English. Platform-specific implementations handle Chinese content extraction automatically. English-only contributors can confidently submit bug fixes, infrastructure improvements, and new platform support.
What Python version does MediaCrawler require?
The project uses modern Python features. Check pyproject.toml in the repository root for the exact version specification. As of the latest commit, Python 3.10+ is typically required to support async/await patterns and type hint syntax used throughout base/base_crawler.py and platform implementations.
How do I test my crawler without hitting rate limits?
The cache/ layer provides mechanisms to store and reuse login sessions. For development, configure ENABLE_CDP_MODE in your platform config to use local browser automation rather than direct API calls. Additionally, the proxy/proxy_ip_pool.py system can route requests through rotating proxies during integration testing. Use small LIMIT values in config files to minimize requests while validating functionality.
Can I contribute documentation or translations instead of code?
Yes. Documentation contributions follow the same workflow: fork, branch, commit, and pull request. The README.md and docs/ directory welcome improvements. For translations, maintain parallel structure to preserve deep links. Code comments in main.py and base/base_crawler.py also benefit from clarification—submit these as dedicated "docs:" commits for easy review.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →