# How to Contribute New Platform Crawlers to MediaCrawler: A Step-by-Step Guide

> Learn how to contribute new platform crawlers to the MediaCrawler project. This step-by-step guide covers implementing the AbstractCrawler interface and registering your crawler.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-03

---

**To add a new platform crawler to MediaCrawler, you must create a package under `media_platform/`, implement the `AbstractCrawler` interface from [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py), register the class in the `CrawlerFactory` located in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py), and provide platform-specific configuration, storage, and tests.**

MediaCrawler is an open-source social media scraping framework built on a clean, plugin-based architecture that makes extending support to new platforms straightforward. By following the established patterns in the codebase, you can contribute to the development of new platform crawlers while maintaining consistency with existing implementations like Tieba and Zhihu. This guide walks you through the concrete steps required to integrate a new platform, referencing actual source files and interfaces from the NanmiCoder/MediaCrawler repository.

## Understanding the Architecture

Before writing code, you need to understand how MediaCrawler structures its components. The framework separates concerns into distinct layers: an abstract interface that all crawlers must implement, platform-specific logic for navigation and data extraction, and generic utilities for storage and browser management.

### The AbstractCrawler Interface

Every platform crawler must inherit from **`AbstractCrawler`**, defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This class establishes the asynchronous contract that the main execution loop expects, including the `start()`, `search()`, and `launch_browser()` methods. When you contribute to the development of new platform crawlers, your concrete implementation must override these methods to handle the specific login flows, API endpoints, and pagination logic of your target site.

### Platform-Specific Components

Each platform lives in its own package under `media_platform/<platform_name>/`. A typical implementation includes:

- **[`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py)** – The main crawler class that orchestrates the browser and API client
- **[`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py)** – An async HTTP wrapper or Playwright page controller that handles raw requests
- **[`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py)** – Encapsulates authentication flows (QR codes, mobile login, cookie injection)
- **[`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py)** – Contains extractor functions that parse raw HTML/JSON into structured data

### Storage and Configuration

Data persistence follows the **`AbstractStore`** pattern. You can reuse generic implementations like `AsyncFileWriter` for CSV/JSONL output, or create a custom store in `store/<platform>/_store_impl.py` if you need specialized database schemas. Platform-specific settings (endpoints, rate limits, headers) belong in `config/<platform>_config.py`.

## Step-by-Step Implementation Guide

Follow these steps to add a new platform crawler to the MediaCrawler ecosystem.

### Step 1: Create the Platform Package

Create a new directory for your platform:

```bash
mkdir media_platform/myplatform
touch media_platform/myplatform/__init__.py

```

Inside this package, you will create [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py), and [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py).

### Step 2: Implement the Crawler Class

In [`media_platform/myplatform/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/myplatform/core.py), define a class that inherits from `AbstractCrawler`:

```python
from base.base_crawler import AbstractCrawler
from tools.browser_launcher import launch_browser
from .client import MyPlatformClient
from .login import MyPlatformLogin
from .help import MyPlatformExtractor
from store import myplatform as myplatform_store
import config
import utils

class MyPlatformCrawler(AbstractCrawler):
    def __init__(self) -> None:
        self.base_url = "https://myplatform.example.com"
        self.user_agent = utils.get_user_agent()
        self._extractor = MyPlatformExtractor()
        self.client: MyPlatformClient | None = None
        self.browser_context = None

    async def start(self) -> None:
        async with async_playwright() as pw:
            self.browser_context = await self.launch_browser(
                pw.chromium, None, self.user_agent, headless=config.HEADLESS
            )
            page = await self.browser_context.new_page()
            await page.goto(self.base_url)

            if not await self._is_logged_in():
                login = MyPlatformLogin(page, self.browser_context)
                await login.begin()

            self.client = MyPlatformClient(httpx_proxy=None)

            if config.CRAWLER_TYPE == "search":
                await self.search()
            elif config.CRAWLER_TYPE == "detail":
                await self.get_detail()

    async def search(self) -> None:
        for kw in config.KEYWORDS.split(","):
            posts = await self.client.search_posts(keyword=kw, page=1)
            for post in posts:
                content = self._extractor.extract_content(post)
                await myplatform_store.MyPlatformCsvStoreImplement().store_content(content)

```

### Step 3: Add Helper Modules

Create the supporting infrastructure for your crawler:

- **Client** ([`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py)): Wrap the platform's HTTP API or Playwright interactions
- **Login** ([`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py)): Handle authentication using cookies, QR codes, or mobile verification
- **Extractor** ([`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py)): Transform raw responses into the data models defined in `model/`

### Step 4: Define Data Models

If your platform uses unique data structures, add models to `model/` following the pattern of [`model/m_baidu_tieba.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_baidu_tieba.py). These dataclasses or Pydantic models standardize the schema for content, comments, and user profiles across the codebase.

### Step 5: Implement Storage

You can reuse existing storage utilities or create a custom implementation. For example, to support CSV output, create a store class in [`store/myplatform/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/myplatform/_store_impl.py):

```python
from base.base_store import AbstractStore

class MyPlatformCsvStoreImplement(AbstractStore):
    async def store_content(self, content_item):
        # Implementation for persisting content to CSV

        pass

```

### Step 6: Register with CrawlerFactory

Edit [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) to register your crawler in the `CrawlerFactory`:

```python
from media_platform.myplatform.core import MyPlatformCrawler

class CrawlerFactory:
    CRAWLERS = {
        "tieba": TieBaCrawler,
        "zhihu": ZhihuCrawler,
        "myplatform": MyPlatformCrawler,  # Add your entry here

    }

```

This registration allows the CLI and API entry points to instantiate your crawler by name.

### Step 7: Add Configuration

Copy an existing configuration file (e.g., [`config/tieba_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/tieba_config.py)) to [`config/myplatform_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/myplatform_config.py) and adjust the constants for your platform's endpoints, rate limits, and authentication credentials. Ensure the file is imported in [`config/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/__init__.py) so it loads at runtime.

### Step 8: Write Tests

Place unit and functional tests under `tests/` to verify your client, login flow, and storage implementations. Use fixtures from [`conftest.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/conftest.py) to mock HTTP responses or spin up in-memory databases:

```python
import pytest
from media_platform.myplatform.client import MyPlatformClient

@pytest.mark.asyncio
async def test_search_posts(httpx_mock):
    httpx_mock.add_response(
        url="https://api.myplatform.example.com/search?kw=test&page=1",
        json={"posts": [{"id": "1", "title": "Demo"}]}
    )
    client = MyPlatformClient()
    posts = await client.search_posts(keyword="test", page=1)
    assert len(posts) == 1
    assert posts[0]["title"] == "Demo"

```

Run `pytest -q` to ensure all tests pass before submitting your contribution.

## Key Implementation Files

When contributing to the development of new platform crawlers, reference these existing files to understand the patterns:

- **[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)** – Defines the `AbstractCrawler` interface that all implementations must extend
- **[`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py)** – Complete reference implementation showing browser management and data extraction
- **[`store/tieba/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/tieba/_store_impl.py)** – Example of how to wire storage implementations to the crawler
- **[`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py)** – Contains the `CrawlerFactory` that maps platform names to crawler classes
- **[`tools/browser_launcher.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/browser_launcher.py)** and **[`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py)** – Helpers for launching Playwright browsers in standard or CDP mode

## Summary

Contributing a new platform crawler to MediaCrawler requires implementing a structured interface and following the repository's architectural conventions:

- **Inherit from `AbstractCrawler`** in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) and implement the `start()` and `search()` methods
- **Organize platform code** into `media_platform/<name>/` with separate modules for the client, login, and extraction logic
- **Register the crawler** in the `CrawlerFactory` dictionary within [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py)
- **Provide configuration** in `config/<platform>_config.py` for endpoints and runtime settings
- **Implement tests** under `tests/` and ensure the full suite passes with `pytest -q`

## Frequently Asked Questions

### What is the minimum code required to add a new platform crawler?

You must create a class inheriting from `AbstractCrawler` with implemented `start()` and `search()` methods, register it in `CrawlerFactory` (located in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py)), and provide a configuration file. While client, login, and storage modules are optional for minimal functionality, they are strongly recommended for production-ready implementations.

### How does MediaCrawler handle browser automation for new platforms?

The `AbstractCrawler` base class provides `launch_browser()` and `launch_browser_with_cdp()` methods that wrap Playwright. You can use the helpers in [`tools/browser_launcher.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/browser_launcher.py) for standard headless/headed modes or [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) for Chrome DevTools Protocol connections. Your crawler should call these in the `start()` method to initialize the browser context before navigating to the target site.

### Where should I store scraped data when contributing a new crawler?

You can reuse the generic `AsyncFileWriter` for CSV or JSONL output, or implement a custom store by subclassing `AbstractStore` in a new file under `store/<platform>/_store_impl.py`. The crawler instantiates the store class and calls its methods (e.g., `store_content()`) during the extraction phase.

### How do I test my new platform crawler before submitting a pull request?

Write unit tests for your client and extractor functions in `tests/`, using `pytest` and fixtures from [`conftest.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/conftest.py) to mock HTTP responses. Run the full test suite with `pytest -q` to ensure your changes do not break existing crawlers. Additionally, test the integration manually by running the crawler via the CLI with your platform name as the target.