# How to Add Support for a New Platform to NanmiCoder/MediaCrawler: A Step-by-Step Guide

> Easily add a new platform to MediaCrawler by implementing AbstractCrawler, creating client wrappers and store modules, and registering in CrawlerFactory. Follow our step-by-step guide.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-03

---

**You can add support for a new platform to MediaCrawler by implementing the `AbstractCrawler` interface in a new package under `media_platform/<platform>/`, creating a client wrapper and store module, and registering the class in `CrawlerFactory` within [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py).**

MediaCrawler is an open-source social media scraping framework built around a clean, plugin-based architecture. Extending the tool to support additional platforms requires no changes to the core infrastructure, as the shared utilities for browser automation, proxy handling, and data persistence are already abstracted. This guide walks through the concrete implementation steps based on the actual source code structure of the NanmiCoder/MediaCrawler repository.

## Architecture Overview

The foundation of MediaCrawler's extensibility lies in the **AbstractCrawler** base class defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This interface declares the common async methods that every platform crawler must implement, including `start`, `search`, and `launch_browser`.

Platform-specific implementations live in isolated packages under `media_platform/<platform>/`. The **CrawlerFactory** in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) maintains a simple dictionary mapping CLI platform flags to concrete crawler classes, enabling dynamic instantiation at runtime. Because heavy-lifting components like Playwright browser management, CDP fallback handling, async concurrency limits, and Excel/DB output flushing are shared across all crawlers, you only need to write platform-specific API interaction logic.

## Step-by-Step Implementation Guide

Adding a new platform requires creating four primary components and updating two registry files. Follow these steps to integrate a new media source.

### 1. Create the Platform Package Structure

Create a new directory `media_platform/<new_platform>/` containing [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), and optionally [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py). This mirrors the structure used by existing implementations like Zhihu (`media_platform/zhihu/`).

### 2. Implement the Core Crawler Class

In `media_platform/<new_platform>/core.py`, subclass `AbstractCrawler` and implement the required async methods. The `start` method serves as the entry point, handling browser initialization, client creation, authentication, and dispatching to crawling modes.

```python

# media_platform/<new_platform>/core.py

from base.base_crawler import AbstractCrawler
from tools import utils
from var import crawler_type_var, source_keyword_var
from .client import <NewPlatform>Client
from .login import <NewPlatform>Login
from store import <new_platform> as store

class <NewPlatform>Crawler(AbstractCrawler):
    async def start(self) -> None:
        # Launch browser (standard or CDP mode)

        async with async_playwright() as playwright:
            self.browser_context = await self.launch_browser_with_cdp(
                playwright, None, self.user_agent, headless=config.CDP_HEADLESS
            ) if config.ENABLE_CDP_MODE else await self.launch_browser(
                playwright.chromium, None, self.user_agent, headless=config.HEADLESS
            )
            self.context_page = await self.browser_context.new_page()
            await self.context_page.goto(self.index_url)

            # Create API client and handle login

            self.client = await self.create_client()
            if not await self.client.pong():
                login = <NewPlatform>Login(...)
                await login.begin()
                await self.client.update_cookies(self.browser_context, self.cookie_urls)

            # Dispatch to crawling mode

            crawler_type_var.set(config.CRAWLER_TYPE)
            if config.CRAWLER_TYPE == "search":
                await self.search()
            elif config.CRAWLER_TYPE == "detail":
                await self.get_specified_notes()
            elif config.CRAWLER_TYPE == "creator":
                await self.get_creators_and_notes()

```

Implement the `search` method to fetch content lists using `config.KEYWORDS` and `config.CRAWLER_MAX_NOTES_COUNT`, then delegate storage to the platform-specific store module.

### 3. Build the HTTP Client Wrapper

Create `media_platform/<new_platform>/client.py` to handle platform-specific HTTP API calls or Playwright page interactions. The client should expose standard methods like `pong`, `get_note_by_keyword`, `get_note_detail`, and `update_cookies`.

```python

# media_platform/<new_platform>/client.py

class <NewPlatform>Client:
    def __init__(self, *, proxy: Optional[str], headers: Dict[str, str], playwright_page: Page, cookie_dict: Dict):
        self.proxy = proxy
        self.headers = headers
        self.page = playwright_page
        self.cookies = cookie_dict

    async def pong(self) -> bool:
        # Simple health-check request

        resp = await httpx.get("https://api.<new_platform>.com/ping", headers=self.headers, proxy=self.proxy)
        return resp.status_code == 200

```

Reference [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) for the expected interface patterns.

### 4. Implement the Storage Layer

Create `store/<new_platform>/__init__.py` and concrete implementations to handle persistence. Implement the abstract store interfaces `store_content`, `store_comment`, and `store_creator`, following the pattern in [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py).

```python

# store/<new_platform>/__init__.py

async def update_<new_platform>_content(content: Dict) -> None:
    # Convert platform-specific model to generic dict if needed

    await store_base.save(content, platform="<new_platform>")

```

### 5. Add Login Support (If Required)

If the platform requires authentication, create `media_platform/<new_platform>/login.py` implementing the `AbstractLogin` contract. This handles QR-code scanning or credential-based authentication before the crawling session begins.

### 6. Register in CrawlerFactory

Import your new crawler in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) and add it to the `CRAWLERS` dictionary within `CrawlerFactory`. The key represents the value users pass to the `--platform` CLI flag.

```python

# main.py

from media_platform.<new_platform> import <NewPlatform>Crawler

CRAWLERS = {
    "xhs": XiaoHongShuCrawler,
    "dy": DouYinCrawler,
    # ... existing platforms ...

    "<new_platform_key>": <NewPlatform>Crawler,
}

```

### 7. Expose Configuration Options

Create `config/<new_platform>_config.py` for platform-specific settings like API keys or pagination limits. Global options such as `ENABLE_IP_PROXY`, `ENABLE_CDP_MODE`, and `CRAWLER_MAX_SLEEP_SEC` are automatically available via the shared `config` module.

### 8. Write Unit Tests

Add tests under `tests/` that import the new crawler and verify public methods execute without error. Mock network calls during testing, following the patterns in [`test_cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/test_cdp_browser.py) or [`test_static_proxy_provider.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/test_static_proxy_provider.py).

## Key Files and Their Roles

| File | Purpose |
|------|---------|
| [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) | Defines the `AbstractCrawler` interface that every platform must implement. |
| `media_platform/<new_platform>/core.py` | Platform-specific crawler implementation containing the main crawling logic. |
| `media_platform/<new_platform>/client.py` | HTTP client wrapper for platform API interactions. |
| `media_platform/<new_platform>/login.py` | Optional login helper respecting the `AbstractLogin` contract. |
| `store/<new_platform>/` | Persistence layer implementing `store_content`, `store_comment`, etc. |
| [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) | Contains `CrawlerFactory` registry where new platforms are mapped to CLI flags. |
| [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) | Global configuration shared across all platforms. |

## Summary

Adding a new platform to MediaCrawler is straightforward because the framework abstracts away browser management, proxy handling, and data formatting. To extend support:

- Subclass **AbstractCrawler** in `media_platform/<new_platform>/core.py` and implement `start` and `search` methods.
- Create a **client wrapper** in [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) for API communication.
- Implement **store interfaces** in `store/<new_platform>/` for data persistence.
- Register the crawler in **CrawlerFactory** within [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) using a unique CLI key.
- Add platform-specific configuration in `config/<new_platform>_config.py`.

The shared infrastructure handles CDP mode, IP proxy rotation, async throttling, and output generation, meaning you typically only need to write a few hundred lines of platform-specific code to add full support for a new media source.

## Frequently Asked Questions

### Do I need to modify existing core files to add a new platform?

No. MediaCrawler's plugin architecture allows you to add support for a new platform without touching the base infrastructure. You only create new files within `media_platform/<new_platform>/` and `store/<new_platform>/`, then import and register the class in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py). The existing code in `base/` and `tools/` remains unchanged.

### What is the minimum code required to implement a new crawler?

At minimum, you must implement the `AbstractCrawler` interface with a functional `start` method and a `search` method that handles the crawling logic. You also need a minimal client class with a `pong` method for health checks. While login and detailed storage implementations are recommended for production use, you can stub these initially to test the integration pipeline.

### How does the configuration system work for platform-specific settings?

Global configuration options like `ENABLE_CDP_MODE` and `CRAWLER_MAX_NOTES_COUNT` are read from [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and are automatically available to all platforms. For unique requirements such as platform-specific API endpoints or authentication tokens, create a new file `config/<new_platform>_config.py` and import these values in your crawler implementation.

### Is login support mandatory for new platforms?

Login support is only required if the platform restricts content access to authenticated users. If the platform allows public access to the data you intend to scrape, you can omit the [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) implementation and skip the authentication check in your `start` method. However, most social media platforms require some form of session management, so implementing `AbstractLogin` is typically necessary for full functionality.