How to Add Support for a New Platform to NanmiCoder/MediaCrawler: A Step-by-Step Guide

You can add support for a new platform to MediaCrawler by implementing the AbstractCrawler interface in a new package under media_platform/<platform>/, creating a client wrapper and store module, and registering the class in CrawlerFactory within main.py.

MediaCrawler is an open-source social media scraping framework built around a clean, plugin-based architecture. Extending the tool to support additional platforms requires no changes to the core infrastructure, as the shared utilities for browser automation, proxy handling, and data persistence are already abstracted. This guide walks through the concrete implementation steps based on the actual source code structure of the NanmiCoder/MediaCrawler repository.

Architecture Overview

The foundation of MediaCrawler's extensibility lies in the AbstractCrawler base class defined in base/base_crawler.py. This interface declares the common async methods that every platform crawler must implement, including start, search, and launch_browser.

Platform-specific implementations live in isolated packages under media_platform/<platform>/. The CrawlerFactory in main.py maintains a simple dictionary mapping CLI platform flags to concrete crawler classes, enabling dynamic instantiation at runtime. Because heavy-lifting components like Playwright browser management, CDP fallback handling, async concurrency limits, and Excel/DB output flushing are shared across all crawlers, you only need to write platform-specific API interaction logic.

Step-by-Step Implementation Guide

Adding a new platform requires creating four primary components and updating two registry files. Follow these steps to integrate a new media source.

1. Create the Platform Package Structure

Create a new directory media_platform/<new_platform>/ containing core.py, client.py, and optionally login.py. This mirrors the structure used by existing implementations like Zhihu (media_platform/zhihu/).

2. Implement the Core Crawler Class

In media_platform/<new_platform>/core.py, subclass AbstractCrawler and implement the required async methods. The start method serves as the entry point, handling browser initialization, client creation, authentication, and dispatching to crawling modes.


# media_platform/<new_platform>/core.py

from base.base_crawler import AbstractCrawler
from tools import utils
from var import crawler_type_var, source_keyword_var
from .client import <NewPlatform>Client
from .login import <NewPlatform>Login
from store import <new_platform> as store

class <NewPlatform>Crawler(AbstractCrawler):
    async def start(self) -> None:
        # Launch browser (standard or CDP mode)

        async with async_playwright() as playwright:
            self.browser_context = await self.launch_browser_with_cdp(
                playwright, None, self.user_agent, headless=config.CDP_HEADLESS
            ) if config.ENABLE_CDP_MODE else await self.launch_browser(
                playwright.chromium, None, self.user_agent, headless=config.HEADLESS
            )
            self.context_page = await self.browser_context.new_page()
            await self.context_page.goto(self.index_url)

            # Create API client and handle login

            self.client = await self.create_client()
            if not await self.client.pong():
                login = <NewPlatform>Login(...)
                await login.begin()
                await self.client.update_cookies(self.browser_context, self.cookie_urls)

            # Dispatch to crawling mode

            crawler_type_var.set(config.CRAWLER_TYPE)
            if config.CRAWLER_TYPE == "search":
                await self.search()
            elif config.CRAWLER_TYPE == "detail":
                await self.get_specified_notes()
            elif config.CRAWLER_TYPE == "creator":
                await self.get_creators_and_notes()

Implement the search method to fetch content lists using config.KEYWORDS and config.CRAWLER_MAX_NOTES_COUNT, then delegate storage to the platform-specific store module.

3. Build the HTTP Client Wrapper

Create media_platform/<new_platform>/client.py to handle platform-specific HTTP API calls or Playwright page interactions. The client should expose standard methods like pong, get_note_by_keyword, get_note_detail, and update_cookies.


# media_platform/<new_platform>/client.py

class <NewPlatform>Client:
    def __init__(self, *, proxy: Optional[str], headers: Dict[str, str], playwright_page: Page, cookie_dict: Dict):
        self.proxy = proxy
        self.headers = headers
        self.page = playwright_page
        self.cookies = cookie_dict

    async def pong(self) -> bool:
        # Simple health-check request

        resp = await httpx.get("https://api.<new_platform>.com/ping", headers=self.headers, proxy=self.proxy)
        return resp.status_code == 200

Reference media_platform/zhihu/client.py for the expected interface patterns.

4. Implement the Storage Layer

Create store/<new_platform>/__init__.py and concrete implementations to handle persistence. Implement the abstract store interfaces store_content, store_comment, and store_creator, following the pattern in store/zhihu/_store_impl.py.


# store/<new_platform>/__init__.py

async def update_<new_platform>_content(content: Dict) -> None:
    # Convert platform-specific model to generic dict if needed

    await store_base.save(content, platform="<new_platform>")

5. Add Login Support (If Required)

If the platform requires authentication, create media_platform/<new_platform>/login.py implementing the AbstractLogin contract. This handles QR-code scanning or credential-based authentication before the crawling session begins.

6. Register in CrawlerFactory

Import your new crawler in main.py and add it to the CRAWLERS dictionary within CrawlerFactory. The key represents the value users pass to the --platform CLI flag.


# main.py

from media_platform.<new_platform> import <NewPlatform>Crawler

CRAWLERS = {
    "xhs": XiaoHongShuCrawler,
    "dy": DouYinCrawler,
    # ... existing platforms ...

    "<new_platform_key>": <NewPlatform>Crawler,
}

7. Expose Configuration Options

Create config/<new_platform>_config.py for platform-specific settings like API keys or pagination limits. Global options such as ENABLE_IP_PROXY, ENABLE_CDP_MODE, and CRAWLER_MAX_SLEEP_SEC are automatically available via the shared config module.

8. Write Unit Tests

Add tests under tests/ that import the new crawler and verify public methods execute without error. Mock network calls during testing, following the patterns in test_cdp_browser.py or test_static_proxy_provider.py.

Key Files and Their Roles

File Purpose
base/base_crawler.py Defines the AbstractCrawler interface that every platform must implement.
media_platform/<new_platform>/core.py Platform-specific crawler implementation containing the main crawling logic.
media_platform/<new_platform>/client.py HTTP client wrapper for platform API interactions.
media_platform/<new_platform>/login.py Optional login helper respecting the AbstractLogin contract.
store/<new_platform>/ Persistence layer implementing store_content, store_comment, etc.
main.py Contains CrawlerFactory registry where new platforms are mapped to CLI flags.
config/base_config.py Global configuration shared across all platforms.

Summary

Adding a new platform to MediaCrawler is straightforward because the framework abstracts away browser management, proxy handling, and data formatting. To extend support:

  • Subclass AbstractCrawler in media_platform/<new_platform>/core.py and implement start and search methods.
  • Create a client wrapper in client.py for API communication.
  • Implement store interfaces in store/<new_platform>/ for data persistence.
  • Register the crawler in CrawlerFactory within main.py using a unique CLI key.
  • Add platform-specific configuration in config/<new_platform>_config.py.

The shared infrastructure handles CDP mode, IP proxy rotation, async throttling, and output generation, meaning you typically only need to write a few hundred lines of platform-specific code to add full support for a new media source.

Frequently Asked Questions

Do I need to modify existing core files to add a new platform?

No. MediaCrawler's plugin architecture allows you to add support for a new platform without touching the base infrastructure. You only create new files within media_platform/<new_platform>/ and store/<new_platform>/, then import and register the class in main.py. The existing code in base/ and tools/ remains unchanged.

What is the minimum code required to implement a new crawler?

At minimum, you must implement the AbstractCrawler interface with a functional start method and a search method that handles the crawling logic. You also need a minimal client class with a pong method for health checks. While login and detailed storage implementations are recommended for production use, you can stub these initially to test the integration pipeline.

How does the configuration system work for platform-specific settings?

Global configuration options like ENABLE_CDP_MODE and CRAWLER_MAX_NOTES_COUNT are read from config/base_config.py and are automatically available to all platforms. For unique requirements such as platform-specific API endpoints or authentication tokens, create a new file config/<new_platform>_config.py and import these values in your crawler implementation.

Is login support mandatory for new platforms?

Login support is only required if the platform restricts content access to authenticated users. If the platform allows public access to the data you intend to scrape, you can omit the login.py implementation and skip the authentication check in your start method. However, most social media platforms require some form of session management, so implementing AbstractLogin is typically necessary for full functionality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →