How to Add a New Media Source to MediaCrawler: Complete Implementation Guide

To add a new media source to MediaCrawler, create a new package under media_platform/ containing field.py, client.py, help.py, and core.py that implement the abstract contracts defined in base/, then register your crawler class in main.py around line 39.

MediaCrawler, maintained by NanmiCoder, is a modular, open-source content aggregation framework built on a strict plugin architecture. When you add a new media source to MediaCrawler, you are extending the media_platform/ package with concrete implementations of three core concepts—orchestration, transport, and transformation—without modifying the underlying engine.

Understanding the Plugin Architecture

Every supported platform in MediaCrawler adheres to a uniform contract defined by abstract base classes in the base/ directory.

AbstractCrawler in base/base_crawler.py defines the public API surface that all crawlers must implement, including methods like run(), search(), and fetch_detail(). AbstractApiClient in base/base_client.py (often used via mix-ins such as ProxyRefreshMixin) provides common utilities for HTTP handling, proxy rotation, and retry logic.

The framework expects each media source to live as an isolated sub-package under media_platform/<new_source>/, containing four distinct responsibilities:

  • Crawler (core.py): Orchestrates the extraction workflow, pagination, and storage calls.
  • Client (client.py): Low-level wrapper for HTTP, GraphQL, or Playwright interactions.
  • Extractor / Helper (help.py): Parses raw responses into Pydantic models defined in model/.
  • Fields (field.py): Enumerations for platform-specific parameters like search type and sort order.

This separation ensures new sources automatically inherit caching, proxy management, and storage layers.

Step-by-Step Implementation Guide

Step 1: Create the Platform Package Structure

Initialize the directory structure for your new source:

mkdir -p media_platform/<new_source>
touch media_platform/<new_source>/__init__.py

Replace <new_source> with a lowercase identifier for the platform (e.g., zhihu, weibo).

Step 2: Define Search Fields and Enums

Create media_platform/<new_source>/field.py to store platform-specific constants. Copy the pattern from media_platform/zhihu/field.py:

from enum import Enum

class SearchType(Enum):
    KEYWORD = "keyword"
    USER = "user"
    # Add platform-specific types

class SearchSort(Enum):
    RELEVANCE = "relevance"
    NEWEST = "newest"
    # Add sort options

Step 3: Implement the API Client

Create media_platform/<new_source>/client.py to handle remote communication. Inherit from AbstractApiClient and include ProxyRefreshMixin for automatic proxy handling, following the pattern in media_platform/weibo/client.py:

from base.base_client import AbstractApiClient
from proxy.proxy_mixin import ProxyRefreshMixin

class NewSourceClient(AbstractApiClient, ProxyRefreshMixin):
    """Thin wrapper around the platform's HTTP/GraphQL endpoints."""

    async def fetch_page(self, url: str, **kwargs):
        # Implement concrete request logic (httpx, aiohttp, or Playwright)

        ...

    async def get_user_posts(self, uid: str, page: int = 1):
        # High-level helper for the Crawler to consume

        ...

Step 4: Build the Data Extractor

Create media_platform/<new_source>/help.py to transform raw JSON or HTML into typed Pydantic models. Reference media_platform/tieba/help.py for the implementation style:

from model.m_xiaohongshu import XiaohongshuPost  # Select appropriate model

class NewSourceExtractor:
    @staticmethod
    def parse_post(raw: dict) -> XiaohongshuPost:
        return XiaohongshuPost(
            uid=raw["user"]["id"],
            content=raw["text"],
            create_time=raw["timestamp"],
            # Map remaining fields

        )

Step 5: Assemble the Concrete Crawler

Create media_platform/<new_source>/core.py to integrate the components. Extend AbstractCrawler and implement the required async methods, using media_platform/zhihu/core.py as a reference:

from base.base_crawler import AbstractCrawler
from .client import NewSourceClient
from .help import NewSourceExtractor
from .field import SearchType, SearchSort

class NewSourceCrawler(AbstractCrawler):
    """Public entry point used by the CLI and scheduler."""

    def __init__(self, config):
        super().__init__(config)
        self.client = NewSourceClient(config)

    async def search(self, keyword: str, page: int = 1, sort: SearchSort = SearchSort.RELEVANCE):
        raw = await self.client.search(keyword, page, sort)
        return [NewSourceExtractor.parse_post(item) for item in raw["items"]]

    async def fetch_detail(self, uid: str):
        raw = await self.client.get_detail(uid)
        return NewSourceExtractor.parse_post(raw)

Step 6: Expose the Crawler in Package Exports

Update media_platform/<new_source>/__init__.py to expose the class:

from .core import NewSourceCrawler

__all__ = ["NewSourceCrawler"]

Step 7: Register in the Application Entry Point

Edit main.py around line 39 to import and register the new crawler:

from media_platform.<new_source> import NewSourceCrawler

crawlers = {
    "<new_source>": NewSourceCrawler,
    # Existing entries...

}

If the CLI uses argument parsing, add a corresponding flag in cmd_arg/arg.py, following the existing pattern (e.g., the TieBa flag implementation).

Step 8: Add Platform Configuration (Optional)

Create config/<new_source>_config.py inheriting from BaseConfig to store platform-specific settings like API keys, rate limits, and pagination size:

from config.base_config import BaseConfig

class NewSourceConfig(BaseConfig):
    API_KEY: str = ""
    RATE_LIMIT: int = 10
    PAGE_SIZE: int = 20

Import this configuration where the crawler is instantiated.

Step 9: Write Unit Tests

Create tests/test_<new_source>_crawler.py following the pattern in tests/test_tieba_extractor.py. Mock the client responses and verify that the extractor returns correctly typed models:

import pytest
from media_platform.<new_source>.help import NewSourceExtractor

def test_parse_post():
    raw_data = {"user": {"id": "123"}, "text": "Example", "timestamp": 1699999999}
    post = NewSourceExtractor.parse_post(raw_data)
    assert post.uid == "123"
    assert post.content == "Example"

Step 10: Update Documentation

Add the new platform to README.md and docs/项目代码结构.md to document the supported sources and provide usage examples for future contributors.

Minimal Working Example

Below is a consolidated implementation for a hypothetical platform named "Example":

media_platform/example/core.py

from base.base_crawler import AbstractCrawler
from .client import ExampleClient
from .help import ExampleExtractor
from .field import SearchType, SearchSort

class ExampleCrawler(AbstractCrawler):
    def __init__(self, cfg):
        super().__init__(cfg)
        self.client = ExampleClient(cfg)

    async def search(self, keyword: str, page: int = 1, sort: SearchSort = SearchSort.NEWEST):
        raw = await self.client.search(keyword, page, sort)
        return [ExampleExtractor.parse_item(it) for it in raw["list"]]

    async def fetch_detail(self, uid: str):
        raw = await self.client.detail(uid)
        return ExampleExtractor.parse_item(raw)

media_platform/example/__init__.py

from .core import ExampleCrawler

__all__ = ["ExampleCrawler"]

main.py (excerpt)

from media_platform.example import ExampleCrawler

crawlers = {
    "example": ExampleCrawler,
    # existing platforms...

}

Key Files to Reference

Consult these source files when implementing a new media source:

Summary

  • Create a new package under media_platform/<new_source>/ with four files: field.py, client.py, help.py, and core.py
  • Inherit from AbstractCrawler in base/base_crawler.py for the main orchestration class
  • Inherit from AbstractApiClient and mix in ProxyRefreshMixin in the client for transport and proxy handling
  • Parse raw responses into Pydantic models using an extractor class defined in help.py
  • Register the crawler in main.py and optionally add CLI flags in cmd_arg/arg.py
  • Add platform-specific configuration in config/<new_source>_config.py and tests in tests/

Following this pattern ensures full compatibility with MediaCrawler's caching, proxy rotation, and storage infrastructure.

Frequently Asked Questions

Do I need to modify the core framework to add a new media source?

No. MediaCrawler uses a strict plugin architecture. You only need to create files within media_platform/<new_source>/ and add an import statement in main.py. The abstract base classes in base/ provide the integration points automatically.

How do I handle authentication for a new platform?

Implement authentication logic within your client.py file. Store sensitive tokens in your platform-specific configuration class (e.g., config/<new_source>_config.py inheriting from BaseConfig), and initialize the client with these credentials in the crawler's __init__ method.

Can I use Playwright instead of HTTP requests for scraping?

Yes. The AbstractApiClient design is agnostic to the transport mechanism. Implement your client methods using Playwright, Selenium, or any other browser automation tool, provided you return data structures that your extractor in help.py can process.

How do I test my new crawler without hitting the live API?

Create unit tests in tests/test_<new_source>_crawler.py that mock your client methods. Follow the pattern in tests/test_tieba_extractor.py by patching the client's request methods and returning sample JSON data to verify that parse_post() returns correctly typed Pydantic models without external network calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →