# How to Add a New Media Source to MediaCrawler: Complete Implementation Guide

> Learn how to add a new media source to MediaCrawler with this complete implementation guide. Follow our step-by-step instructions to integrate custom crawlers.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**To add a new media source to MediaCrawler, create a new package under `media_platform/` containing [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py), and [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) that implement the abstract contracts defined in `base/`, then register your crawler class in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) around line 39.**

MediaCrawler, maintained by NanmiCoder, is a modular, open-source content aggregation framework built on a strict plugin architecture. When you **add a new media source to MediaCrawler**, you are extending the `media_platform/` package with concrete implementations of three core concepts—orchestration, transport, and transformation—without modifying the underlying engine.

## Understanding the Plugin Architecture

Every supported platform in MediaCrawler adheres to a uniform contract defined by abstract base classes in the `base/` directory.

**`AbstractCrawler`** in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) defines the public API surface that all crawlers must implement, including methods like `run()`, `search()`, and `fetch_detail()`. **`AbstractApiClient`** in [`base/base_client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_client.py) (often used via mix-ins such as `ProxyRefreshMixin`) provides common utilities for HTTP handling, proxy rotation, and retry logic.

The framework expects each media source to live as an isolated sub-package under `media_platform/<new_source>/`, containing four distinct responsibilities:

- **Crawler** ([`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py)): Orchestrates the extraction workflow, pagination, and storage calls.
- **Client** ([`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py)): Low-level wrapper for HTTP, GraphQL, or Playwright interactions.
- **Extractor / Helper** ([`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py)): Parses raw responses into Pydantic models defined in `model/`.
- **Fields** ([`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py)): Enumerations for platform-specific parameters like search type and sort order.

This separation ensures new sources automatically inherit caching, proxy management, and storage layers.

## Step-by-Step Implementation Guide

### Step 1: Create the Platform Package Structure

Initialize the directory structure for your new source:

```bash
mkdir -p media_platform/<new_source>
touch media_platform/<new_source>/__init__.py

```

Replace `<new_source>` with a lowercase identifier for the platform (e.g., `zhihu`, `weibo`).

### Step 2: Define Search Fields and Enums

Create `media_platform/<new_source>/field.py` to store platform-specific constants. Copy the pattern from [`media_platform/zhihu/field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/field.py):

```python
from enum import Enum

class SearchType(Enum):
    KEYWORD = "keyword"
    USER = "user"
    # Add platform-specific types

class SearchSort(Enum):
    RELEVANCE = "relevance"
    NEWEST = "newest"
    # Add sort options

```

### Step 3: Implement the API Client

Create `media_platform/<new_source>/client.py` to handle remote communication. Inherit from `AbstractApiClient` and include `ProxyRefreshMixin` for automatic proxy handling, following the pattern in [`media_platform/weibo/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/client.py):

```python
from base.base_client import AbstractApiClient
from proxy.proxy_mixin import ProxyRefreshMixin

class NewSourceClient(AbstractApiClient, ProxyRefreshMixin):
    """Thin wrapper around the platform's HTTP/GraphQL endpoints."""

    async def fetch_page(self, url: str, **kwargs):
        # Implement concrete request logic (httpx, aiohttp, or Playwright)

        ...

    async def get_user_posts(self, uid: str, page: int = 1):
        # High-level helper for the Crawler to consume

        ...

```

### Step 4: Build the Data Extractor

Create `media_platform/<new_source>/help.py` to transform raw JSON or HTML into typed Pydantic models. Reference [`media_platform/tieba/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/help.py) for the implementation style:

```python
from model.m_xiaohongshu import XiaohongshuPost  # Select appropriate model

class NewSourceExtractor:
    @staticmethod
    def parse_post(raw: dict) -> XiaohongshuPost:
        return XiaohongshuPost(
            uid=raw["user"]["id"],
            content=raw["text"],
            create_time=raw["timestamp"],
            # Map remaining fields

        )

```

### Step 5: Assemble the Concrete Crawler

Create `media_platform/<new_source>/core.py` to integrate the components. Extend `AbstractCrawler` and implement the required async methods, using [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) as a reference:

```python
from base.base_crawler import AbstractCrawler
from .client import NewSourceClient
from .help import NewSourceExtractor
from .field import SearchType, SearchSort

class NewSourceCrawler(AbstractCrawler):
    """Public entry point used by the CLI and scheduler."""

    def __init__(self, config):
        super().__init__(config)
        self.client = NewSourceClient(config)

    async def search(self, keyword: str, page: int = 1, sort: SearchSort = SearchSort.RELEVANCE):
        raw = await self.client.search(keyword, page, sort)
        return [NewSourceExtractor.parse_post(item) for item in raw["items"]]

    async def fetch_detail(self, uid: str):
        raw = await self.client.get_detail(uid)
        return NewSourceExtractor.parse_post(raw)

```

### Step 6: Expose the Crawler in Package Exports

Update `media_platform/<new_source>/__init__.py` to expose the class:

```python
from .core import NewSourceCrawler

__all__ = ["NewSourceCrawler"]

```

### Step 7: Register in the Application Entry Point

Edit [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) around line 39 to import and register the new crawler:

```python
from media_platform.<new_source> import NewSourceCrawler

crawlers = {
    "<new_source>": NewSourceCrawler,
    # Existing entries...

}

```

If the CLI uses argument parsing, add a corresponding flag in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py), following the existing pattern (e.g., the TieBa flag implementation).

### Step 8: Add Platform Configuration (Optional)

Create `config/<new_source>_config.py` inheriting from `BaseConfig` to store platform-specific settings like API keys, rate limits, and pagination size:

```python
from config.base_config import BaseConfig

class NewSourceConfig(BaseConfig):
    API_KEY: str = ""
    RATE_LIMIT: int = 10
    PAGE_SIZE: int = 20

```

Import this configuration where the crawler is instantiated.

### Step 9: Write Unit Tests

Create `tests/test_<new_source>_crawler.py` following the pattern in [`tests/test_tieba_extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_tieba_extractor.py). Mock the client responses and verify that the extractor returns correctly typed models:

```python
import pytest
from media_platform.<new_source>.help import NewSourceExtractor

def test_parse_post():
    raw_data = {"user": {"id": "123"}, "text": "Example", "timestamp": 1699999999}
    post = NewSourceExtractor.parse_post(raw_data)
    assert post.uid == "123"
    assert post.content == "Example"

```

### Step 10: Update Documentation

Add the new platform to [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md) and `docs/项目代码结构.md` to document the supported sources and provide usage examples for future contributors.

## Minimal Working Example

Below is a consolidated implementation for a hypothetical platform named "Example":

**[`media_platform/example/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/example/core.py)**

```python
from base.base_crawler import AbstractCrawler
from .client import ExampleClient
from .help import ExampleExtractor
from .field import SearchType, SearchSort

class ExampleCrawler(AbstractCrawler):
    def __init__(self, cfg):
        super().__init__(cfg)
        self.client = ExampleClient(cfg)

    async def search(self, keyword: str, page: int = 1, sort: SearchSort = SearchSort.NEWEST):
        raw = await self.client.search(keyword, page, sort)
        return [ExampleExtractor.parse_item(it) for it in raw["list"]]

    async def fetch_detail(self, uid: str):
        raw = await self.client.detail(uid)
        return ExampleExtractor.parse_item(raw)

```

**[`media_platform/example/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/example/__init__.py)**

```python
from .core import ExampleCrawler

__all__ = ["ExampleCrawler"]

```

**[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) (excerpt)**

```python
from media_platform.example import ExampleCrawler

crawlers = {
    "example": ExampleCrawler,
    # existing platforms...

}

```

## Key Files to Reference

Consult these source files when implementing a new media source:

- **[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)**: Defines `AbstractCrawler` and required method signatures
- **[`base/base_client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_client.py)**: Defines `AbstractApiClient` and proxy handling utilities
- **[`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py)**: Reference implementation of `ZhihuCrawler`
- **[`media_platform/weibo/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/client.py)**: Example of an API client with `ProxyRefreshMixin`
- **[`media_platform/tieba/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/help.py)**: Extractor logic for parsing raw data into Pydantic models
- **[`media_platform/tieba/field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/field.py)**: Enum definitions for search parameters
- **[`config/zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py)**: Platform-specific configuration pattern
- **[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)**: Central crawler registry and import location
- **[`tests/test_tieba_extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_tieba_extractor.py)**: Unit test template for extractors

## Summary

- Create a new package under `media_platform/<new_source>/` with four files: [`field.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/field.py), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py), [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py), and [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py)
- Inherit from `AbstractCrawler` in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) for the main orchestration class
- Inherit from `AbstractApiClient` and mix in `ProxyRefreshMixin` in the client for transport and proxy handling
- Parse raw responses into Pydantic models using an extractor class defined in [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py)
- Register the crawler in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) and optionally add CLI flags in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py)
- Add platform-specific configuration in `config/<new_source>_config.py` and tests in `tests/`

Following this pattern ensures full compatibility with MediaCrawler's caching, proxy rotation, and storage infrastructure.

## Frequently Asked Questions

### Do I need to modify the core framework to add a new media source?

No. MediaCrawler uses a strict plugin architecture. You only need to create files within `media_platform/<new_source>/` and add an import statement in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py). The abstract base classes in `base/` provide the integration points automatically.

### How do I handle authentication for a new platform?

Implement authentication logic within your [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) file. Store sensitive tokens in your platform-specific configuration class (e.g., `config/<new_source>_config.py` inheriting from `BaseConfig`), and initialize the client with these credentials in the crawler's `__init__` method.

### Can I use Playwright instead of HTTP requests for scraping?

Yes. The `AbstractApiClient` design is agnostic to the transport mechanism. Implement your client methods using Playwright, Selenium, or any other browser automation tool, provided you return data structures that your extractor in [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py) can process.

### How do I test my new crawler without hitting the live API?

Create unit tests in `tests/test_<new_source>_crawler.py` that mock your client methods. Follow the pattern in [`tests/test_tieba_extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_tieba_extractor.py) by patching the client's request methods and returning sample JSON data to verify that `parse_post()` returns correctly typed Pydantic models without external network calls.