# How to Extend MediaCrawler with Custom Crawlers: A Step‑by‑Step Guide

> Extend MediaCrawler with custom crawlers using its simple three-method interface and factory registration pattern. Build your own crawlers easily with this step-by-step guide for NanmiCoder/MediaCrawler.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-12

---

**Yes, MediaCrawler supports custom crawler extensions through a minimal three‑method interface and a factory registration pattern.**

The [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) project is architected around a clean abstraction layer that lets developers add new platform support without modifying core engine code. By implementing the `AbstractCrawler` base class and registering your subclass in the `CrawlerFactory`, you can integrate any media source into the existing data pipeline.

## The Core Abstraction: `AbstractCrawler`

All crawlers in MediaCrawler inherit from `AbstractCrawler`, defined in [[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This abstract base class enforces a three‑method contract that every crawler must fulfill:

- `start()` – Entry point that orchestrates the entire crawling workflow.
- `search()` – Platform‑specific implementation for discovering and fetching content.
- `launch_browser()` – Creates a Playwright `BrowserContext` (or optionally uses Chrome DevTools Protocol mode).

These async methods provide exactly enough structure for the framework to invoke your crawler generically, while leaving all platform‑specific details to your implementation.

## How the CrawlerFactory Works

The `CrawlerFactory` in [[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) maintains a dictionary named `CRAWLERS` that maps short platform identifiers to concrete crawler classes:

```python
class CrawlerFactory:
    CRAWLERS: dict[str, Type[AbstractCrawler]] = {
        "xhs": XiaoHongShuCrawler,
        "dy": DouYinCrawler,
        "ks": KuaishouCrawler,
        "bili": BilibiliCrawler,
        "wb": WeiboCrawler,
        "tieba": TieBaCrawler,
        "zhihu": ZhihuCrawler,
    }

```

When you run MediaCrawler with `--platform <identifier>`, the factory looks up the key, instantiates the corresponding class, and calls `start()`. Runtime resolution means you can add new entries without touching any other framework code.

## Step‑by‑Step: Building a Custom MediaCrawler Extension

### Step 1: Create Your Crawler Subclass

Implement the three required methods in a new file. Below is a minimal, runnable template placed at [`media_platform/mysite/mysite_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/mysite/mysite_crawler.py):

```python

# media_platform/mysite/mysite_crawler.py

from base.base_crawler import AbstractCrawler
from playwright.async_api import BrowserContext, BrowserType, Playwright
from typing import Optional, Dict

class MySiteCrawler(AbstractCrawler):
    """Crawler for the fictional MySite platform."""

    def __init__(self) -> None:
        self.base_url = "https://www.mysite.com"
        self.user_agent = (
            "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
            "AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36"
        )

    async def start(self) -> None:
        async with async_playwright() as pw:
            self.browser_context = await self.launch_browser(
                pw.chromium, None, self.user_agent, headless=True
            )
            page = await self.browser_context.new_page()
            await page.goto(self.base_url)
            await self.search()

    async def search(self) -> None:
        print("[MySiteCrawler] Performing dummy search…")
        # Replace with: pagination logic, API calls, DOM scraping, etc.

    async def launch_browser(
        self,
        chromium: BrowserType,
        playwright_proxy: Optional[Dict],
        user_agent: Optional[str],
        headless: bool = True,
    ) -> BrowserContext:
        browser = await chromium.launch(headless=headless, proxy=playwright_proxy)
        return await browser.new_context(
            viewport={"width": 1920, "height": 1080},
            user_agent=user_agent
        )

```

### Step 2: Register in the Factory

Modify [[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to import your class and add it to `CRAWLERS`:

```python

# Add import near existing platform imports

from media_platform.mysite.mysite_crawler import MySiteCrawler

class CrawlerFactory:
    CRAWLERS: dict[str, Type[AbstractCrawler]] = {
        "xhs": XiaoHongShuCrawler,
        "dy": DouYinCrawler,
        "ks": KuaishouCrawler,
        "bili": BilibiliCrawler,
        "wb": WeiboCrawler,
        "tieba": TieBaCrawler,
        "zhihu": ZhihuCrawler,
        "mysite": MySiteCrawler,  # ← your custom crawler

    }

```

### Step 3: Execute with Your Platform Identifier

Run MediaCrawler using either the CLI flag or configuration:

```bash

# CLI method

python -m MediaCrawler.main --platform mysite

# Or set in config/base_config.py

PLATFORM = "mysite"

```

The factory instantiates `MySiteCrawler`, invokes `start()`, and downstream features—database storage, Excel export, word‑cloud generation—function unchanged.

## Key Files for Custom Crawler Development

| File | Purpose |
|------|---------|
| [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) | Defines `AbstractCrawler`, `AbstractLogin`, `AbstractStore`, and other base interfaces. |
| [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) | Reference implementation showing a production crawler (Zhihu). |
| [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) | Entry point containing `CrawlerFactory` and the `CRAWLERS` registry. |
| [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) | Runtime configuration; set `PLATFORM` for config‑driven selection. |

Studying existing implementations in `media_platform/*/` reveals patterns for handling authentication, rate limiting, and data extraction that you can adapt for your custom crawler.

## Summary

- **AbstractCrawler interface** – Three async methods (`start`, `search`, `launch_browser`) define the contract.
- **Factory registration** – Add a dictionary entry in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to wire up your class.
- **Zero core changes** – The framework resolves crawlers at runtime; no engine modifications needed.
- **Full pipeline compatibility** – Standard features like storage and visualization work immediately.

MediaCrawler's extension model follows classic factory‑pattern design, making it straightforward to integrate any new media platform.

## Frequently Asked Questions

### What happens if I don't implement all three abstract methods?

Python raises `TypeError` at instantiation because `AbstractCrawler` uses `abc.ABC` enforcement. You must provide concrete implementations of `start()`, `search()`, and `launch_browser()` or the class cannot be instantiated.

### Can I reuse existing login or storage implementations?

Yes. The same [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) file declares `AbstractLogin` and `AbstractStore` interfaces. Reference how [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) composes these to add persistent sessions and database output to your custom crawler without reinventing infrastructure.

### Is headless browser mode mandatory for custom crawlers?

No. The `launch_browser()` signature accepts a `headless` parameter defaulting to `True`. You can override this in your subclass or pass `headless=False` for debugging. CDP (Chrome DevTools Protocol) mode is also supported by returning an appropriate `BrowserContext` configuration.

### Do I need to modify the database schema for new platforms?

Typically no. MediaCrawler's `AbstractStore` implementations use generic field mappings. As long as your crawler yields data structures compatible with existing store methods, records flow into the same tables. Platform‑specific fields can be stored in JSON columns or extended through subclassing `AbstractStore` if needed.