How to Extend MediaCrawler with Custom Crawlers: A Step‑by‑Step Guide

Yes, MediaCrawler supports custom crawler extensions through a minimal three‑method interface and a factory registration pattern.

The NanmiCoder/MediaCrawler project is architected around a clean abstraction layer that lets developers add new platform support without modifying core engine code. By implementing the AbstractCrawler base class and registering your subclass in the CrawlerFactory, you can integrate any media source into the existing data pipeline.

The Core Abstraction: AbstractCrawler

All crawlers in MediaCrawler inherit from AbstractCrawler, defined in [base/base_crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This abstract base class enforces a three‑method contract that every crawler must fulfill:

  • start() – Entry point that orchestrates the entire crawling workflow.
  • search() – Platform‑specific implementation for discovering and fetching content.
  • launch_browser() – Creates a Playwright BrowserContext (or optionally uses Chrome DevTools Protocol mode).

These async methods provide exactly enough structure for the framework to invoke your crawler generically, while leaving all platform‑specific details to your implementation.

How the CrawlerFactory Works

The CrawlerFactory in [main.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) maintains a dictionary named CRAWLERS that maps short platform identifiers to concrete crawler classes:

class CrawlerFactory:
    CRAWLERS: dict[str, Type[AbstractCrawler]] = {
        "xhs": XiaoHongShuCrawler,
        "dy": DouYinCrawler,
        "ks": KuaishouCrawler,
        "bili": BilibiliCrawler,
        "wb": WeiboCrawler,
        "tieba": TieBaCrawler,
        "zhihu": ZhihuCrawler,
    }

When you run MediaCrawler with --platform <identifier>, the factory looks up the key, instantiates the corresponding class, and calls start(). Runtime resolution means you can add new entries without touching any other framework code.

Step‑by‑Step: Building a Custom MediaCrawler Extension

Step 1: Create Your Crawler Subclass

Implement the three required methods in a new file. Below is a minimal, runnable template placed at media_platform/mysite/mysite_crawler.py:


# media_platform/mysite/mysite_crawler.py

from base.base_crawler import AbstractCrawler
from playwright.async_api import BrowserContext, BrowserType, Playwright
from typing import Optional, Dict

class MySiteCrawler(AbstractCrawler):
    """Crawler for the fictional MySite platform."""

    def __init__(self) -> None:
        self.base_url = "https://www.mysite.com"
        self.user_agent = (
            "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
            "AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36"
        )

    async def start(self) -> None:
        async with async_playwright() as pw:
            self.browser_context = await self.launch_browser(
                pw.chromium, None, self.user_agent, headless=True
            )
            page = await self.browser_context.new_page()
            await page.goto(self.base_url)
            await self.search()

    async def search(self) -> None:
        print("[MySiteCrawler] Performing dummy search…")
        # Replace with: pagination logic, API calls, DOM scraping, etc.

    async def launch_browser(
        self,
        chromium: BrowserType,
        playwright_proxy: Optional[Dict],
        user_agent: Optional[str],
        headless: bool = True,
    ) -> BrowserContext:
        browser = await chromium.launch(headless=headless, proxy=playwright_proxy)
        return await browser.new_context(
            viewport={"width": 1920, "height": 1080},
            user_agent=user_agent
        )

Step 2: Register in the Factory

Modify [main.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to import your class and add it to CRAWLERS:


# Add import near existing platform imports

from media_platform.mysite.mysite_crawler import MySiteCrawler

class CrawlerFactory:
    CRAWLERS: dict[str, Type[AbstractCrawler]] = {
        "xhs": XiaoHongShuCrawler,
        "dy": DouYinCrawler,
        "ks": KuaishouCrawler,
        "bili": BilibiliCrawler,
        "wb": WeiboCrawler,
        "tieba": TieBaCrawler,
        "zhihu": ZhihuCrawler,
        "mysite": MySiteCrawler,  # ← your custom crawler

    }

Step 3: Execute with Your Platform Identifier

Run MediaCrawler using either the CLI flag or configuration:


# CLI method

python -m MediaCrawler.main --platform mysite

# Or set in config/base_config.py

PLATFORM = "mysite"

The factory instantiates MySiteCrawler, invokes start(), and downstream features—database storage, Excel export, word‑cloud generation—function unchanged.

Key Files for Custom Crawler Development

File Purpose
base/base_crawler.py Defines AbstractCrawler, AbstractLogin, AbstractStore, and other base interfaces.
media_platform/zhihu/core.py Reference implementation showing a production crawler (Zhihu).
main.py Entry point containing CrawlerFactory and the CRAWLERS registry.
config/base_config.py Runtime configuration; set PLATFORM for config‑driven selection.

Studying existing implementations in media_platform/*/ reveals patterns for handling authentication, rate limiting, and data extraction that you can adapt for your custom crawler.

Summary

  • AbstractCrawler interface – Three async methods (start, search, launch_browser) define the contract.
  • Factory registration – Add a dictionary entry in main.py to wire up your class.
  • Zero core changes – The framework resolves crawlers at runtime; no engine modifications needed.
  • Full pipeline compatibility – Standard features like storage and visualization work immediately.

MediaCrawler's extension model follows classic factory‑pattern design, making it straightforward to integrate any new media platform.

Frequently Asked Questions

What happens if I don't implement all three abstract methods?

Python raises TypeError at instantiation because AbstractCrawler uses abc.ABC enforcement. You must provide concrete implementations of start(), search(), and launch_browser() or the class cannot be instantiated.

Can I reuse existing login or storage implementations?

Yes. The same base/base_crawler.py file declares AbstractLogin and AbstractStore interfaces. Reference how media_platform/zhihu/core.py composes these to add persistent sessions and database output to your custom crawler without reinventing infrastructure.

Is headless browser mode mandatory for custom crawlers?

No. The launch_browser() signature accepts a headless parameter defaulting to True. You can override this in your subclass or pass headless=False for debugging. CDP (Chrome DevTools Protocol) mode is also supported by returning an appropriate BrowserContext configuration.

Do I need to modify the database schema for new platforms?

Typically no. MediaCrawler's AbstractStore implementations use generic field mappings. As long as your crawler yields data structures compatible with existing store methods, records flow into the same tables. Platform‑specific fields can be stored in JSON columns or extended through subclassing AbstractStore if needed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →