How the CrawlerFactory Pattern Creates Platform-Specific Crawler Instances in MediaCrawler

MediaCrawler uses a classic Factory pattern via CrawlerFactory class to decouple crawler instantiation from business logic, mapping platform strings like "xhs" or "dy" to concrete crawler classes through a static registry dictionary.

The CrawlerFactory pattern in NanmiCoder/MediaCrawler provides a clean mechanism for creating platform-specific crawler instances without hard-coding class references throughout the application. This design enables seamless switching between social media platforms through configuration changes alone.

Abstract Interface Design

All platform crawlers inherit from AbstractCrawler, defined in [base/base_crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This abstract base class establishes the contract that every concrete crawler must fulfill:

  • async def start(self) – entry point for crawler execution
  • async def search(self) – handles search-based data collection
  • async def launch_browser(self) – manages browser automation setup

By programming against this interface rather than concrete implementations, the application achieves interface-based polymorphism that the factory exploits.

Concrete Crawler Implementations

Each supported platform provides its own subclass of AbstractCrawler within the media_platform package:

Platform File Class Name
XiaoHongShu (小红书) [media_platform/xhs/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) XiaoHongShuCrawler
DouYin (抖音) [media_platform/douyin/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) DouYinCrawler
Zhihu (知乎) [media_platform/zhihu/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) ZhihuCrawler
Kuaishou (快手) media_platform/ks/core.py KuaishouCrawler
Bilibili media_platform/bilibili/core.py BilibiliCrawler
Weibo media_platform/weibo/core.py WeiboCrawler
TieBa media_platform/tieba/core.py TieBaCrawler

Each class independently implements the three abstract methods according to platform-specific requirements.

Factory Registry Implementation

The factory itself resides in [main.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) as the CrawlerFactory class. It maintains a static dictionary mapping platform identifiers to crawler classes:

class CrawlerFactory:
    CRAWLERS: dict[str, Type[AbstractCrawler]] = {
        "xhs": XiaoHongShuCrawler,
        "dy": DouYinCrawler,
        "ks": KuaishouCrawler,
        "bili": BilibiliCrawler,
        "wb": WeiboCrawler,
        "tieba": TieBaCrawler,
        "zhihu": ZhihuCrawler,
    }

This registry enables runtime crawler selection based on string configuration rather than compile-time class references.

Factory Method: create_crawler()

The factory exposes a single static method for crawler instantiation:

@staticmethod
def create_crawler(platform: str) -> AbstractCrawler:
    crawler_class = CrawlerFactory.CRAWLERS.get(platform)
    if not crawler_class:
        supported = ", ".join(sorted(CrawlerFactory.CRAWLERS))
        raise ValueError(
            f"Invalid media platform: {platform!r}. Supported: {supported}"
        )
    return crawler_class()

Key characteristics of this implementation:

  • Defensive validation – raises ValueError with enumerated supported platforms for invalid keys
  • Instance creation – returns a new instance, not the class itself
  • Return type – AbstractCrawler ensures callers depend on the interface, not concrete types

Runtime Usage in the Application

The entry point demonstrates clean factory consumption:


# From main.py

crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)
await crawler.start()

The config.PLATFORM value (set in [config/base_config.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)) determines which crawler executes. Callers remain oblivious to the concrete class—only the interface matters.

Practical Code Examples

Direct Factory Usage

from main import CrawlerFactory
from config import PLATFORM  # e.g., "xhs"

crawler = CrawlerFactory.create_crawler(PLATFORM)
await crawler.start()  # Executes XiaoHongShuCrawler.start()

Extending with a Custom Platform

Adding new platforms requires no modifications to existing code:


# Step 1: Implement AbstractCrawler subclass

# media_platform/custom/core.py

from base.base_crawler import AbstractCrawler

class CustomCrawler(AbstractCrawler):
    async def start(self): ...
    async def search(self): ...
    async def launch_browser(self): ...

# Step 2: Register in factory

from media_platform.custom.core import CustomCrawler
CrawlerFactory.CRAWLERS["custom"] = CustomCrawler

# Step 3: Execute

crawler = CrawlerFactory.create_crawler("custom")
await crawler.start()

Extensibility Architecture

The factory pattern delivers three architectural benefits:

  1. Open/Closed Principle – new platforms extend without modifying factory logic or caller code
  2. Single Responsibility – creation logic isolated from execution logic
  3. Testability – interface-based design enables mock crawlers for unit testing

Adding a platform follows a three-step process: implement the abstract class, import it in main.py, and register in CRAWLERS. No other changes propagate through the codebase.

Summary

  • Abstract base: AbstractCrawler in base/base_crawler.py defines the crawler contract
  • Factory location: CrawlerFactory class in main.py with static CRAWLERS registry
  • Registry keys: Platform strings ("xhs", "dy", "zhihu", etc.) map to concrete classes
  • Creation method: create_crawler(platform) validates input and returns new instances
  • Runtime binding: config.PLATFORM selects the crawler without hard-coded references
  • Extension mechanism: new platforms register by adding dictionary entries

Frequently Asked Questions

What happens if I pass an unsupported platform string to create_crawler()?

create_crawler() raises a ValueError with a descriptive message listing all supported platforms: f"Invalid media platform: {platform!r}. Supported: {supported}". The supported string is dynamically generated from CRAWLERS.keys().

Can I instantiate crawlers directly without the factory?

Yes, but this bypasses the abstraction layer and couples your code to concrete implementations. Direct instantiation (XiaoHongShuCrawler()) works for testing but defeats the factory's decoupling benefits in production code.

How do I add a new platform without modifying main.py?

Currently, registration requires modifying main.py to import the new class and update CRAWLERS. For fully dynamic registration, you could implement a decorator-based or entry-point-based discovery system, though this would extend beyond the existing factory pattern.

Is CrawlerFactory thread-safe for concurrent crawler creation?

The factory's create_crawler() method is stateless and read-only except for the static dictionary lookup. Dictionary reads in Python are thread-safe, making concurrent calls safe. However, the returned crawler instance's start() method must handle its own concurrency constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →