# How the CrawlerFactory Pattern Creates Platform-Specific Crawler Instances in MediaCrawler

> Learn how the CrawlerFactory pattern in MediaCrawler efficiently creates platform-specific crawler instances using a static registry to decouple instantiation from business logic.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: internals
- Published: 2026-08-14

---

**MediaCrawler uses a classic Factory pattern via `CrawlerFactory` class to decouple crawler instantiation from business logic, mapping platform strings like "xhs" or "dy" to concrete crawler classes through a static registry dictionary.**

The `CrawlerFactory` pattern in [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) provides a clean mechanism for creating platform-specific crawler instances without hard-coding class references throughout the application. This design enables seamless switching between social media platforms through configuration changes alone.

## Abstract Interface Design

All platform crawlers inherit from `AbstractCrawler`, defined in [[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This abstract base class establishes the contract that every concrete crawler must fulfill:

- `async def start(self)` – entry point for crawler execution
- `async def search(self)` – handles search-based data collection
- `async def launch_browser(self)` – manages browser automation setup

By programming against this interface rather than concrete implementations, the application achieves **interface-based polymorphism** that the factory exploits.

## Concrete Crawler Implementations

Each supported platform provides its own subclass of `AbstractCrawler` within the `media_platform` package:

| Platform | File | Class Name |
|----------|------|------------|
| XiaoHongShu (小红书) | [[`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) | `XiaoHongShuCrawler` |
| DouYin (抖音) | [[`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) | `DouYinCrawler` |
| Zhihu (知乎) | [[`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) | `ZhihuCrawler` |
| Kuaishou (快手) | [`media_platform/ks/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/ks/core.py) | `KuaishouCrawler` |
| Bilibili | [`media_platform/bilibili/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/bilibili/core.py) | `BilibiliCrawler` |
| Weibo | [`media_platform/weibo/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/weibo/core.py) | `WeiboCrawler` |
| TieBa | [`media_platform/tieba/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/tieba/core.py) | `TieBaCrawler` |

Each class independently implements the three abstract methods according to platform-specific requirements.

## Factory Registry Implementation

The factory itself resides in [[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) as the `CrawlerFactory` class. It maintains a static dictionary mapping platform identifiers to crawler classes:

```python
class CrawlerFactory:
    CRAWLERS: dict[str, Type[AbstractCrawler]] = {
        "xhs": XiaoHongShuCrawler,
        "dy": DouYinCrawler,
        "ks": KuaishouCrawler,
        "bili": BilibiliCrawler,
        "wb": WeiboCrawler,
        "tieba": TieBaCrawler,
        "zhihu": ZhihuCrawler,
    }

```

This registry enables **runtime crawler selection** based on string configuration rather than compile-time class references.

## Factory Method: `create_crawler()`

The factory exposes a single static method for crawler instantiation:

```python
@staticmethod
def create_crawler(platform: str) -> AbstractCrawler:
    crawler_class = CrawlerFactory.CRAWLERS.get(platform)
    if not crawler_class:
        supported = ", ".join(sorted(CrawlerFactory.CRAWLERS))
        raise ValueError(
            f"Invalid media platform: {platform!r}. Supported: {supported}"
        )
    return crawler_class()

```

Key characteristics of this implementation:

- **Defensive validation** – raises `ValueError` with enumerated supported platforms for invalid keys
- **Instance creation** – returns a new instance, not the class itself
- **Return type** – `AbstractCrawler` ensures callers depend on the interface, not concrete types

## Runtime Usage in the Application

The entry point demonstrates clean factory consumption:

```python

# From main.py

crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)
await crawler.start()

```

The `config.PLATFORM` value (set in [[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)) determines which crawler executes. Callers remain oblivious to the concrete class—only the interface matters.

## Practical Code Examples

### Direct Factory Usage

```python
from main import CrawlerFactory
from config import PLATFORM  # e.g., "xhs"

crawler = CrawlerFactory.create_crawler(PLATFORM)
await crawler.start()  # Executes XiaoHongShuCrawler.start()

```

### Extending with a Custom Platform

Adding new platforms requires no modifications to existing code:

```python

# Step 1: Implement AbstractCrawler subclass

# media_platform/custom/core.py

from base.base_crawler import AbstractCrawler

class CustomCrawler(AbstractCrawler):
    async def start(self): ...
    async def search(self): ...
    async def launch_browser(self): ...

# Step 2: Register in factory

from media_platform.custom.core import CustomCrawler
CrawlerFactory.CRAWLERS["custom"] = CustomCrawler

# Step 3: Execute

crawler = CrawlerFactory.create_crawler("custom")
await crawler.start()

```

## Extensibility Architecture

The factory pattern delivers three architectural benefits:

1. **Open/Closed Principle** – new platforms extend without modifying factory logic or caller code
2. **Single Responsibility** – creation logic isolated from execution logic
3. **Testability** – interface-based design enables mock crawlers for unit testing

Adding a platform follows a three-step process: implement the abstract class, import it in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py), and register in `CRAWLERS`. No other changes propagate through the codebase.

## Summary

- **Abstract base**: `AbstractCrawler` in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) defines the crawler contract
- **Factory location**: `CrawlerFactory` class in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) with static `CRAWLERS` registry
- **Registry keys**: Platform strings (`"xhs"`, `"dy"`, `"zhihu"`, etc.) map to concrete classes
- **Creation method**: `create_crawler(platform)` validates input and returns new instances
- **Runtime binding**: `config.PLATFORM` selects the crawler without hard-coded references
- **Extension mechanism**: new platforms register by adding dictionary entries

## Frequently Asked Questions

### What happens if I pass an unsupported platform string to `create_crawler()`?

`create_crawler()` raises a `ValueError` with a descriptive message listing all supported platforms: `f"Invalid media platform: {platform!r}. Supported: {supported}"`. The `supported` string is dynamically generated from `CRAWLERS.keys()`.

### Can I instantiate crawlers directly without the factory?

Yes, but this bypasses the abstraction layer and couples your code to concrete implementations. Direct instantiation (`XiaoHongShuCrawler()`) works for testing but defeats the factory's decoupling benefits in production code.

### How do I add a new platform without modifying [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)?

Currently, registration requires modifying [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) to import the new class and update `CRAWLERS`. For fully dynamic registration, you could implement a decorator-based or entry-point-based discovery system, though this would extend beyond the existing factory pattern.

### Is `CrawlerFactory` thread-safe for concurrent crawler creation?

The factory's `create_crawler()` method is stateless and read-only except for the static dictionary lookup. Dictionary reads in Python are thread-safe, making concurrent calls safe. However, the returned crawler instance's `start()` method must handle its own concurrency constraints.