How the CrawlerFactory Pattern Creates Platform-Specific Crawler Instances in MediaCrawler
MediaCrawler uses a classic Factory pattern via CrawlerFactory class to decouple crawler instantiation from business logic, mapping platform strings like "xhs" or "dy" to concrete crawler classes through a static registry dictionary.
The CrawlerFactory pattern in NanmiCoder/MediaCrawler provides a clean mechanism for creating platform-specific crawler instances without hard-coding class references throughout the application. This design enables seamless switching between social media platforms through configuration changes alone.
Abstract Interface Design
All platform crawlers inherit from AbstractCrawler, defined in [base/base_crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py). This abstract base class establishes the contract that every concrete crawler must fulfill:
async def start(self)– entry point for crawler executionasync def search(self)– handles search-based data collectionasync def launch_browser(self)– manages browser automation setup
By programming against this interface rather than concrete implementations, the application achieves interface-based polymorphism that the factory exploits.
Concrete Crawler Implementations
Each supported platform provides its own subclass of AbstractCrawler within the media_platform package:
| Platform | File | Class Name |
|---|---|---|
| XiaoHongShu (小红书) | [media_platform/xhs/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) |
XiaoHongShuCrawler |
| DouYin (抖音) | [media_platform/douyin/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py) |
DouYinCrawler |
| Zhihu (知乎) | [media_platform/zhihu/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) |
ZhihuCrawler |
| Kuaishou (快手) | media_platform/ks/core.py |
KuaishouCrawler |
| Bilibili | media_platform/bilibili/core.py |
BilibiliCrawler |
media_platform/weibo/core.py |
WeiboCrawler |
|
| TieBa | media_platform/tieba/core.py |
TieBaCrawler |
Each class independently implements the three abstract methods according to platform-specific requirements.
Factory Registry Implementation
The factory itself resides in [main.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) as the CrawlerFactory class. It maintains a static dictionary mapping platform identifiers to crawler classes:
class CrawlerFactory:
CRAWLERS: dict[str, Type[AbstractCrawler]] = {
"xhs": XiaoHongShuCrawler,
"dy": DouYinCrawler,
"ks": KuaishouCrawler,
"bili": BilibiliCrawler,
"wb": WeiboCrawler,
"tieba": TieBaCrawler,
"zhihu": ZhihuCrawler,
}
This registry enables runtime crawler selection based on string configuration rather than compile-time class references.
Factory Method: create_crawler()
The factory exposes a single static method for crawler instantiation:
@staticmethod
def create_crawler(platform: str) -> AbstractCrawler:
crawler_class = CrawlerFactory.CRAWLERS.get(platform)
if not crawler_class:
supported = ", ".join(sorted(CrawlerFactory.CRAWLERS))
raise ValueError(
f"Invalid media platform: {platform!r}. Supported: {supported}"
)
return crawler_class()
Key characteristics of this implementation:
- Defensive validation – raises
ValueErrorwith enumerated supported platforms for invalid keys - Instance creation – returns a new instance, not the class itself
- Return type –
AbstractCrawlerensures callers depend on the interface, not concrete types
Runtime Usage in the Application
The entry point demonstrates clean factory consumption:
# From main.py
crawler = CrawlerFactory.create_crawler(platform=config.PLATFORM)
await crawler.start()
The config.PLATFORM value (set in [config/base_config.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)) determines which crawler executes. Callers remain oblivious to the concrete class—only the interface matters.
Practical Code Examples
Direct Factory Usage
from main import CrawlerFactory
from config import PLATFORM # e.g., "xhs"
crawler = CrawlerFactory.create_crawler(PLATFORM)
await crawler.start() # Executes XiaoHongShuCrawler.start()
Extending with a Custom Platform
Adding new platforms requires no modifications to existing code:
# Step 1: Implement AbstractCrawler subclass
# media_platform/custom/core.py
from base.base_crawler import AbstractCrawler
class CustomCrawler(AbstractCrawler):
async def start(self): ...
async def search(self): ...
async def launch_browser(self): ...
# Step 2: Register in factory
from media_platform.custom.core import CustomCrawler
CrawlerFactory.CRAWLERS["custom"] = CustomCrawler
# Step 3: Execute
crawler = CrawlerFactory.create_crawler("custom")
await crawler.start()
Extensibility Architecture
The factory pattern delivers three architectural benefits:
- Open/Closed Principle – new platforms extend without modifying factory logic or caller code
- Single Responsibility – creation logic isolated from execution logic
- Testability – interface-based design enables mock crawlers for unit testing
Adding a platform follows a three-step process: implement the abstract class, import it in main.py, and register in CRAWLERS. No other changes propagate through the codebase.
Summary
- Abstract base:
AbstractCrawlerinbase/base_crawler.pydefines the crawler contract - Factory location:
CrawlerFactoryclass inmain.pywith staticCRAWLERSregistry - Registry keys: Platform strings (
"xhs","dy","zhihu", etc.) map to concrete classes - Creation method:
create_crawler(platform)validates input and returns new instances - Runtime binding:
config.PLATFORMselects the crawler without hard-coded references - Extension mechanism: new platforms register by adding dictionary entries
Frequently Asked Questions
What happens if I pass an unsupported platform string to create_crawler()?
create_crawler() raises a ValueError with a descriptive message listing all supported platforms: f"Invalid media platform: {platform!r}. Supported: {supported}". The supported string is dynamically generated from CRAWLERS.keys().
Can I instantiate crawlers directly without the factory?
Yes, but this bypasses the abstraction layer and couples your code to concrete implementations. Direct instantiation (XiaoHongShuCrawler()) works for testing but defeats the factory's decoupling benefits in production code.
How do I add a new platform without modifying main.py?
Currently, registration requires modifying main.py to import the new class and update CRAWLERS. For fully dynamic registration, you could implement a decorator-based or entry-point-based discovery system, though this would extend beyond the existing factory pattern.
Is CrawlerFactory thread-safe for concurrent crawler creation?
The factory's create_crawler() method is stateless and read-only except for the static dictionary lookup. Dictionary reads in Python are thread-safe, making concurrent calls safe. However, the returned crawler instance's start() method must handle its own concurrency constraints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →