How to Add a New Media Platform to MediaCrawler's Crawler Factory
To add a new media platform to MediaCrawler, create a concrete crawler class that inherits from AbstractCrawler defined in base/base_crawler.py, place it in media_platform/<platform>/, expose it via __init__.py, and register the class mapping in CrawlerFactory.CRAWLERS inside main.py.
MediaCrawler is an open-source multi-platform content aggregation framework that discovers concrete implementations through a centralized factory pattern. Adding support for new media sources requires implementing three core abstract methods and registering your class without modifying existing platform logic. This guide demonstrates the exact implementation steps using the actual source code from the NanmiCoder/MediaCrawler repository.
Understanding the CrawlerFactory Architecture
The CrawlerFactory class defined in main.py (lines 50‑60) serves as the central registry for all platform implementations. It maintains a static dictionary named CRAWLERS that maps short platform identifiers—such as "bili" for Bilibili or "dy" for DouYin—to concrete crawler classes.
Each registered class must implement the AbstractCrawler interface defined in base/base_crawler.py (lines 26‑34). This interface requires three asynchronous methods: start() for entry point orchestration, search() for content retrieval, and launch_browser() for browser context initialization. The factory's create_crawler() method instantiates the appropriate class based on the platform identifier, raising a ValueError if the identifier is not found in the registry (lines 61‑66 in main.py).
Step-by-Step Implementation Guide
Follow these steps to integrate a new media platform into the factory.
1. Create the Platform Directory Structure
Create a new directory under media_platform/ named after your platform. Mirror the structure of existing implementations like media_platform/bilibili/:
media_platform/
├── newplatform/
│ ├── __init__.py
│ └── core.py
2. Implement the AbstractCrawler Interface
In core.py, define a class that inherits from AbstractCrawler and implements the three required abstract methods:
# media_platform/newplatform/core.py
from base.base_crawler import AbstractCrawler
from typing import Dict, Optional
from playwright.async_api import BrowserContext, BrowserType
class NewPlatformCrawler(AbstractCrawler):
async def start(self):
"""Entry point that orchestrates the crawling workflow."""
await self.search()
# Add post-processing logic here
async def search(self):
"""Platform-specific search implementation."""
# Implement content retrieval logic
print("Searching new platform...")
async def launch_browser(
self,
chromium: BrowserType,
playwright_proxy: Optional[Dict],
user_agent: Optional[str],
headless: bool = True,
) -> BrowserContext:
"""Initialize and return a Playwright browser context."""
from tools.crawler_util import launch_browser
return await launch_browser(
chromium, playwright_proxy, user_agent, headless
)
3. Expose the Class via init.py
Make the crawler importable by updating the package initialization:
# media_platform/newplatform/__init__.py
from .core import NewPlatformCrawler
__all__ = ["NewPlatformCrawler"]
4. Register with CrawlerFactory
Import your class in main.py and add it to the CRAWLERS dictionary:
# main.py
from media_platform.newplatform import NewPlatformCrawler
class CrawlerFactory:
CRAWLERS: dict[str, Type[AbstractCrawler]] = {
"xhs": XiaoHongShuCrawler,
"dy": DouYinCrawler,
"ks": KuaishouCrawler,
"bili": BilibiliCrawler,
"wb": WeiboCrawler,
"tieba": TieBaCrawler,
"zhihu": ZhihuCrawler,
"new": NewPlatformCrawler, # Your new platform
}
5. Update Configuration (Optional)
To make the new platform selectable via configuration, update the default in config/base_config.py:
# config/base_config.py
PLATFORM = "new" # Now defaults to your new crawler
Testing Your New Platform Integration
Verify your implementation by creating a simple test that exercises the factory:
# tests/test_new_platform.py
from main import CrawlerFactory
from media_platform.newplatform import NewPlatformCrawler
def test_factory_creates_new_platform():
crawler = CrawlerFactory.create_crawler("new")
assert isinstance(crawler, NewPlatformCrawler)
assert hasattr(crawler, 'start')
assert hasattr(crawler, 'search')
Existing tests such as tests/test_store_factory.py demonstrate the factory validation pattern used throughout the codebase.
Summary
- Create a crawler class in
media_platform/<platform>/core.pythat inherits fromAbstractCrawler. - Implement three required methods:
start(),search(), andlaunch_browser(). - Expose the class via
__init__.pyto make it importable. - Register the platform identifier and class in
CrawlerFactory.CRAWLERSinmain.py. - Update
config/base_config.pyto set the new platform as default if desired. - Test using the factory to ensure proper instantiation and interface compliance.
Frequently Asked Questions
What methods must I implement when adding a new platform to MediaCrawler?
You must implement three abstract methods defined in base/base_crawler.py: start() for workflow orchestration, search() for content retrieval, and launch_browser() for browser initialization. These methods ensure your crawler follows the expected interface used by the rest of the framework.
Where does MediaCrawler store the platform-to-crawler mapping?
The mapping lives in the CrawlerFactory class within main.py as a static dictionary called CRAWLERS. This dictionary maps string identifiers like "bili" or "new" to concrete crawler classes that inherit from AbstractCrawler.
Do I need to modify existing code to add a new media platform?
No existing platform code requires modification. You only need to create new files in media_platform/<your_platform>/, import your class in main.py, and add one entry to the CRAWLERS dictionary. This follows the open-closed principle, allowing extension without modification.
How does the factory handle unsupported platform identifiers?
If you request a platform not present in CrawlerFactory.CRAWLERS, the create_crawler() method raises a ValueError (see lines 61‑66 in main.py). This prevents the system from attempting to instantiate undefined crawlers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →