How to Contribute New Platform Crawlers to MediaCrawler: A Step-by-Step Guide
To add a new platform crawler to MediaCrawler, you must create a package under media_platform/, implement the AbstractCrawler interface from base/base_crawler.py, register the class in the CrawlerFactory located in main/main.py, and provide platform-specific configuration, storage, and tests.
MediaCrawler is an open-source social media scraping framework built on a clean, plugin-based architecture that makes extending support to new platforms straightforward. By following the established patterns in the codebase, you can contribute to the development of new platform crawlers while maintaining consistency with existing implementations like Tieba and Zhihu. This guide walks you through the concrete steps required to integrate a new platform, referencing actual source files and interfaces from the NanmiCoder/MediaCrawler repository.
Understanding the Architecture
Before writing code, you need to understand how MediaCrawler structures its components. The framework separates concerns into distinct layers: an abstract interface that all crawlers must implement, platform-specific logic for navigation and data extraction, and generic utilities for storage and browser management.
The AbstractCrawler Interface
Every platform crawler must inherit from AbstractCrawler, defined in base/base_crawler.py. This class establishes the asynchronous contract that the main execution loop expects, including the start(), search(), and launch_browser() methods. When you contribute to the development of new platform crawlers, your concrete implementation must override these methods to handle the specific login flows, API endpoints, and pagination logic of your target site.
Platform-Specific Components
Each platform lives in its own package under media_platform/<platform_name>/. A typical implementation includes:
core.py– The main crawler class that orchestrates the browser and API clientclient.py– An async HTTP wrapper or Playwright page controller that handles raw requestslogin.py– Encapsulates authentication flows (QR codes, mobile login, cookie injection)help.py– Contains extractor functions that parse raw HTML/JSON into structured data
Storage and Configuration
Data persistence follows the AbstractStore pattern. You can reuse generic implementations like AsyncFileWriter for CSV/JSONL output, or create a custom store in store/<platform>/_store_impl.py if you need specialized database schemas. Platform-specific settings (endpoints, rate limits, headers) belong in config/<platform>_config.py.
Step-by-Step Implementation Guide
Follow these steps to add a new platform crawler to the MediaCrawler ecosystem.
Step 1: Create the Platform Package
Create a new directory for your platform:
mkdir media_platform/myplatform
touch media_platform/myplatform/__init__.py
Inside this package, you will create core.py, client.py, login.py, and help.py.
Step 2: Implement the Crawler Class
In media_platform/myplatform/core.py, define a class that inherits from AbstractCrawler:
from base.base_crawler import AbstractCrawler
from tools.browser_launcher import launch_browser
from .client import MyPlatformClient
from .login import MyPlatformLogin
from .help import MyPlatformExtractor
from store import myplatform as myplatform_store
import config
import utils
class MyPlatformCrawler(AbstractCrawler):
def __init__(self) -> None:
self.base_url = "https://myplatform.example.com"
self.user_agent = utils.get_user_agent()
self._extractor = MyPlatformExtractor()
self.client: MyPlatformClient | None = None
self.browser_context = None
async def start(self) -> None:
async with async_playwright() as pw:
self.browser_context = await self.launch_browser(
pw.chromium, None, self.user_agent, headless=config.HEADLESS
)
page = await self.browser_context.new_page()
await page.goto(self.base_url)
if not await self._is_logged_in():
login = MyPlatformLogin(page, self.browser_context)
await login.begin()
self.client = MyPlatformClient(httpx_proxy=None)
if config.CRAWLER_TYPE == "search":
await self.search()
elif config.CRAWLER_TYPE == "detail":
await self.get_detail()
async def search(self) -> None:
for kw in config.KEYWORDS.split(","):
posts = await self.client.search_posts(keyword=kw, page=1)
for post in posts:
content = self._extractor.extract_content(post)
await myplatform_store.MyPlatformCsvStoreImplement().store_content(content)
Step 3: Add Helper Modules
Create the supporting infrastructure for your crawler:
- Client (
client.py): Wrap the platform's HTTP API or Playwright interactions - Login (
login.py): Handle authentication using cookies, QR codes, or mobile verification - Extractor (
help.py): Transform raw responses into the data models defined inmodel/
Step 4: Define Data Models
If your platform uses unique data structures, add models to model/ following the pattern of model/m_baidu_tieba.py. These dataclasses or Pydantic models standardize the schema for content, comments, and user profiles across the codebase.
Step 5: Implement Storage
You can reuse existing storage utilities or create a custom implementation. For example, to support CSV output, create a store class in store/myplatform/_store_impl.py:
from base.base_store import AbstractStore
class MyPlatformCsvStoreImplement(AbstractStore):
async def store_content(self, content_item):
# Implementation for persisting content to CSV
pass
Step 6: Register with CrawlerFactory
Edit main/main.py to register your crawler in the CrawlerFactory:
from media_platform.myplatform.core import MyPlatformCrawler
class CrawlerFactory:
CRAWLERS = {
"tieba": TieBaCrawler,
"zhihu": ZhihuCrawler,
"myplatform": MyPlatformCrawler, # Add your entry here
}
This registration allows the CLI and API entry points to instantiate your crawler by name.
Step 7: Add Configuration
Copy an existing configuration file (e.g., config/tieba_config.py) to config/myplatform_config.py and adjust the constants for your platform's endpoints, rate limits, and authentication credentials. Ensure the file is imported in config/__init__.py so it loads at runtime.
Step 8: Write Tests
Place unit and functional tests under tests/ to verify your client, login flow, and storage implementations. Use fixtures from conftest.py to mock HTTP responses or spin up in-memory databases:
import pytest
from media_platform.myplatform.client import MyPlatformClient
@pytest.mark.asyncio
async def test_search_posts(httpx_mock):
httpx_mock.add_response(
url="https://api.myplatform.example.com/search?kw=test&page=1",
json={"posts": [{"id": "1", "title": "Demo"}]}
)
client = MyPlatformClient()
posts = await client.search_posts(keyword="test", page=1)
assert len(posts) == 1
assert posts[0]["title"] == "Demo"
Run pytest -q to ensure all tests pass before submitting your contribution.
Key Implementation Files
When contributing to the development of new platform crawlers, reference these existing files to understand the patterns:
base/base_crawler.py– Defines theAbstractCrawlerinterface that all implementations must extendmedia_platform/tieba/core.py– Complete reference implementation showing browser management and data extractionstore/tieba/_store_impl.py– Example of how to wire storage implementations to the crawlermain/main.py– Contains theCrawlerFactorythat maps platform names to crawler classestools/browser_launcher.pyandtools/cdp_browser.py– Helpers for launching Playwright browsers in standard or CDP mode
Summary
Contributing a new platform crawler to MediaCrawler requires implementing a structured interface and following the repository's architectural conventions:
- Inherit from
AbstractCrawlerinbase/base_crawler.pyand implement thestart()andsearch()methods - Organize platform code into
media_platform/<name>/with separate modules for the client, login, and extraction logic - Register the crawler in the
CrawlerFactorydictionary withinmain/main.py - Provide configuration in
config/<platform>_config.pyfor endpoints and runtime settings - Implement tests under
tests/and ensure the full suite passes withpytest -q
Frequently Asked Questions
What is the minimum code required to add a new platform crawler?
You must create a class inheriting from AbstractCrawler with implemented start() and search() methods, register it in CrawlerFactory (located in main/main.py), and provide a configuration file. While client, login, and storage modules are optional for minimal functionality, they are strongly recommended for production-ready implementations.
How does MediaCrawler handle browser automation for new platforms?
The AbstractCrawler base class provides launch_browser() and launch_browser_with_cdp() methods that wrap Playwright. You can use the helpers in tools/browser_launcher.py for standard headless/headed modes or tools/cdp_browser.py for Chrome DevTools Protocol connections. Your crawler should call these in the start() method to initialize the browser context before navigating to the target site.
Where should I store scraped data when contributing a new crawler?
You can reuse the generic AsyncFileWriter for CSV or JSONL output, or implement a custom store by subclassing AbstractStore in a new file under store/<platform>/_store_impl.py. The crawler instantiates the store class and calls its methods (e.g., store_content()) during the extraction phase.
How do I test my new platform crawler before submitting a pull request?
Write unit tests for your client and extractor functions in tests/, using pytest and fixtures from conftest.py to mock HTTP responses. Run the full test suite with pytest -q to ensure your changes do not break existing crawlers. Additionally, test the integration manually by running the crawler via the CLI with your platform name as the target.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →