Common Interfaces and Base Classes for Platform Crawlers in MediaCrawler
MediaCrawler defines a comprehensive hierarchy of abstract base classes in base/base_crawler.py that standardize crawling, authentication, storage, and API interactions across all supported social media platforms.
The MediaCrawler project implements a modular, extensible architecture that simplifies adding support for new social media platforms. By defining common interfaces for platform crawlers, the repository ensures consistent behavior whether scraping Zhihu, Weibo, TikTok, or Bilibili. These abstract base classes establish a strict contract that every platform-specific implementation must follow, enabling the framework to treat all crawlers as interchangeable components.
Core Abstract Base Classes in base/base_crawler.py
The foundation of MediaCrawler's architecture resides in base/base_crawler.py, which defines five primary abstract classes that separate concerns across the crawling lifecycle.
AbstractCrawler
The AbstractCrawler class defines the core workflow for browser automation and content discovery. It requires concrete implementations to provide start(), search(), and launch_browser() methods, with an optional launch_browser_with_cdp() for Chrome DevTools Protocol support. Platform implementations like ZhihuCrawler and DouyinCrawler inherit from this class to handle site-specific navigation logic while adhering to the universal execution pattern defined in the base class.
AbstractLogin
Authentication strategies are standardized through the AbstractLogin interface, which mandates begin(), login_by_qrcode(), login_by_mobile(), and login_by_cookies() methods. This abstraction allows the framework to support multiple login mechanisms without coupling the crawler logic to specific authentication flows. Concrete classes such as ZhihuLogin and WeiboLogin implement these methods to handle platform-specific authentication while presenting a uniform interface to the crawler manager.
AbstractStore and Media Storage
Data persistence is handled through AbstractStore, which requires store_content(), store_comment(), and store_creator() implementations. For platforms handling media files, the architecture extends to AbstractStoreImage and AbstractStoreVideo with corresponding store_image() and store_video() methods, currently leveraged by Weibo implementations. Concrete classes in store/<platform>/_store_impl.py provide backend-specific logic for CSV, JSON, SQLite, and MongoDB storage.
AbstractApiClient
Low-level HTTP communication is abstracted through AbstractApiClient, which specifies request() and update_cookies() methods. Platform-specific clients like ZhihuApiClient and DouyinApiClient implement these in media_platform/<platform>/client.py to handle authentication headers, rate limiting, and endpoint-specific logic while exposing a consistent interface for data retrieval.
Implementation Architecture
Concrete platform crawlers reside in media_platform/<platform>/core.py and implement the abstract methods defined in base/base_crawler.py. This structure ensures that api/services/crawler_manager.py can instantiate and orchestrate any crawler without platform-specific knowledge. The FastAPI routes in api/routers/crawler.py operate on these common interfaces, making the REST API agnostic to whether it's controlling a Weibo or Bilibili crawler.
Practical Implementation Examples
The following examples demonstrate how platform-specific implementations extend the base classes to create functional crawlers.
Zhihu crawler implementation:
from base.base_crawler import AbstractCrawler, AbstractLogin, AbstractStore
from media_platform.zhihu.client import ZhihuApiClient
from store.zhihu._store_impl import ZhihuJsonStoreImplement
class ZhihuCrawler(AbstractCrawler):
async def launch_browser(self, chromium, playwright_proxy, user_agent, headless=True):
# concrete browser‑launch logic (omitted for brevity)
pass
async def start(self):
# orchestrate login → search → store
login = ZhihuLogin()
await login.begin()
await self.search()
async def search(self):
client = ZhihuApiClient()
results = await client.request('GET', '/api/v4/search', params={'q': 'AI'})
await ZhihuJsonStoreImplement().store_content(results)
class ZhihuLogin(AbstractLogin):
async def begin(self):
# choose login method automatically
await self.login_by_qrcode()
async def login_by_qrcode(self):
pass
async def login_by_mobile(self):
pass
async def login_by_cookies(self):
pass
Generic JSON store implementation:
from base.base_crawler import AbstractStore
import json
import aiofiles
class JsonFileStore(AbstractStore):
async def store_content(self, content_item):
async with aiofiles.open('content.json', 'a') as f:
await f.write(json.dumps(content_item) + '\n')
async def store_comment(self, comment_item):
async with aiofiles.open('comments.json', 'a') as f:
await f.write(json.dumps(comment_item) + '\n')
async def store_creator(self, creator):
async with aiofiles.open('creators.json', 'a') as f:
await f.write(json.dumps(creator) + '\n')
Summary
- AbstractCrawler in
base/base_crawler.pydefines the core crawling workflow withstart(),search(), andlaunch_browser()methods that all platform implementations must override. - AbstractLogin standardizes authentication across QR codes, mobile numbers, and cookies through a consistent interface implemented by platform-specific login classes.
- AbstractStore, AbstractStoreImage, and AbstractStoreVideo provide uniform data persistence contracts for content, comments, creators, and media files.
- AbstractApiClient abstracts HTTP interactions, allowing platform-specific API clients to handle authentication and rate limiting while presenting a consistent
request()interface. - The architecture enables
api/services/crawler_manager.pyto treat all platform crawlers as interchangeable components, simplifying the addition of new social media platforms.
Frequently Asked Questions
What file contains the base class definitions for MediaCrawler platform implementations?
All abstract base classes are defined in base/base_crawler.py. This file contains AbstractCrawler, AbstractLogin, AbstractStore, AbstractStoreImage, AbstractStoreVideo, and AbstractApiClient, which establish the contracts that every platform-specific crawler must implement.
How does MediaCrawler ensure consistency across different social media platforms?
The repository enforces consistency through Python's abstract base class mechanism. By inheriting from the abstract classes in base/base_crawler.py, concrete implementations like ZhihuCrawler or WeiboCrawler are required to implement specific methods such as start(), search(), and store_content(). This guarantees that the crawler manager can invoke the same methods regardless of the underlying platform.
Can I add a new platform crawler without modifying existing code?
Yes. You can create a new platform crawler by implementing the abstract methods defined in base/base_crawler.py within a new module under media_platform/<new_platform>/. As long as your concrete classes inherit from AbstractCrawler, AbstractLogin, and AbstractStore, the existing crawler manager and API routers will automatically support the new platform without requiring changes to the core framework.
What storage backends are supported through the AbstractStore interface?
The AbstractStore interface supports multiple backends including CSV, JSON, SQLite, and MongoDB. Concrete implementations are typically located in store/<platform>/_store_impl.py and can be swapped by configuring the appropriate store class that implements the required store_content(), store_comment(), and store_creator() methods.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →