How to Extend MediaCrawler to Support a New Social Media Platform: A Complete Developer Guide
Extend MediaCrawler by implementing four abstract base classes—AbstractCrawler, AbstractLogin, AbstractStore, and AbstractApiClient—and registering your platform in the CLI parser and crawler manager.
MediaCrawler's modular architecture makes it straightforward to add support for a new social media platform without modifying existing code. The project separates concerns through abstract base classes defined in base/base_crawler.py, allowing platform-specific implementations to plug into the crawling, login, storage, and API client layers.
This guide walks through the complete implementation process using concrete file paths and working code examples from the NanmiCoder/MediaCrawler repository.
Core Architecture and Extension Points
MediaCrawler defines four core abstractions in base/base_crawler.py that every platform must implement:
| Component | Responsibility | File Location |
|---|---|---|
| AbstractCrawler | Async start(), search(), launch_browser() methods |
base/base_crawler.py |
| AbstractLogin | login_by_qrcode(), login_by_mobile(), login_by_cookies() |
base/base_crawler.py |
| AbstractStore | store_content(), store_comment(), store_creator() persistence |
base/base_crawler.py |
| AbstractApiClient | HTTP requests and cookie synchronization | base/base_crawler.py |
Concrete implementations live in namespaced modules under media_platform/<platform>/. The main entry point (main.py) delegates to a factory pattern that instantiates the correct client based on the --platform CLI argument.
Reference implementations exist for Zhihu (media_platform/zhihu/), Douyin, Xiaohongshu, and others—use these as templates when adding your own platform.
Step-by-Step: Adding a New Platform
1. Create the Module Namespace
Create a new directory under media_platform/ with the standard structure:
mkdir -p media_platform/twitter
touch media_platform/twitter/__init__.py
touch media_platform/twitter/client.py
touch media_platform/twitter/login.py
2. Implement the Crawler Client
Subclass AbstractCrawler and provide concrete implementations for the three required async methods. The launch_browser() method can delegate to the parent class implementation in base/base_crawler.py.
# media_platform/twitter/client.py
from base.base_crawler import AbstractCrawler
from playwright.async_api import BrowserContext, Page
from media_platform.twitter.login import TwitterLogin
from media_platform.twitter.api_client import TwitterApiClient
from store.twitter._store_impl import TwitterStore
from config.twitter_config import TwitterConfig
class TwitterCrawler(AbstractCrawler):
def __init__(self, cfg: TwitterConfig):
self.cfg = cfg
self.login = TwitterLogin(cfg)
self.store = TwitterStore(cfg)
self.api_client = TwitterApiClient(cfg)
self.browser_context: BrowserContext | None = None
async def start(self):
"""Initialize browser, authenticate, and begin crawling workflow."""
self.browser_context = await self.launch_browser(
chromium=self.playwright.chromium,
playwright_proxy=self.cfg.proxy,
user_agent=self.cfg.user_agent,
headless=self.cfg.headless,
)
await self.login.begin()
await self.login.login_by_cookies(self.browser_context)
async def search(self, keyword: str, limit: int = 20):
"""Search for content and store results."""
page: Page = await self.browser_context.new_page()
await page.goto(
f"https://twitter.com/search?q={keyword}&f=live",
wait_until="networkidle"
)
tweets = await self._extract_tweets(page, limit)
for tweet in tweets:
await self.store.store_content(tweet)
await page.close()
async def launch_browser(self, chromium, playwright_proxy, user_agent, headless=True):
"""Reuse the base implementation for browser initialization."""
return await super().launch_browser(
chromium, playwright_proxy, user_agent, headless
)
async def _extract_tweets(self, page: Page, limit: int) -> list[dict]:
# Platform-specific scraping logic
pass
3. Implement Login Flows
Subclass AbstractLogin and implement the three required methods. Raise NotImplementedError for unsupported login methods.
# media_platform/twitter/login.py
from base.base_crawler import AbstractLogin
from playwright.async_api import BrowserContext
class TwitterLogin(AbstractLogin):
def __init__(self, cfg):
self.cfg = cfg
self.context = None
async def begin(self):
"""Optional pre-flight initialization."""
pass
async def login_by_qrcode(self):
"""Twitter does not support QR-code login."""
raise NotImplementedError("Twitter does not support QR-code login")
async def login_by_mobile(self):
"""Twitter mobile/SMS login not implemented."""
raise NotImplementedError("Twitter mobile login not implemented")
async def login_by_cookies(self, context: BrowserContext):
"""
Authenticate using cookies from config or saved session.
Called automatically by the crawler after browser launch.
"""
if self.cfg.cookies:
await context.add_cookies(self.cfg.cookies)
await context.storage_state(path=self.cfg.cookie_save_path)
4. Create the Store Implementation
Subclass AbstractStore to persist scraped data. Implement optional AbstractStoreImage and AbstractStoreVideo if your platform returns media files.
# store/twitter/_store_impl.py
import json
import asyncio
from pathlib import Path
from base.base_crawler import AbstractStore, AbstractStoreImage
class TwitterStore(AbstractStore, AbstractStoreImage):
def __init__(self, cfg):
self.cfg = cfg
self.output_dir = Path(cfg.output_path)
self.output_dir.mkdir(parents=True, exist_ok=True)
self.content_file = open(
self.output_dir / "tweets.jsonl",
"w",
encoding="utf-8"
)
self._lock = asyncio.Lock()
async def store_content(self, tweet: dict):
"""Persist a tweet to JSON Lines format."""
async with self._lock:
self.content_file.write(json.dumps(tweet, ensure_ascii=False) + "\n")
async def store_comment(self, reply: dict):
"""Twitter replies are stored as content with type annotation."""
reply["_type"] = "reply"
await self.store_content(reply)
async def store_creator(self, user: dict):
"""Store user profile metadata."""
user["_type"] = "creator"
await self.store_content(user)
async def store_image(self, image_url: str, tweet_id: str):
"""Optional: download and store attached images."""
# Implementation using aiohttp or similar
pass
def __del__(self):
"""Ensure file handle closure."""
self.content_file.close()
5. Add Platform Configuration
Create a dataclass that inherits from BaseConfig with platform-specific fields.
# config/twitter_config.py
from dataclasses import dataclass, field
from base.base_crawler import BaseConfig
@dataclass
class TwitterConfig(BaseConfig):
"""Twitter-specific configuration."""
# Inherited from BaseConfig: proxy, user_agent, headless, etc.
# Platform-specific fields
api_key: str = ""
api_secret: str = ""
bearer_token: str = ""
rate_limit_per_hour: int = 300
# Cookie storage
cookie_save_path: str = "data/twitter_cookies.json"
cookies: list[dict] = field(default_factory=list)
# Output configuration
output_path: str = "data/twitter"
download_media: bool = True
6. Register in the CLI Parser
Add your platform to the choices in cmd_arg/arg.py:
# cmd_arg/arg.py
import argparse
def get_arg_parser():
parser = argparse.ArgumentParser(description="MediaCrawler")
parser.add_argument(
"--platform",
choices=["xhs", "dy", "ks", "zhihu", "bilibili", "weibo", "twitter"],
required=True,
help="Target social media platform to crawl"
)
parser.add_argument(
"--type",
choices=["search", "detail", "creator"],
default="search",
help="Crawling mode"
)
parser.add_argument(
"--keywords",
nargs="+",
help="Search keywords"
)
# ... additional arguments ...
return parser
7. Register in the Crawler Manager
Import and instantiate your crawler in api/services/crawler_manager.py:
# api/services/crawler_manager.py
from config.twitter_config import TwitterConfig
from media_platform.twitter.client import TwitterCrawler
# ... existing imports ...
class CrawlerManager:
def __init__(self):
self._crawlers = {}
def get_crawler(self, platform: str, config_override: dict = None):
"""Factory method returning appropriate crawler instance."""
if platform == "twitter":
cfg = TwitterConfig(**config_override) if config_override else TwitterConfig()
return TwitterCrawler(cfg)
elif platform == "zhihu":
from media_platform.zhihu.client import ZhihuCrawler
from config.zhihu_config import ZhihuConfig
cfg = ZhihuConfig(**config_override) if config_override else ZhihuConfig()
return ZhihuCrawler(cfg)
# ... existing platform branches ...
else:
raise ValueError(f"Unsupported platform: {platform}")
async def run(self, platform: str, **kwargs):
"""Execute crawling workflow for specified platform."""
crawler = self.get_crawler(platform, kwargs.get("config"))
await crawler.start()
if kwargs.get("type") == "search":
await crawler.search(
keyword=kwargs.get("keywords", [""])[0],
limit=kwargs.get("limit", 100)
)
8. Add Optional Constants
For static mappings (URL patterns, API endpoints, selectors), create a constants file:
# constant/twitter.py
"""Twitter-specific constants and selectors."""
BASE_URL = "https://twitter.com"
API_ENDPOINTS = {
"search": "/search",
"user_profile": "/i/api/graphql/UserByScreenName",
"tweet_detail": "/i/api/graphql/TweetDetail",
}
SELECTORS = {
"tweet_article": "article[data-testid='tweet']",
"tweet_text": "[data-testid='tweetText']",
"tweet_time": "time",
"user_handle": "[data-testid='User-Names'] a[role='link']",
}
HEADERS = {
"authorization": "Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA",
"x-twitter-active-user": "yes",
"x-twitter-client-language": "en",
}
Factory Pattern and Loose Coupling
MediaCrawler uses factory patterns in two key locations to maintain loose coupling:
cache/cache_factory.py: ProvidesCacheFactory.get_cache()for platform-agnostic cachingstore/__init__.py: ExposesStoreFactory.get_store()abstracted from concrete implementations
This design means existing core code never changes when adding a new platform. Only new files are created, and two registration points are updated (CLI parser and crawler manager).
Testing Your Implementation
Follow the pattern in tests/ for validation:
# tests/test_twitter_crawler.py
import pytest
from media_platform.twitter.client import TwitterCrawler
from config.twitter_config import TwitterConfig
@pytest.fixture
def mock_config():
return TwitterConfig(
headless=True,
output_path="/tmp/test_output"
)
@pytest.mark.asyncio
async def test_twitter_crawler_init(mock_config):
crawler = TwitterCrawler(mock_config)
assert crawler.cfg == mock_config
assert crawler.login is not None
assert crawler.store is not None
@pytest.mark.asyncio
async def test_login_by_cookies_not_implemented(mock_config):
from media_platform.twitter.login import TwitterLogin
login = TwitterLogin(mock_config)
with pytest.raises(NotImplementedError):
await login.login_by_qrcode()
with pytest.raises(NotImplementedError):
await login.login_by_mobile()
Key Reference Files
| Purpose | Path | Description |
|---|---|---|
| Base abstractions | base/base_crawler.py |
AbstractCrawler, AbstractLogin, AbstractStore definitions |
| Reference client | media_platform/zhihu/client.py |
Working implementation of all three required methods |
| Reference login | media_platform/zhihu/login.py |
QR, mobile, and cookie login flows |
| Reference store | store/zhihu/_store_impl.py |
CSV/JSON/DB persistence patterns |
| Reference config | config/zhihu_config.py |
Dataclass with platform-specific fields |
| Main entry | main.py |
Orchestrates platform selection and execution |
| CLI arguments | cmd_arg/arg.py |
Argument parser with platform choices |
| Crawler manager | api/services/crawler_manager.py |
Factory dispatch logic |
Summary
-
MediaCrawler extension requires implementing four abstract classes:
AbstractCrawler,AbstractLogin,AbstractStore, and optionallyAbstractApiClient. -
File organization follows the pattern
media_platform/<platform>/for client code,store/<platform>/for persistence, andconfig/<platform>_config.pyfor settings. -
Registration requires two edits: add the platform name to
cmd_arg/arg.pychoices, and add a factory branch inapi/services/crawler_manager.py. -
No core modifications are needed—the architecture is designed for additive-only extensions through loose coupling and factory patterns.
-
Reference implementations in
media_platform/zhihu/provide working templates for browser automation, login handling, and data storage.
Frequently Asked Questions
What programming patterns does MediaCrawler use for platform extensibility?
MediaCrawler uses the Abstract Factory and Template Method patterns. Abstract base classes in base/base_crawler.py define the template structure, while concrete implementations in media_platform/<platform>/ provide platform-specific behavior. Factory classes (CrawlerManager, CacheFactory) handle runtime instantiation without hardcoding class names.
Can I implement only one login method if my platform supports only cookies?
Yes. The AbstractLogin interface requires all three methods, but you can raise NotImplementedError for unsupported flows. As shown in the Twitter example, login_by_qrcode() and login_by_mobile() raise exceptions while login_by_cookies() performs actual authentication. The main workflow in AbstractCrawler.start() catches and handles these exceptions gracefully.
How does MediaCrawler handle browser instances across platforms?
The AbstractCrawler.launch_browser() method in base/base_crawler.py provides a standard Playwright initialization that subclasses can call via super().launch_browser(). Each platform client stores the returned BrowserContext and passes it to login and scraping methods. This ensures consistent proxy handling, user-agent rotation, and headless configuration across all platforms.
What storage formats does MediaCrawler support out of the box?
The AbstractStore interface is format-agnostic. Existing implementations in store/zhihu/_store_impl.py demonstrate JSON Lines, CSV, and SQLite persistence. You can implement any storage backend—Elasticsearch, PostgreSQL, S3—by overriding store_content(), store_comment(), and store_creator(). The optional AbstractStoreImage and AbstractStoreVideo mixins add media download capabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →