How to Add a New Social Media Platform to MediaCrawler: A Complete Implementation Guide
To add a new social media platform to MediaCrawler, you must create a concrete crawler class extending AbstractCrawler, implement the corresponding store module and API client, add a configuration file, and register the mapping in CrawlerFactory.CRAWLERS within main/main.py.
MediaCrawler is a modular async scraping framework that unifies data collection across multiple social networks through a plugin-based architecture. Adding a new social media platform to MediaCrawler involves implementing four distinct layers—configuration, persistent storage, HTTP client, and the crawler itself—while adhering to the async patterns established in the existing Zhihu and Douyin implementations. The following guide references the actual source file paths and method signatures required to integrate a new platform seamlessly.
Architecture Overview
MediaCrawler separates concerns into distinct layers to ensure consistency across platforms:
- Abstract Base:
base/base_crawler.pydefines theAbstractCrawlerinterface with required async methods likestart(),search(), andget_specified_notes(). - Factory Pattern:
main/main.pycontainsCrawlerFactorywhich maps platform name strings (e.g.,"zhihu","douyin") to concrete crawler classes via theCRAWLERSdictionary. - Platform Implementation: Each platform lives under
media_platform/<platform>/withcore.py(crawler logic),client.py(API wrapper), andexception.py(custom errors). - Storage Layer:
store/<platform>/_store_impl.pyhandles database persistence using platform-specific models. - Configuration:
config/<platform>_config.pyextendsBaseConfigto define platform-specific constants likeCRAWLER_TYPEandKEYWORDS.
Step-by-Step Implementation
1. Create Platform Configuration
Create a new configuration file under config/ that inherits from BaseConfig. This file defines how the crawler behaves for this platform.
# config/twitter_config.py
from .base_config import BaseConfig
class TwitterConfig(BaseConfig):
PLATFORM = "twitter"
CRAWLER_TYPE = "search" # Options: search | detail | creator
KEYWORDS = "python,asyncio"
MAX_PAGES = 10
ENABLE_IP_PROXY = False
Reference the pattern used in [config/zhihu_config.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py).
2. Define Platform Constants (Optional)
For API endpoints and field mappings, create a constants module under constant/:
# constant/twitter.py
SEARCH_ENDPOINT = "https://api.twitter.com/2/tweets/search/recent"
TWEET_FIELDS = "id,text,author_id,created_at,public_metrics"
3. Implement the Storage Layer
Create a store implementation to persist scraped data. The file must expose async functions that accept your platform's Pydantic models.
# store/twitter/_store_impl.py
from typing import List
from model.m_twitter import TwitterTweet, TwitterComment
from . import twitter_store
async def update_twitter_tweet(tweet: TwitterTweet) -> None:
"""Persist a single tweet."""
await twitter_store.save_tweet(tweet)
async def batch_update_twitter_comments(comments: List[TwitterComment]) -> None:
"""Bulk insert comments."""
await twitter_store.save_comments(comments)
This mirrors the structure found in [store/zhihu/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py).
4. Build the API Client
Implement an async client class to handle HTTP requests and authentication. The client is typically instantiated within the crawler's create_<platform>_client() method.
# media_platform/twitter/client.py
import httpx
from typing import Optional, Dict, Any
from tools.utils import logger
class TwitterClient:
def __init__(self, proxy: Optional[str], headers: Dict[str, str]):
self.http = httpx.AsyncClient(
proxy=proxy,
headers=headers,
timeout=30.0
)
async def search_tweets(self, keyword: str, next_token: Optional[str] = None) -> Dict[str, Any]:
params = {
"query": keyword,
"max_results": 20,
"tweet.fields": "created_at,author_id,public_metrics"
}
if next_token:
params["next_token"] = next_token
resp = await self.http.get(constant.twitter.SEARCH_ENDPOINT, params=params)
resp.raise_for_status()
return resp.json()
async def close(self):
await self.http.aclose()
Follow the pattern in [media_platform/zhihu/client.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py).
5. Create the Crawler Class
Extend AbstractCrawler in media_platform/<platform>/core.py. Implement the required lifecycle methods and handle the crawler_type_var context variable to support different execution modes.
# media_platform/twitter/core.py
from base.base_crawler import AbstractCrawler
from .client import TwitterClient
from .exception import DataFetchError
from var import crawler_type_var, source_keyword_var
from store.twitter import update_twitter_tweet, batch_update_twitter_comments
from tools import utils
import config
class TwitterCrawler(AbstractCrawler):
async def start(self) -> None:
"""Entry point invoked by CrawlerFactory."""
self.twitter_client = await self.create_twitter_client()
# Set crawler type in context
crawler_type_var.set(config.CRAWLER_TYPE)
if config.CRAWLER_TYPE == "search":
await self.search()
elif config.CRAWLER_TYPE == "detail":
await self.get_specified_notes()
elif config.CRAWLER_TYPE == "creator":
await self.get_creators_and_notes()
async def create_twitter_client(self) -> TwitterClient:
"""Factory method for client instantiation."""
return TwitterClient(
proxy=config.IP_PROXY if config.ENABLE_IP_PROXY else None,
headers={"User-Agent": "MediaCrawler/1.0", "Authorization": f"Bearer {config.TWITTER_BEARER_TOKEN}"}
)
async def search(self) -> None:
"""Handle search-based crawling."""
for keyword in config.KEYWORDS.split(","):
source_keyword_var.set(keyword)
next_token = None
for page in range(config.MAX_PAGES):
try:
data = await self.twitter_client.search_tweets(keyword, next_token)
tweets = data.get("data", [])
if not tweets:
break
# Persist tweets
for tweet in tweets:
await update_twitter_tweet(tweet)
# Handle pagination
next_token = data.get("meta", {}).get("next_token")
if not next_token:
break
await utils.sleep_random(1.0, 3.0)
except Exception as e:
logger.error(f"Error fetching tweets for {keyword}: {e}")
break
async def get_specified_notes(self) -> None:
"""Implement for detail mode (single tweet lookup)."""
pass
async def get_creators_and_notes(self) -> None:
"""Implement for creator mode (user timeline scraping)."""
pass
Reference the complete implementations in [media_platform/zhihu/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) and [media_platform/douyin/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py).
6. Register in CrawlerFactory
Open [main/main.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) and import your crawler class, then add it to the CRAWLERS mapping:
from media_platform.twitter.core import TwitterCrawler
class CrawlerFactory:
CRAWLERS = {
"zhihu": ZhihuCrawler,
"douyin": DouYinCrawler,
"twitter": TwitterCrawler, # <-- New platform registered
# ... existing platforms
}
The factory's create_crawler() method uses this mapping to instantiate the correct class based on the CLI argument.
7. Add CLI Enumeration
Update the platform enum in [cmd_arg/arg.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) to expose the new option:
from enum import Enum
class CrawlerTypeEnum(str, Enum):
ZHIHU = "zhihu"
DOUYIN = "douyin"
XIAOHONGSHU = "xhs"
TWITTER = "twitter" # <-- Added
Users can now invoke the crawler with python main.py --platform twitter.
8. Create Platform Models (Optional)
Define Pydantic models for type safety in model/m_<platform>.py:
# model/m_twitter.py
from pydantic import BaseModel
from datetime import datetime
class TwitterTweet(BaseModel):
id: str
text: str
author_id: str
created_at: datetime
retweet_count: int = 0
like_count: int = 0
Summary
- Configuration: Create
config/<platform>_config.pyinheriting fromBaseConfigto defineCRAWLER_TYPE,KEYWORDS, and proxy settings. - Storage: Implement
store/<platform>/_store_impl.pywith async save functions for your data models. - Client: Build an async client in
media_platform/<platform>/client.pyto handle platform API authentication and requests. - Crawler: Extend
AbstractCrawlerinmedia_platform/<platform>/core.py, implementingstart(),search(), and other required methods while usingcrawler_type_varfor mode detection. - Registration: Import and map the class in
main/main.py'sCrawlerFactory.CRAWLERSdictionary. - CLI: Add the platform identifier to
CrawlerTypeEnumincmd_arg/arg.pyto enable command-line usage.
Frequently Asked Questions
What methods are required when extending AbstractCrawler?
You must implement start() as the entry point, which should call create_<platform>_client() and dispatch to mode-specific methods like search(), get_specified_notes(), or get_creators_and_notes() based on config.CRAWLER_TYPE. The AbstractCrawler base class in [base/base_crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) defines these abstract methods.
How do I handle platform-specific authentication tokens?
Store sensitive credentials in your platform's config file (e.g., TWITTER_BEARER_TOKEN in config/twitter_config.py) and access them via config imports within your client class. Never hardcode tokens; use environment variables loaded through the config class.
Can I implement a crawler without creating a separate client class?
Yes, but it is not recommended. While you could embed HTTP logic directly in the crawler class, the architecture separates concerns for maintainability. The client class pattern in [media_platform/zhihu/client.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) provides reusable request handling, rate limiting hooks, and connection pooling that should be preserved.
Where does MediaCrawler handle proxy configuration for new platforms?
Proxy settings are read from your platform's config file (e.g., ENABLE_IP_PROXY and IP_PROXY values) and passed to the client constructor during instantiation in the crawler's create_<platform>_client() method. The httpx.AsyncClient accepts the proxy parameter directly, as shown in the Twitter client example above.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →