How to Add a New Social Media Platform to MediaCrawler: A Complete Implementation Guide

To add a new social media platform to MediaCrawler, you must create a concrete crawler class extending AbstractCrawler, implement the corresponding store module and API client, add a configuration file, and register the mapping in CrawlerFactory.CRAWLERS within main/main.py.

MediaCrawler is a modular async scraping framework that unifies data collection across multiple social networks through a plugin-based architecture. Adding a new social media platform to MediaCrawler involves implementing four distinct layers—configuration, persistent storage, HTTP client, and the crawler itself—while adhering to the async patterns established in the existing Zhihu and Douyin implementations. The following guide references the actual source file paths and method signatures required to integrate a new platform seamlessly.

Architecture Overview

MediaCrawler separates concerns into distinct layers to ensure consistency across platforms:

  • Abstract Base: base/base_crawler.py defines the AbstractCrawler interface with required async methods like start(), search(), and get_specified_notes().
  • Factory Pattern: main/main.py contains CrawlerFactory which maps platform name strings (e.g., "zhihu", "douyin") to concrete crawler classes via the CRAWLERS dictionary.
  • Platform Implementation: Each platform lives under media_platform/<platform>/ with core.py (crawler logic), client.py (API wrapper), and exception.py (custom errors).
  • Storage Layer: store/<platform>/_store_impl.py handles database persistence using platform-specific models.
  • Configuration: config/<platform>_config.py extends BaseConfig to define platform-specific constants like CRAWLER_TYPE and KEYWORDS.

Step-by-Step Implementation

1. Create Platform Configuration

Create a new configuration file under config/ that inherits from BaseConfig. This file defines how the crawler behaves for this platform.


# config/twitter_config.py

from .base_config import BaseConfig

class TwitterConfig(BaseConfig):
    PLATFORM = "twitter"
    CRAWLER_TYPE = "search"  # Options: search | detail | creator

    KEYWORDS = "python,asyncio"
    MAX_PAGES = 10
    ENABLE_IP_PROXY = False

Reference the pattern used in [config/zhihu_config.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py).

2. Define Platform Constants (Optional)

For API endpoints and field mappings, create a constants module under constant/:


# constant/twitter.py

SEARCH_ENDPOINT = "https://api.twitter.com/2/tweets/search/recent"
TWEET_FIELDS = "id,text,author_id,created_at,public_metrics"

3. Implement the Storage Layer

Create a store implementation to persist scraped data. The file must expose async functions that accept your platform's Pydantic models.


# store/twitter/_store_impl.py

from typing import List
from model.m_twitter import TwitterTweet, TwitterComment
from . import twitter_store

async def update_twitter_tweet(tweet: TwitterTweet) -> None:
    """Persist a single tweet."""
    await twitter_store.save_tweet(tweet)

async def batch_update_twitter_comments(comments: List[TwitterComment]) -> None:
    """Bulk insert comments."""
    await twitter_store.save_comments(comments)

This mirrors the structure found in [store/zhihu/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py).

4. Build the API Client

Implement an async client class to handle HTTP requests and authentication. The client is typically instantiated within the crawler's create_<platform>_client() method.


# media_platform/twitter/client.py

import httpx
from typing import Optional, Dict, Any
from tools.utils import logger

class TwitterClient:
    def __init__(self, proxy: Optional[str], headers: Dict[str, str]):
        self.http = httpx.AsyncClient(
            proxy=proxy,
            headers=headers,
            timeout=30.0
        )
    
    async def search_tweets(self, keyword: str, next_token: Optional[str] = None) -> Dict[str, Any]:
        params = {
            "query": keyword,
            "max_results": 20,
            "tweet.fields": "created_at,author_id,public_metrics"
        }
        if next_token:
            params["next_token"] = next_token
        
        resp = await self.http.get(constant.twitter.SEARCH_ENDPOINT, params=params)
        resp.raise_for_status()
        return resp.json()
    
    async def close(self):
        await self.http.aclose()

Follow the pattern in [media_platform/zhihu/client.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py).

5. Create the Crawler Class

Extend AbstractCrawler in media_platform/<platform>/core.py. Implement the required lifecycle methods and handle the crawler_type_var context variable to support different execution modes.


# media_platform/twitter/core.py

from base.base_crawler import AbstractCrawler
from .client import TwitterClient
from .exception import DataFetchError
from var import crawler_type_var, source_keyword_var
from store.twitter import update_twitter_tweet, batch_update_twitter_comments
from tools import utils
import config

class TwitterCrawler(AbstractCrawler):
    async def start(self) -> None:
        """Entry point invoked by CrawlerFactory."""
        self.twitter_client = await self.create_twitter_client()
        
        # Set crawler type in context

        crawler_type_var.set(config.CRAWLER_TYPE)
        
        if config.CRAWLER_TYPE == "search":
            await self.search()
        elif config.CRAWLER_TYPE == "detail":
            await self.get_specified_notes()
        elif config.CRAWLER_TYPE == "creator":
            await self.get_creators_and_notes()
    
    async def create_twitter_client(self) -> TwitterClient:
        """Factory method for client instantiation."""
        return TwitterClient(
            proxy=config.IP_PROXY if config.ENABLE_IP_PROXY else None,
            headers={"User-Agent": "MediaCrawler/1.0", "Authorization": f"Bearer {config.TWITTER_BEARER_TOKEN}"}
        )
    
    async def search(self) -> None:
        """Handle search-based crawling."""
        for keyword in config.KEYWORDS.split(","):
            source_keyword_var.set(keyword)
            next_token = None
            
            for page in range(config.MAX_PAGES):
                try:
                    data = await self.twitter_client.search_tweets(keyword, next_token)
                    tweets = data.get("data", [])
                    
                    if not tweets:
                        break
                    
                    # Persist tweets

                    for tweet in tweets:
                        await update_twitter_tweet(tweet)
                    
                    # Handle pagination

                    next_token = data.get("meta", {}).get("next_token")
                    if not next_token:
                        break
                        
                    await utils.sleep_random(1.0, 3.0)
                    
                except Exception as e:
                    logger.error(f"Error fetching tweets for {keyword}: {e}")
                    break
    
    async def get_specified_notes(self) -> None:
        """Implement for detail mode (single tweet lookup)."""
        pass
    
    async def get_creators_and_notes(self) -> None:
        """Implement for creator mode (user timeline scraping)."""
        pass

Reference the complete implementations in [media_platform/zhihu/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) and [media_platform/douyin/core.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py).

6. Register in CrawlerFactory

Open [main/main.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) and import your crawler class, then add it to the CRAWLERS mapping:

from media_platform.twitter.core import TwitterCrawler

class CrawlerFactory:
    CRAWLERS = {
        "zhihu": ZhihuCrawler,
        "douyin": DouYinCrawler,
        "twitter": TwitterCrawler,  # <-- New platform registered

        # ... existing platforms

    }

The factory's create_crawler() method uses this mapping to instantiate the correct class based on the CLI argument.

7. Add CLI Enumeration

Update the platform enum in [cmd_arg/arg.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) to expose the new option:

from enum import Enum

class CrawlerTypeEnum(str, Enum):
    ZHIHU = "zhihu"
    DOUYIN = "douyin"
    XIAOHONGSHU = "xhs"
    TWITTER = "twitter"  # <-- Added

Users can now invoke the crawler with python main.py --platform twitter.

8. Create Platform Models (Optional)

Define Pydantic models for type safety in model/m_<platform>.py:


# model/m_twitter.py

from pydantic import BaseModel
from datetime import datetime

class TwitterTweet(BaseModel):
    id: str
    text: str
    author_id: str
    created_at: datetime
    retweet_count: int = 0
    like_count: int = 0

Summary

  • Configuration: Create config/<platform>_config.py inheriting from BaseConfig to define CRAWLER_TYPE, KEYWORDS, and proxy settings.
  • Storage: Implement store/<platform>/_store_impl.py with async save functions for your data models.
  • Client: Build an async client in media_platform/<platform>/client.py to handle platform API authentication and requests.
  • Crawler: Extend AbstractCrawler in media_platform/<platform>/core.py, implementing start(), search(), and other required methods while using crawler_type_var for mode detection.
  • Registration: Import and map the class in main/main.py's CrawlerFactory.CRAWLERS dictionary.
  • CLI: Add the platform identifier to CrawlerTypeEnum in cmd_arg/arg.py to enable command-line usage.

Frequently Asked Questions

What methods are required when extending AbstractCrawler?

You must implement start() as the entry point, which should call create_<platform>_client() and dispatch to mode-specific methods like search(), get_specified_notes(), or get_creators_and_notes() based on config.CRAWLER_TYPE. The AbstractCrawler base class in [base/base_crawler.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) defines these abstract methods.

How do I handle platform-specific authentication tokens?

Store sensitive credentials in your platform's config file (e.g., TWITTER_BEARER_TOKEN in config/twitter_config.py) and access them via config imports within your client class. Never hardcode tokens; use environment variables loaded through the config class.

Can I implement a crawler without creating a separate client class?

Yes, but it is not recommended. While you could embed HTTP logic directly in the crawler class, the architecture separates concerns for maintainability. The client class pattern in [media_platform/zhihu/client.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) provides reusable request handling, rate limiting hooks, and connection pooling that should be preserved.

Where does MediaCrawler handle proxy configuration for new platforms?

Proxy settings are read from your platform's config file (e.g., ENABLE_IP_PROXY and IP_PROXY values) and passed to the client constructor during instantiation in the crawler's create_<platform>_client() method. The httpx.AsyncClient accepts the proxy parameter directly, as shown in the Twitter client example above.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →