# How to Add a New Social Media Platform to MediaCrawler: A Complete Implementation Guide

> Learn how to add a new social media platform to MediaCrawler. This guide covers implementing crawler classes, store modules, API clients, and configuration for seamless integration.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**To add a new social media platform to MediaCrawler, you must create a concrete crawler class extending `AbstractCrawler`, implement the corresponding store module and API client, add a configuration file, and register the mapping in `CrawlerFactory.CRAWLERS` within [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py).**

MediaCrawler is a modular async scraping framework that unifies data collection across multiple social networks through a plugin-based architecture. Adding a new social media platform to MediaCrawler involves implementing four distinct layers—configuration, persistent storage, HTTP client, and the crawler itself—while adhering to the async patterns established in the existing Zhihu and Douyin implementations. The following guide references the actual source file paths and method signatures required to integrate a new platform seamlessly.

## Architecture Overview

MediaCrawler separates concerns into distinct layers to ensure consistency across platforms:

- **Abstract Base**: [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) defines the `AbstractCrawler` interface with required async methods like `start()`, `search()`, and `get_specified_notes()`.
- **Factory Pattern**: [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) contains `CrawlerFactory` which maps platform name strings (e.g., `"zhihu"`, `"douyin"`) to concrete crawler classes via the `CRAWLERS` dictionary.
- **Platform Implementation**: Each platform lives under `media_platform/<platform>/` with [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py) (crawler logic), [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) (API wrapper), and [`exception.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/exception.py) (custom errors).
- **Storage Layer**: `store/<platform>/_store_impl.py` handles database persistence using platform-specific models.
- **Configuration**: `config/<platform>_config.py` extends `BaseConfig` to define platform-specific constants like `CRAWLER_TYPE` and `KEYWORDS`.

## Step-by-Step Implementation

### 1. Create Platform Configuration

Create a new configuration file under `config/` that inherits from `BaseConfig`. This file defines how the crawler behaves for this platform.

```python

# config/twitter_config.py

from .base_config import BaseConfig

class TwitterConfig(BaseConfig):
    PLATFORM = "twitter"
    CRAWLER_TYPE = "search"  # Options: search | detail | creator

    KEYWORDS = "python,asyncio"
    MAX_PAGES = 10
    ENABLE_IP_PROXY = False

```

Reference the pattern used in [[`config/zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py).

### 2. Define Platform Constants (Optional)

For API endpoints and field mappings, create a constants module under `constant/`:

```python

# constant/twitter.py

SEARCH_ENDPOINT = "https://api.twitter.com/2/tweets/search/recent"
TWEET_FIELDS = "id,text,author_id,created_at,public_metrics"

```

### 3. Implement the Storage Layer

Create a store implementation to persist scraped data. The file must expose async functions that accept your platform's Pydantic models.

```python

# store/twitter/_store_impl.py

from typing import List
from model.m_twitter import TwitterTweet, TwitterComment
from . import twitter_store

async def update_twitter_tweet(tweet: TwitterTweet) -> None:
    """Persist a single tweet."""
    await twitter_store.save_tweet(tweet)

async def batch_update_twitter_comments(comments: List[TwitterComment]) -> None:
    """Bulk insert comments."""
    await twitter_store.save_comments(comments)

```

This mirrors the structure found in [[`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py).

### 4. Build the API Client

Implement an async client class to handle HTTP requests and authentication. The client is typically instantiated within the crawler's `create_<platform>_client()` method.

```python

# media_platform/twitter/client.py

import httpx
from typing import Optional, Dict, Any
from tools.utils import logger

class TwitterClient:
    def __init__(self, proxy: Optional[str], headers: Dict[str, str]):
        self.http = httpx.AsyncClient(
            proxy=proxy,
            headers=headers,
            timeout=30.0
        )
    
    async def search_tweets(self, keyword: str, next_token: Optional[str] = None) -> Dict[str, Any]:
        params = {
            "query": keyword,
            "max_results": 20,
            "tweet.fields": "created_at,author_id,public_metrics"
        }
        if next_token:
            params["next_token"] = next_token
        
        resp = await self.http.get(constant.twitter.SEARCH_ENDPOINT, params=params)
        resp.raise_for_status()
        return resp.json()
    
    async def close(self):
        await self.http.aclose()

```

Follow the pattern in [[`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py).

### 5. Create the Crawler Class

Extend `AbstractCrawler` in `media_platform/<platform>/core.py`. Implement the required lifecycle methods and handle the `crawler_type_var` context variable to support different execution modes.

```python

# media_platform/twitter/core.py

from base.base_crawler import AbstractCrawler
from .client import TwitterClient
from .exception import DataFetchError
from var import crawler_type_var, source_keyword_var
from store.twitter import update_twitter_tweet, batch_update_twitter_comments
from tools import utils
import config

class TwitterCrawler(AbstractCrawler):
    async def start(self) -> None:
        """Entry point invoked by CrawlerFactory."""
        self.twitter_client = await self.create_twitter_client()
        
        # Set crawler type in context

        crawler_type_var.set(config.CRAWLER_TYPE)
        
        if config.CRAWLER_TYPE == "search":
            await self.search()
        elif config.CRAWLER_TYPE == "detail":
            await self.get_specified_notes()
        elif config.CRAWLER_TYPE == "creator":
            await self.get_creators_and_notes()
    
    async def create_twitter_client(self) -> TwitterClient:
        """Factory method for client instantiation."""
        return TwitterClient(
            proxy=config.IP_PROXY if config.ENABLE_IP_PROXY else None,
            headers={"User-Agent": "MediaCrawler/1.0", "Authorization": f"Bearer {config.TWITTER_BEARER_TOKEN}"}
        )
    
    async def search(self) -> None:
        """Handle search-based crawling."""
        for keyword in config.KEYWORDS.split(","):
            source_keyword_var.set(keyword)
            next_token = None
            
            for page in range(config.MAX_PAGES):
                try:
                    data = await self.twitter_client.search_tweets(keyword, next_token)
                    tweets = data.get("data", [])
                    
                    if not tweets:
                        break
                    
                    # Persist tweets

                    for tweet in tweets:
                        await update_twitter_tweet(tweet)
                    
                    # Handle pagination

                    next_token = data.get("meta", {}).get("next_token")
                    if not next_token:
                        break
                        
                    await utils.sleep_random(1.0, 3.0)
                    
                except Exception as e:
                    logger.error(f"Error fetching tweets for {keyword}: {e}")
                    break
    
    async def get_specified_notes(self) -> None:
        """Implement for detail mode (single tweet lookup)."""
        pass
    
    async def get_creators_and_notes(self) -> None:
        """Implement for creator mode (user timeline scraping)."""
        pass

```

Reference the complete implementations in [[`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py) and [[`media_platform/douyin/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/douyin/core.py).

### 6. Register in CrawlerFactory

Open [[`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) and import your crawler class, then add it to the `CRAWLERS` mapping:

```python
from media_platform.twitter.core import TwitterCrawler

class CrawlerFactory:
    CRAWLERS = {
        "zhihu": ZhihuCrawler,
        "douyin": DouYinCrawler,
        "twitter": TwitterCrawler,  # <-- New platform registered

        # ... existing platforms

    }

```

The factory's `create_crawler()` method uses this mapping to instantiate the correct class based on the CLI argument.

### 7. Add CLI Enumeration

Update the platform enum in [[`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) to expose the new option:

```python
from enum import Enum

class CrawlerTypeEnum(str, Enum):
    ZHIHU = "zhihu"
    DOUYIN = "douyin"
    XIAOHONGSHU = "xhs"
    TWITTER = "twitter"  # <-- Added

```

Users can now invoke the crawler with `python main.py --platform twitter`.

### 8. Create Platform Models (Optional)

Define Pydantic models for type safety in `model/m_<platform>.py`:

```python

# model/m_twitter.py

from pydantic import BaseModel
from datetime import datetime

class TwitterTweet(BaseModel):
    id: str
    text: str
    author_id: str
    created_at: datetime
    retweet_count: int = 0
    like_count: int = 0

```

## Summary

- **Configuration**: Create `config/<platform>_config.py` inheriting from `BaseConfig` to define `CRAWLER_TYPE`, `KEYWORDS`, and proxy settings.
- **Storage**: Implement `store/<platform>/_store_impl.py` with async save functions for your data models.
- **Client**: Build an async client in `media_platform/<platform>/client.py` to handle platform API authentication and requests.
- **Crawler**: Extend `AbstractCrawler` in `media_platform/<platform>/core.py`, implementing `start()`, `search()`, and other required methods while using `crawler_type_var` for mode detection.
- **Registration**: Import and map the class in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py)'s `CrawlerFactory.CRAWLERS` dictionary.
- **CLI**: Add the platform identifier to `CrawlerTypeEnum` in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) to enable command-line usage.

## Frequently Asked Questions

### What methods are required when extending AbstractCrawler?

You must implement `start()` as the entry point, which should call `create_<platform>_client()` and dispatch to mode-specific methods like `search()`, `get_specified_notes()`, or `get_creators_and_notes()` based on `config.CRAWLER_TYPE`. The `AbstractCrawler` base class in [[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) defines these abstract methods.

### How do I handle platform-specific authentication tokens?

Store sensitive credentials in your platform's config file (e.g., `TWITTER_BEARER_TOKEN` in [`config/twitter_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/twitter_config.py)) and access them via `config` imports within your client class. Never hardcode tokens; use environment variables loaded through the config class.

### Can I implement a crawler without creating a separate client class?

Yes, but it is not recommended. While you could embed HTTP logic directly in the crawler class, the architecture separates concerns for maintainability. The client class pattern in [[`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) provides reusable request handling, rate limiting hooks, and connection pooling that should be preserved.

### Where does MediaCrawler handle proxy configuration for new platforms?

Proxy settings are read from your platform's config file (e.g., `ENABLE_IP_PROXY` and `IP_PROXY` values) and passed to the client constructor during instantiation in the crawler's `create_<platform>_client()` method. The `httpx.AsyncClient` accepts the proxy parameter directly, as shown in the Twitter client example above.