# How to Extend MediaCrawler to Support a New Social Media Platform: A Complete Developer Guide

> Learn to extend MediaCrawler for new platforms by implementing four abstract classes and registering your platform. A complete developer guide for NanmiCoder/MediaCrawler.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-08-14

---

**Extend MediaCrawler by implementing four abstract base classes—AbstractCrawler, AbstractLogin, AbstractStore, and AbstractApiClient—and registering your platform in the CLI parser and crawler manager.**

MediaCrawler's modular architecture makes it straightforward to add **support for a new social media platform** without modifying existing code. The project separates concerns through abstract base classes defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py), allowing platform-specific implementations to plug into the crawling, login, storage, and API client layers.

This guide walks through the complete implementation process using concrete file paths and working code examples from the NanmiCoder/MediaCrawler repository.

## Core Architecture and Extension Points

MediaCrawler defines **four core abstractions** in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) that every platform must implement:

| Component | Responsibility | File Location |
|-----------|--------------|---------------|
| **AbstractCrawler** | Async `start()`, `search()`, `launch_browser()` methods | [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) |
| **AbstractLogin** | `login_by_qrcode()`, `login_by_mobile()`, `login_by_cookies()` | [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) |
| **AbstractStore** | `store_content()`, `store_comment()`, `store_creator()` persistence | [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) |
| **AbstractApiClient** | HTTP requests and cookie synchronization | [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) |

Concrete implementations live in **namespaced modules** under `media_platform/<platform>/`. The main entry point ([`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)) delegates to a **factory pattern** that instantiates the correct client based on the `--platform` CLI argument.

Reference implementations exist for **Zhihu** (`media_platform/zhihu/`), **Douyin**, **Xiaohongshu**, and others—use these as templates when adding your own platform.

## Step-by-Step: Adding a New Platform

### 1. Create the Module Namespace

Create a new directory under `media_platform/` with the standard structure:

```bash
mkdir -p media_platform/twitter
touch media_platform/twitter/__init__.py
touch media_platform/twitter/client.py
touch media_platform/twitter/login.py

```

### 2. Implement the Crawler Client

Subclass `AbstractCrawler` and provide concrete implementations for the three required async methods. The `launch_browser()` method can delegate to the parent class implementation in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py).

```python

# media_platform/twitter/client.py

from base.base_crawler import AbstractCrawler
from playwright.async_api import BrowserContext, Page

from media_platform.twitter.login import TwitterLogin
from media_platform.twitter.api_client import TwitterApiClient
from store.twitter._store_impl import TwitterStore
from config.twitter_config import TwitterConfig


class TwitterCrawler(AbstractCrawler):
    def __init__(self, cfg: TwitterConfig):
        self.cfg = cfg
        self.login = TwitterLogin(cfg)
        self.store = TwitterStore(cfg)
        self.api_client = TwitterApiClient(cfg)
        self.browser_context: BrowserContext | None = None

    async def start(self):
        """Initialize browser, authenticate, and begin crawling workflow."""
        self.browser_context = await self.launch_browser(
            chromium=self.playwright.chromium,
            playwright_proxy=self.cfg.proxy,
            user_agent=self.cfg.user_agent,
            headless=self.cfg.headless,
        )
        await self.login.begin()
        await self.login.login_by_cookies(self.browser_context)

    async def search(self, keyword: str, limit: int = 20):
        """Search for content and store results."""
        page: Page = await self.browser_context.new_page()
        await page.goto(
            f"https://twitter.com/search?q={keyword}&f=live",
            wait_until="networkidle"
        )
        
        tweets = await self._extract_tweets(page, limit)
        for tweet in tweets:
            await self.store.store_content(tweet)
        
        await page.close()

    async def launch_browser(self, chromium, playwright_proxy, user_agent, headless=True):
        """Reuse the base implementation for browser initialization."""
        return await super().launch_browser(
            chromium, playwright_proxy, user_agent, headless
        )

    async def _extract_tweets(self, page: Page, limit: int) -> list[dict]:
        # Platform-specific scraping logic

        pass

```

### 3. Implement Login Flows

Subclass `AbstractLogin` and implement the three required methods. Raise `NotImplementedError` for unsupported login methods.

```python

# media_platform/twitter/login.py

from base.base_crawler import AbstractLogin
from playwright.async_api import BrowserContext


class TwitterLogin(AbstractLogin):
    def __init__(self, cfg):
        self.cfg = cfg
        self.context = None

    async def begin(self):
        """Optional pre-flight initialization."""
        pass

    async def login_by_qrcode(self):
        """Twitter does not support QR-code login."""
        raise NotImplementedError("Twitter does not support QR-code login")

    async def login_by_mobile(self):
        """Twitter mobile/SMS login not implemented."""
        raise NotImplementedError("Twitter mobile login not implemented")

    async def login_by_cookies(self, context: BrowserContext):
        """
        Authenticate using cookies from config or saved session.
        Called automatically by the crawler after browser launch.
        """
        if self.cfg.cookies:
            await context.add_cookies(self.cfg.cookies)
            await context.storage_state(path=self.cfg.cookie_save_path)

```

### 4. Create the Store Implementation

Subclass `AbstractStore` to persist scraped data. Implement optional `AbstractStoreImage` and `AbstractStoreVideo` if your platform returns media files.

```python

# store/twitter/_store_impl.py

import json
import asyncio
from pathlib import Path

from base.base_crawler import AbstractStore, AbstractStoreImage


class TwitterStore(AbstractStore, AbstractStoreImage):
    def __init__(self, cfg):
        self.cfg = cfg
        self.output_dir = Path(cfg.output_path)
        self.output_dir.mkdir(parents=True, exist_ok=True)
        
        self.content_file = open(
            self.output_dir / "tweets.jsonl", 
            "w", 
            encoding="utf-8"
        )
        self._lock = asyncio.Lock()

    async def store_content(self, tweet: dict):
        """Persist a tweet to JSON Lines format."""
        async with self._lock:
            self.content_file.write(json.dumps(tweet, ensure_ascii=False) + "\n")

    async def store_comment(self, reply: dict):
        """Twitter replies are stored as content with type annotation."""
        reply["_type"] = "reply"
        await self.store_content(reply)

    async def store_creator(self, user: dict):
        """Store user profile metadata."""
        user["_type"] = "creator"
        await self.store_content(user)

    async def store_image(self, image_url: str, tweet_id: str):
        """Optional: download and store attached images."""
        # Implementation using aiohttp or similar

        pass

    def __del__(self):
        """Ensure file handle closure."""
        self.content_file.close()

```

### 5. Add Platform Configuration

Create a dataclass that inherits from `BaseConfig` with platform-specific fields.

```python

# config/twitter_config.py

from dataclasses import dataclass, field
from base.base_crawler import BaseConfig


@dataclass
class TwitterConfig(BaseConfig):
    """Twitter-specific configuration."""
    
    # Inherited from BaseConfig: proxy, user_agent, headless, etc.

    
    # Platform-specific fields

    api_key: str = ""
    api_secret: str = ""
    bearer_token: str = ""
    rate_limit_per_hour: int = 300
    
    # Cookie storage

    cookie_save_path: str = "data/twitter_cookies.json"
    cookies: list[dict] = field(default_factory=list)
    
    # Output configuration

    output_path: str = "data/twitter"
    download_media: bool = True

```

### 6. Register in the CLI Parser

Add your platform to the choices in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py):

```python

# cmd_arg/arg.py

import argparse


def get_arg_parser():
    parser = argparse.ArgumentParser(description="MediaCrawler")
    
    parser.add_argument(
        "--platform",
        choices=["xhs", "dy", "ks", "zhihu", "bilibili", "weibo", "twitter"],
        required=True,
        help="Target social media platform to crawl"
    )
    
    parser.add_argument(
        "--type",
        choices=["search", "detail", "creator"],
        default="search",
        help="Crawling mode"
    )
    
    parser.add_argument(
        "--keywords",
        nargs="+",
        help="Search keywords"
    )
    
    # ... additional arguments ...

    
    return parser

```

### 7. Register in the Crawler Manager

Import and instantiate your crawler in [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py):

```python

# api/services/crawler_manager.py

from config.twitter_config import TwitterConfig
from media_platform.twitter.client import TwitterCrawler

# ... existing imports ...


class CrawlerManager:
    def __init__(self):
        self._crawlers = {}

    def get_crawler(self, platform: str, config_override: dict = None):
        """Factory method returning appropriate crawler instance."""
        
        if platform == "twitter":
            cfg = TwitterConfig(**config_override) if config_override else TwitterConfig()
            return TwitterCrawler(cfg)
        
        elif platform == "zhihu":
            from media_platform.zhihu.client import ZhihuCrawler
            from config.zhihu_config import ZhihuConfig
            cfg = ZhihuConfig(**config_override) if config_override else ZhihuConfig()
            return ZhihuCrawler(cfg)
        
        # ... existing platform branches ...

        
        else:
            raise ValueError(f"Unsupported platform: {platform}")

    async def run(self, platform: str, **kwargs):
        """Execute crawling workflow for specified platform."""
        crawler = self.get_crawler(platform, kwargs.get("config"))
        await crawler.start()
        
        if kwargs.get("type") == "search":
            await crawler.search(
                keyword=kwargs.get("keywords", [""])[0],
                limit=kwargs.get("limit", 100)
            )

```

### 8. Add Optional Constants

For static mappings (URL patterns, API endpoints, selectors), create a constants file:

```python

# constant/twitter.py

"""Twitter-specific constants and selectors."""

BASE_URL = "https://twitter.com"

API_ENDPOINTS = {
    "search": "/search",
    "user_profile": "/i/api/graphql/UserByScreenName",
    "tweet_detail": "/i/api/graphql/TweetDetail",
}

SELECTORS = {
    "tweet_article": "article[data-testid='tweet']",
    "tweet_text": "[data-testid='tweetText']",
    "tweet_time": "time",
    "user_handle": "[data-testid='User-Names'] a[role='link']",
}

HEADERS = {
    "authorization": "Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA",
    "x-twitter-active-user": "yes",
    "x-twitter-client-language": "en",
}

```

## Factory Pattern and Loose Coupling

MediaCrawler uses **factory patterns** in two key locations to maintain loose coupling:

- **[`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py)**: Provides `CacheFactory.get_cache()` for platform-agnostic caching
- **[`store/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/__init__.py)**: Exposes `StoreFactory.get_store()` abstracted from concrete implementations

This design means **existing core code never changes** when adding a new platform. Only new files are created, and two registration points are updated (CLI parser and crawler manager).

## Testing Your Implementation

Follow the pattern in `tests/` for validation:

```python

# tests/test_twitter_crawler.py

import pytest
from media_platform.twitter.client import TwitterCrawler
from config.twitter_config import TwitterConfig


@pytest.fixture
def mock_config():
    return TwitterConfig(
        headless=True,
        output_path="/tmp/test_output"
    )


@pytest.mark.asyncio
async def test_twitter_crawler_init(mock_config):
    crawler = TwitterCrawler(mock_config)
    assert crawler.cfg == mock_config
    assert crawler.login is not None
    assert crawler.store is not None


@pytest.mark.asyncio
async def test_login_by_cookies_not_implemented(mock_config):
    from media_platform.twitter.login import TwitterLogin
    
    login = TwitterLogin(mock_config)
    
    with pytest.raises(NotImplementedError):
        await login.login_by_qrcode()
    
    with pytest.raises(NotImplementedError):
        await login.login_by_mobile()

```

## Key Reference Files

| Purpose | Path | Description |
|---------|------|-------------|
| Base abstractions | [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) | `AbstractCrawler`, `AbstractLogin`, `AbstractStore` definitions |
| Reference client | [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py) | Working implementation of all three required methods |
| Reference login | [`media_platform/zhihu/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/login.py) | QR, mobile, and cookie login flows |
| Reference store | [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) | CSV/JSON/DB persistence patterns |
| Reference config | [`config/zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py) | Dataclass with platform-specific fields |
| Main entry | [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) | Orchestrates platform selection and execution |
| CLI arguments | [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) | Argument parser with platform choices |
| Crawler manager | [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py) | Factory dispatch logic |

## Summary

- **MediaCrawler extension** requires implementing four abstract classes: `AbstractCrawler`, `AbstractLogin`, `AbstractStore`, and optionally `AbstractApiClient`.

- **File organization** follows the pattern `media_platform/<platform>/` for client code, `store/<platform>/` for persistence, and `config/<platform>_config.py` for settings.

- **Registration** requires two edits: add the platform name to [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py) choices, and add a factory branch in [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py).

- **No core modifications** are needed—the architecture is designed for additive-only extensions through loose coupling and factory patterns.

- **Reference implementations** in `media_platform/zhihu/` provide working templates for browser automation, login handling, and data storage.

## Frequently Asked Questions

### What programming patterns does MediaCrawler use for platform extensibility?

MediaCrawler uses the **Abstract Factory** and **Template Method** patterns. Abstract base classes in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) define the template structure, while concrete implementations in `media_platform/<platform>/` provide platform-specific behavior. Factory classes (`CrawlerManager`, `CacheFactory`) handle runtime instantiation without hardcoding class names.

### Can I implement only one login method if my platform supports only cookies?

Yes. The `AbstractLogin` interface requires all three methods, but you can raise `NotImplementedError` for unsupported flows. As shown in the Twitter example, `login_by_qrcode()` and `login_by_mobile()` raise exceptions while `login_by_cookies()` performs actual authentication. The main workflow in `AbstractCrawler.start()` catches and handles these exceptions gracefully.

### How does MediaCrawler handle browser instances across platforms?

The `AbstractCrawler.launch_browser()` method in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) provides a standard Playwright initialization that subclasses can call via `super().launch_browser()`. Each platform client stores the returned `BrowserContext` and passes it to login and scraping methods. This ensures consistent proxy handling, user-agent rotation, and headless configuration across all platforms.

### What storage formats does MediaCrawler support out of the box?

The `AbstractStore` interface is format-agnostic. Existing implementations in [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) demonstrate JSON Lines, CSV, and SQLite persistence. You can implement any storage backend—Elasticsearch, PostgreSQL, S3—by overriding `store_content()`, `store_comment()`, and `store_creator()`. The optional `AbstractStoreImage` and `AbstractStoreVideo` mixins add media download capabilities.