# How to Customize MediaCrawler's Scraping Logic: A Developer's Guide to Extending the Crawler

> Customize MediaCrawler scraping logic by subclassing AbstractCrawler overriding platform-specific methods and registering your custom class via CrawlerFactory. Extend MediaCrawler easily.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-29

---

**You customize MediaCrawler by subclassing `AbstractCrawler` in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py), overriding platform-specific methods like `search()`, and plugging in custom extractors or client logic, then registering your new class in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py) via the `CrawlerFactory`.**

MediaCrawler is an open-source multi-platform content crawler that uses a modular, plugin-based architecture to separate generic crawling workflows from platform-specific implementations. Whether you need to add new data fields, modify search filters, or integrate an entirely new social media platform, the codebase provides clear extension points through its abstract base classes and factory pattern. This guide walks through the concrete steps to customize the scraping logic according to the NanmiCoder/MediaCrawler source code.

## Understand the Plugin Architecture

Before modifying behavior, you need to understand how the components interact. The architecture separates concerns into distinct layers:

- **[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)** – Contains `AbstractCrawler`, the async interface defining `start()`, `search()`, and `launch_browser()`. Every platform crawler must implement these methods.
- **`media_platform/{platform}/core.py`** – Houses concrete crawlers like `ZhihuCrawler` or `TieBaCrawler` that inherit from `AbstractCrawler` and implement platform-specific login, pagination, and extraction workflows.
- **`media_platform/{platform}/client.py`** – Client classes (e.g., `ZhiHuClient`, `BaiduTieBaClient`) wrap HTTP requests or Playwright browser automation, handling cookies, headers, and anti-detection measures.
- **`media_platform/{platform}/help.py`** – Extractor classes parse raw HTML/JSON into structured models (e.g., `ZhihuContent`, `TiebaNote`).
- **`store/{platform}/_store_impl.py`** – Persistence layer using MongoDB, Excel, or JSON.
- **`config/{platform}_config.py`** – Centralized configuration for keywords, concurrency limits, and sleep intervals.

## Key Customization Points

### Modify Search Behavior and Filters

The `search()` method in concrete crawlers controls what content is fetched. In [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py), the `ZhihuCrawler.search()` method iterates over `config.KEYWORDS` and calls `self.zhihu_client.get_note_by_keyword()`.

To customize:
1. Subclass the existing crawler (e.g., `class MyZhihuCrawler(ZhihuCrawler)`)
2. Override `async def search(self)` to inject pre-filters, change API endpoints, or add post-processing logic:

```python
from media_platform.zhihu.core import ZhihuCrawler
from tools import utils
import config

class FilteredZhihuCrawler(ZhihuCrawler):
    async def search(self) -> None:
        utils.logger.info("[FilteredZhihuCrawler] Starting custom search")
        for keyword in config.KEYWORDS.split(","):
            # Skip low-priority keywords

            if "spam" in keyword.lower():
                continue
            # Use parent method but with custom parameters

            notes = await self.zhihu_client.get_note_by_keyword(keyword=keyword)
            for note in notes:
                if note.get("score", 0) > 50:  # Custom quality filter

                    await self.zhihu_store.update_zhihu_content(note)

```

### Adjust Pagination and Concurrency

Control flow limits by modifying constants in the platform-specific config files. In [`config/zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py), locate:

- `CRAWLER_MAX_NOTES_COUNT` – Maximum items to fetch per run
- `CRAWLER_MAX_SLEEP_SEC` – Delay between requests (anti-ban protection)
- `MAX_CONCURRENCY_NUM` – Async semaphore limit for concurrent connections

Change these at runtime before initializing the crawler:

```python
import config.zhihu_config as zhihu_config

zhihu_config.CRAWLER_MAX_SLEEP_SEC = 1.5  # Faster crawling

zhihu_config.CRAWLER_MAX_NOTES_COUNT = 500  # Deeper scrape

```

### Implement Custom Data Extractors

Extractors in `media_platform/{platform}/help.py` transform raw API responses into model objects. To capture additional fields:

1. Create a subclass of the platform's extractor
2. Override the `extract()` method to parse extra JSON keys
3. Assign it in your crawler's `__init__`

```python
from media_platform.zhihu.help import ZhihuExtractor
from media_platform.zhihu.models import ZhihuContent

class RichMetadataExtractor(ZhihuExtractor):
    def extract(self, raw_json: dict) -> ZhihuContent:
        content = super().extract(raw_json)
        # Extract custom field from API response

        content.extra_metadata = raw_json.get("adJson", {}).get("creativeId")
        return content

# In your custom crawler:

class EnhancedZhihuCrawler(ZhihuCrawler):
    def __init__(self):
        super().__init__()
        self._extractor = RichMetadataExtractor()  # Replace default extractor

```

### Override Request Headers and Client Logic

The crawler instantiates clients with default headers in `create_zhihu_client()` (or equivalent). Override this method to inject custom user-agents, authentication tokens, or proxy configurations:

```python
class StealthZhihuCrawler(ZhihuCrawler):
    async def create_zhihu_client(self):
        client = await super().create_zhihu_client()
        # Add custom headers for API versioning

        client.headers.update({
            "X-API-Version": "v2",
            "X-Custom-Auth": "Bearer token_here"
        })
        return client

```

## Step-by-Step Customization Examples

### Example 1: Adding a New Platform

To add support for a new site, create four components and register them:

**Step 1:** Create [`media_platform/newsite/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/newsite/core.py):

```python
from base.base_crawler import AbstractCrawler
from tools import utils

class NewSiteCrawler(AbstractCrawler):
    async def start(self):
        await self.search()
    
    async def search(self):
        utils.logger.info("[NewSiteCrawler] Starting scrape")
        # Implementation here

    
    async def launch_browser(self, ...):
        # Playwright setup

        pass

```

**Step 2:** Create matching client in [`media_platform/newsite/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/newsite/client.py)

**Step 3:** Register in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py):

```python
class CrawlerFactory:
    CRAWLERS = {
        "zhihu": ZhihuCrawler,
        "tieba": TieBaCrawler,
        "newsite": NewSiteCrawler,  # Your new crawler

    }

```

### Example 2: Custom Store Implementation

To save data to a different database, extend the store layer in [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py):

```python
async def save_to_postgres(content):
    import asyncpg
    conn = await asyncpg.connect(DATABASE_URL)
    await conn.execute("""
        INSERT INTO zhihu_content (content_id, title, content, created_at)
        VALUES ($1, $2, $3, $4)
        ON CONFLICT (content_id) DO UPDATE SET
            title = EXCLUDED.title,
            content = EXCLUDED.content
    """, content.content_id, content.title, content.content, content.created_at)
    await conn.close()

```

Then call `save_to_postgres()` inside your crawler's `search()` method instead of the default MongoDB store.

### Example 3: Command-Line Mode Switching

MediaCrawler supports different modes (search, detail, creator) through `config.CRAWLER_TYPE`. Set this dynamically based on CLI arguments handled in [`cmd_arg/arg.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cmd_arg/arg.py), or modify your crawler to branch logic:

```python
class MultiModeZhihuCrawler(ZhihuCrawler):
    async def start(self):
        if config.CRAWLER_TYPE == "creator":
            await self.creator_scrape()
        else:
            await super().start()
    
    async def creator_scrape(self):
        # Custom logic for scraping specific creator profiles

        pass

```

## Registering Your Custom Crawler

After implementing your subclass, you must register it to make it available via the CLI. In [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py), locate the `CrawlerFactory` class and update the `CRAWLERS` dictionary:

```python
from media_platform.zhihu.core import FilteredZhihuCrawler

class CrawlerFactory:
    CRAWLERS = {
        "zhihu": FilteredZhihuCrawler,  # Replace default or add alongside

        "tieba": TieBaCrawler,
        # ...

    }

```

Now running `python main.py --platform zhihu` executes your customized logic instead of the default implementation.

## Summary

- **Inherit from `AbstractCrawler`** in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) to create new platform support or modify existing behavior.
- **Override `search()`** in concrete crawlers (e.g., [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py)) to customize what content is fetched and how it is filtered.
- **Swap extractors** by assigning a custom class to `self._extractor` that implements the same `extract(raw) -> model` signature.
- **Modify `config/*.py` constants** like `CRAWLER_MAX_SLEEP_SEC` and `CRAWLER_MAX_NOTES_COUNT` to tune performance and limits without changing code logic.
- **Register custom classes** in [`main/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main/main.py)'s `CrawlerFactory.CRAWLERS` dictionary to activate them via command-line arguments.
- **Extend store implementations** in `store/{platform}/_store_impl.py` to persist custom fields or change output formats.

## Frequently Asked Questions

### How do I change the delay between requests to avoid being blocked?

Adjust the `CRAWLER_MAX_SLEEP_SEC` value in the platform-specific config file (e.g., [`config/zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/zhihu_config.py)). This constant controls the `await asyncio.sleep()` duration used between paginated requests throughout the crawler.

### Can I add new data fields without modifying the database schema?

Yes. Extend the extractor class in `media_platform/{platform}/help.py` to parse additional JSON keys, then update the corresponding model class. If using MongoDB (the default), new fields are accepted automatically without schema migration. For SQL stores, you must add columns to the table definition in your custom store implementation.

### Where do I change the user-agent string for all requests?

Override the `create_*_client()` method in your crawler subclass (found in files like [`media_platform/zhihu/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/core.py)). Return a client instance with modified headers, or modify [`tools/utils.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/utils.py) where the default user-agent is generated, though per-crawler overrides provide better isolation.

### How do I scrape only specific content types (e.g., only videos or only articles)?

Implement a filter inside the overridden `search()` method. After fetching the raw data from the client, check the content type field (usually present in the API response JSON) and only call the store update methods for matching items. This keeps your storage clean and reduces processing overhead.