How to Customize MediaCrawler's Scraping Logic: A Developer's Guide to Extending the Crawler

You customize MediaCrawler by subclassing AbstractCrawler in base/base_crawler.py, overriding platform-specific methods like search(), and plugging in custom extractors or client logic, then registering your new class in main/main.py via the CrawlerFactory.

MediaCrawler is an open-source multi-platform content crawler that uses a modular, plugin-based architecture to separate generic crawling workflows from platform-specific implementations. Whether you need to add new data fields, modify search filters, or integrate an entirely new social media platform, the codebase provides clear extension points through its abstract base classes and factory pattern. This guide walks through the concrete steps to customize the scraping logic according to the NanmiCoder/MediaCrawler source code.

Understand the Plugin Architecture

Before modifying behavior, you need to understand how the components interact. The architecture separates concerns into distinct layers:

  • base/base_crawler.py – Contains AbstractCrawler, the async interface defining start(), search(), and launch_browser(). Every platform crawler must implement these methods.
  • media_platform/{platform}/core.py – Houses concrete crawlers like ZhihuCrawler or TieBaCrawler that inherit from AbstractCrawler and implement platform-specific login, pagination, and extraction workflows.
  • media_platform/{platform}/client.py – Client classes (e.g., ZhiHuClient, BaiduTieBaClient) wrap HTTP requests or Playwright browser automation, handling cookies, headers, and anti-detection measures.
  • media_platform/{platform}/help.py – Extractor classes parse raw HTML/JSON into structured models (e.g., ZhihuContent, TiebaNote).
  • store/{platform}/_store_impl.py – Persistence layer using MongoDB, Excel, or JSON.
  • config/{platform}_config.py – Centralized configuration for keywords, concurrency limits, and sleep intervals.

Key Customization Points

Modify Search Behavior and Filters

The search() method in concrete crawlers controls what content is fetched. In media_platform/zhihu/core.py, the ZhihuCrawler.search() method iterates over config.KEYWORDS and calls self.zhihu_client.get_note_by_keyword().

To customize:

  1. Subclass the existing crawler (e.g., class MyZhihuCrawler(ZhihuCrawler))
  2. Override async def search(self) to inject pre-filters, change API endpoints, or add post-processing logic:
from media_platform.zhihu.core import ZhihuCrawler
from tools import utils
import config

class FilteredZhihuCrawler(ZhihuCrawler):
    async def search(self) -> None:
        utils.logger.info("[FilteredZhihuCrawler] Starting custom search")
        for keyword in config.KEYWORDS.split(","):
            # Skip low-priority keywords

            if "spam" in keyword.lower():
                continue
            # Use parent method but with custom parameters

            notes = await self.zhihu_client.get_note_by_keyword(keyword=keyword)
            for note in notes:
                if note.get("score", 0) > 50:  # Custom quality filter

                    await self.zhihu_store.update_zhihu_content(note)

Adjust Pagination and Concurrency

Control flow limits by modifying constants in the platform-specific config files. In config/zhihu_config.py, locate:

  • CRAWLER_MAX_NOTES_COUNT – Maximum items to fetch per run
  • CRAWLER_MAX_SLEEP_SEC – Delay between requests (anti-ban protection)
  • MAX_CONCURRENCY_NUM – Async semaphore limit for concurrent connections

Change these at runtime before initializing the crawler:

import config.zhihu_config as zhihu_config

zhihu_config.CRAWLER_MAX_SLEEP_SEC = 1.5  # Faster crawling

zhihu_config.CRAWLER_MAX_NOTES_COUNT = 500  # Deeper scrape

Implement Custom Data Extractors

Extractors in media_platform/{platform}/help.py transform raw API responses into model objects. To capture additional fields:

  1. Create a subclass of the platform's extractor
  2. Override the extract() method to parse extra JSON keys
  3. Assign it in your crawler's __init__
from media_platform.zhihu.help import ZhihuExtractor
from media_platform.zhihu.models import ZhihuContent

class RichMetadataExtractor(ZhihuExtractor):
    def extract(self, raw_json: dict) -> ZhihuContent:
        content = super().extract(raw_json)
        # Extract custom field from API response

        content.extra_metadata = raw_json.get("adJson", {}).get("creativeId")
        return content

# In your custom crawler:

class EnhancedZhihuCrawler(ZhihuCrawler):
    def __init__(self):
        super().__init__()
        self._extractor = RichMetadataExtractor()  # Replace default extractor

Override Request Headers and Client Logic

The crawler instantiates clients with default headers in create_zhihu_client() (or equivalent). Override this method to inject custom user-agents, authentication tokens, or proxy configurations:

class StealthZhihuCrawler(ZhihuCrawler):
    async def create_zhihu_client(self):
        client = await super().create_zhihu_client()
        # Add custom headers for API versioning

        client.headers.update({
            "X-API-Version": "v2",
            "X-Custom-Auth": "Bearer token_here"
        })
        return client

Step-by-Step Customization Examples

Example 1: Adding a New Platform

To add support for a new site, create four components and register them:

Step 1: Create media_platform/newsite/core.py:

from base.base_crawler import AbstractCrawler
from tools import utils

class NewSiteCrawler(AbstractCrawler):
    async def start(self):
        await self.search()
    
    async def search(self):
        utils.logger.info("[NewSiteCrawler] Starting scrape")
        # Implementation here

    
    async def launch_browser(self, ...):
        # Playwright setup

        pass

Step 2: Create matching client in media_platform/newsite/client.py

Step 3: Register in main/main.py:

class CrawlerFactory:
    CRAWLERS = {
        "zhihu": ZhihuCrawler,
        "tieba": TieBaCrawler,
        "newsite": NewSiteCrawler,  # Your new crawler

    }

Example 2: Custom Store Implementation

To save data to a different database, extend the store layer in store/zhihu/_store_impl.py:

async def save_to_postgres(content):
    import asyncpg
    conn = await asyncpg.connect(DATABASE_URL)
    await conn.execute("""
        INSERT INTO zhihu_content (content_id, title, content, created_at)
        VALUES ($1, $2, $3, $4)
        ON CONFLICT (content_id) DO UPDATE SET
            title = EXCLUDED.title,
            content = EXCLUDED.content
    """, content.content_id, content.title, content.content, content.created_at)
    await conn.close()

Then call save_to_postgres() inside your crawler's search() method instead of the default MongoDB store.

Example 3: Command-Line Mode Switching

MediaCrawler supports different modes (search, detail, creator) through config.CRAWLER_TYPE. Set this dynamically based on CLI arguments handled in cmd_arg/arg.py, or modify your crawler to branch logic:

class MultiModeZhihuCrawler(ZhihuCrawler):
    async def start(self):
        if config.CRAWLER_TYPE == "creator":
            await self.creator_scrape()
        else:
            await super().start()
    
    async def creator_scrape(self):
        # Custom logic for scraping specific creator profiles

        pass

Registering Your Custom Crawler

After implementing your subclass, you must register it to make it available via the CLI. In main/main.py, locate the CrawlerFactory class and update the CRAWLERS dictionary:

from media_platform.zhihu.core import FilteredZhihuCrawler

class CrawlerFactory:
    CRAWLERS = {
        "zhihu": FilteredZhihuCrawler,  # Replace default or add alongside

        "tieba": TieBaCrawler,
        # ...

    }

Now running python main.py --platform zhihu executes your customized logic instead of the default implementation.

Summary

  • Inherit from AbstractCrawler in base/base_crawler.py to create new platform support or modify existing behavior.
  • Override search() in concrete crawlers (e.g., media_platform/zhihu/core.py) to customize what content is fetched and how it is filtered.
  • Swap extractors by assigning a custom class to self._extractor that implements the same extract(raw) -> model signature.
  • Modify config/*.py constants like CRAWLER_MAX_SLEEP_SEC and CRAWLER_MAX_NOTES_COUNT to tune performance and limits without changing code logic.
  • Register custom classes in main/main.py's CrawlerFactory.CRAWLERS dictionary to activate them via command-line arguments.
  • Extend store implementations in store/{platform}/_store_impl.py to persist custom fields or change output formats.

Frequently Asked Questions

How do I change the delay between requests to avoid being blocked?

Adjust the CRAWLER_MAX_SLEEP_SEC value in the platform-specific config file (e.g., config/zhihu_config.py). This constant controls the await asyncio.sleep() duration used between paginated requests throughout the crawler.

Can I add new data fields without modifying the database schema?

Yes. Extend the extractor class in media_platform/{platform}/help.py to parse additional JSON keys, then update the corresponding model class. If using MongoDB (the default), new fields are accepted automatically without schema migration. For SQL stores, you must add columns to the table definition in your custom store implementation.

Where do I change the user-agent string for all requests?

Override the create_*_client() method in your crawler subclass (found in files like media_platform/zhihu/core.py). Return a client instance with modified headers, or modify tools/utils.py where the default user-agent is generated, though per-crawler overrides provide better isolation.

How do I scrape only specific content types (e.g., only videos or only articles)?

Implement a filter inside the overridden search() method. After fetching the raw data from the client, check the content type field (usually present in the API response JSON) and only call the store update methods for matching items. This keeps your storage clean and reduces processing overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →