How to Customize MediaCrawler's Scraping Logic: A Developer's Guide to Extending the Crawler
You customize MediaCrawler by subclassing AbstractCrawler in base/base_crawler.py, overriding platform-specific methods like search(), and plugging in custom extractors or client logic, then registering your new class in main/main.py via the CrawlerFactory.
MediaCrawler is an open-source multi-platform content crawler that uses a modular, plugin-based architecture to separate generic crawling workflows from platform-specific implementations. Whether you need to add new data fields, modify search filters, or integrate an entirely new social media platform, the codebase provides clear extension points through its abstract base classes and factory pattern. This guide walks through the concrete steps to customize the scraping logic according to the NanmiCoder/MediaCrawler source code.
Understand the Plugin Architecture
Before modifying behavior, you need to understand how the components interact. The architecture separates concerns into distinct layers:
base/base_crawler.py– ContainsAbstractCrawler, the async interface definingstart(),search(), andlaunch_browser(). Every platform crawler must implement these methods.media_platform/{platform}/core.py– Houses concrete crawlers likeZhihuCrawlerorTieBaCrawlerthat inherit fromAbstractCrawlerand implement platform-specific login, pagination, and extraction workflows.media_platform/{platform}/client.py– Client classes (e.g.,ZhiHuClient,BaiduTieBaClient) wrap HTTP requests or Playwright browser automation, handling cookies, headers, and anti-detection measures.media_platform/{platform}/help.py– Extractor classes parse raw HTML/JSON into structured models (e.g.,ZhihuContent,TiebaNote).store/{platform}/_store_impl.py– Persistence layer using MongoDB, Excel, or JSON.config/{platform}_config.py– Centralized configuration for keywords, concurrency limits, and sleep intervals.
Key Customization Points
Modify Search Behavior and Filters
The search() method in concrete crawlers controls what content is fetched. In media_platform/zhihu/core.py, the ZhihuCrawler.search() method iterates over config.KEYWORDS and calls self.zhihu_client.get_note_by_keyword().
To customize:
- Subclass the existing crawler (e.g.,
class MyZhihuCrawler(ZhihuCrawler)) - Override
async def search(self)to inject pre-filters, change API endpoints, or add post-processing logic:
from media_platform.zhihu.core import ZhihuCrawler
from tools import utils
import config
class FilteredZhihuCrawler(ZhihuCrawler):
async def search(self) -> None:
utils.logger.info("[FilteredZhihuCrawler] Starting custom search")
for keyword in config.KEYWORDS.split(","):
# Skip low-priority keywords
if "spam" in keyword.lower():
continue
# Use parent method but with custom parameters
notes = await self.zhihu_client.get_note_by_keyword(keyword=keyword)
for note in notes:
if note.get("score", 0) > 50: # Custom quality filter
await self.zhihu_store.update_zhihu_content(note)
Adjust Pagination and Concurrency
Control flow limits by modifying constants in the platform-specific config files. In config/zhihu_config.py, locate:
CRAWLER_MAX_NOTES_COUNT– Maximum items to fetch per runCRAWLER_MAX_SLEEP_SEC– Delay between requests (anti-ban protection)MAX_CONCURRENCY_NUM– Async semaphore limit for concurrent connections
Change these at runtime before initializing the crawler:
import config.zhihu_config as zhihu_config
zhihu_config.CRAWLER_MAX_SLEEP_SEC = 1.5 # Faster crawling
zhihu_config.CRAWLER_MAX_NOTES_COUNT = 500 # Deeper scrape
Implement Custom Data Extractors
Extractors in media_platform/{platform}/help.py transform raw API responses into model objects. To capture additional fields:
- Create a subclass of the platform's extractor
- Override the
extract()method to parse extra JSON keys - Assign it in your crawler's
__init__
from media_platform.zhihu.help import ZhihuExtractor
from media_platform.zhihu.models import ZhihuContent
class RichMetadataExtractor(ZhihuExtractor):
def extract(self, raw_json: dict) -> ZhihuContent:
content = super().extract(raw_json)
# Extract custom field from API response
content.extra_metadata = raw_json.get("adJson", {}).get("creativeId")
return content
# In your custom crawler:
class EnhancedZhihuCrawler(ZhihuCrawler):
def __init__(self):
super().__init__()
self._extractor = RichMetadataExtractor() # Replace default extractor
Override Request Headers and Client Logic
The crawler instantiates clients with default headers in create_zhihu_client() (or equivalent). Override this method to inject custom user-agents, authentication tokens, or proxy configurations:
class StealthZhihuCrawler(ZhihuCrawler):
async def create_zhihu_client(self):
client = await super().create_zhihu_client()
# Add custom headers for API versioning
client.headers.update({
"X-API-Version": "v2",
"X-Custom-Auth": "Bearer token_here"
})
return client
Step-by-Step Customization Examples
Example 1: Adding a New Platform
To add support for a new site, create four components and register them:
Step 1: Create media_platform/newsite/core.py:
from base.base_crawler import AbstractCrawler
from tools import utils
class NewSiteCrawler(AbstractCrawler):
async def start(self):
await self.search()
async def search(self):
utils.logger.info("[NewSiteCrawler] Starting scrape")
# Implementation here
async def launch_browser(self, ...):
# Playwright setup
pass
Step 2: Create matching client in media_platform/newsite/client.py
Step 3: Register in main/main.py:
class CrawlerFactory:
CRAWLERS = {
"zhihu": ZhihuCrawler,
"tieba": TieBaCrawler,
"newsite": NewSiteCrawler, # Your new crawler
}
Example 2: Custom Store Implementation
To save data to a different database, extend the store layer in store/zhihu/_store_impl.py:
async def save_to_postgres(content):
import asyncpg
conn = await asyncpg.connect(DATABASE_URL)
await conn.execute("""
INSERT INTO zhihu_content (content_id, title, content, created_at)
VALUES ($1, $2, $3, $4)
ON CONFLICT (content_id) DO UPDATE SET
title = EXCLUDED.title,
content = EXCLUDED.content
""", content.content_id, content.title, content.content, content.created_at)
await conn.close()
Then call save_to_postgres() inside your crawler's search() method instead of the default MongoDB store.
Example 3: Command-Line Mode Switching
MediaCrawler supports different modes (search, detail, creator) through config.CRAWLER_TYPE. Set this dynamically based on CLI arguments handled in cmd_arg/arg.py, or modify your crawler to branch logic:
class MultiModeZhihuCrawler(ZhihuCrawler):
async def start(self):
if config.CRAWLER_TYPE == "creator":
await self.creator_scrape()
else:
await super().start()
async def creator_scrape(self):
# Custom logic for scraping specific creator profiles
pass
Registering Your Custom Crawler
After implementing your subclass, you must register it to make it available via the CLI. In main/main.py, locate the CrawlerFactory class and update the CRAWLERS dictionary:
from media_platform.zhihu.core import FilteredZhihuCrawler
class CrawlerFactory:
CRAWLERS = {
"zhihu": FilteredZhihuCrawler, # Replace default or add alongside
"tieba": TieBaCrawler,
# ...
}
Now running python main.py --platform zhihu executes your customized logic instead of the default implementation.
Summary
- Inherit from
AbstractCrawlerinbase/base_crawler.pyto create new platform support or modify existing behavior. - Override
search()in concrete crawlers (e.g.,media_platform/zhihu/core.py) to customize what content is fetched and how it is filtered. - Swap extractors by assigning a custom class to
self._extractorthat implements the sameextract(raw) -> modelsignature. - Modify
config/*.pyconstants likeCRAWLER_MAX_SLEEP_SECandCRAWLER_MAX_NOTES_COUNTto tune performance and limits without changing code logic. - Register custom classes in
main/main.py'sCrawlerFactory.CRAWLERSdictionary to activate them via command-line arguments. - Extend store implementations in
store/{platform}/_store_impl.pyto persist custom fields or change output formats.
Frequently Asked Questions
How do I change the delay between requests to avoid being blocked?
Adjust the CRAWLER_MAX_SLEEP_SEC value in the platform-specific config file (e.g., config/zhihu_config.py). This constant controls the await asyncio.sleep() duration used between paginated requests throughout the crawler.
Can I add new data fields without modifying the database schema?
Yes. Extend the extractor class in media_platform/{platform}/help.py to parse additional JSON keys, then update the corresponding model class. If using MongoDB (the default), new fields are accepted automatically without schema migration. For SQL stores, you must add columns to the table definition in your custom store implementation.
Where do I change the user-agent string for all requests?
Override the create_*_client() method in your crawler subclass (found in files like media_platform/zhihu/core.py). Return a client instance with modified headers, or modify tools/utils.py where the default user-agent is generated, though per-crawler overrides provide better isolation.
How do I scrape only specific content types (e.g., only videos or only articles)?
Implement a filter inside the overridden search() method. After fetching the raw data from the client, check the content type field (usually present in the API response JSON) and only call the store update methods for matching items. This keeps your storage clean and reduces processing overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →