MediaCrawler Data Deduplication Strategies for Database Storage Mode
MediaCrawler uses upsert-based deduplication and unique MongoDB indexes to prevent duplicate records when storing crawled media content.
When scraping media at scale, storing duplicates wastes storage and skews analytics. MediaCrawler addresses this through two layered mechanisms that work together in database storage mode. This article explains exactly how the framework prevents duplicate records based on its MongoDB-based architecture.
Upsert-Based Deduplication in MongoDBStoreBase
The primary deduplication strategy centers on upsert operations in database/mongodb_store_base.py. The save_or_update method implements this pattern:
# From mongodb_store_base.py lines 5-12 (conceptual structure)
async def save_or_update(
self,
collection_suffix: str,
query: Dict,
data: Dict
) -> None:
collection = self._get_collection(collection_suffix)
await collection.update_one(
query, # natural key lookup
{"$set": data}, # update existing or set new
upsert=True # insert if not exists
)
How it works: The method accepts a query dictionary that defines the natural key (e.g., {"post_id": "12345"}). MongoDB's update_one with upsert=True either updates the matching document or inserts a new one if no match exists. This atomic operation guarantees single-record semantics without explicit existence checks.
This pattern appears across all platform-specific stores. For Zhihu content, the store passes {"post_id": post_id} as the query; for Douyin, it might use {"aweme_id": aweme_id}.
Unique Index Enforcement at the Database Layer
The second safeguard involves unique indexes created during store initialization. Each platform store defines which fields constitute uniqueness for its content types.
In store/zhihu/_store_impl.py (lines 45-48), the Zhihu store creates indexes like:
# Unique index creation example from _store_impl.py
await self.create_index(
collection_suffix="article",
keys=[("post_id", 1)], # ascending index on natural key
unique=True # reject duplicate values
)
Why both mechanisms? The upsert handles the common case efficiently at the application layer. The unique index acts as a failsafe against race conditions, manual inserts, or bugs that might bypass the upsert logic. If a duplicate insertion somehow occurs, MongoDB raises DuplicateKeyError, which the store catches and logs rather than corrupting the dataset.
Practical Implementation Example
Here's the complete workflow for storing Zhihu articles without duplicates:
from store.zhihu._store_impl import ZhihuStoreImpl
store = ZhihuStoreImpl()
article = {
"post_id": "67890",
"title": "Reinforcement Learning Advances",
"content": "...extracted text...",
"crawl_time": "2024-01-15T10:30:00Z"
}
# Natural key query prevents duplicates
query = {"post_id": article["post_id"]}
# Upsert: updates if exists, inserts if new
await store.save_or_update(
collection_suffix="article",
query=query,
data=article
)
For initial setup, ensure unique indexes exist:
# Run once per collection during application startup
await store.create_index(
collection_suffix="article",
keys=[("post_id", 1)],
unique=True
)
Deduplication for Alternative Storage Backends
MediaCrawler also provides store/excel_store_base.py for Excel export scenarios. This implementation uses an in-memory set to track seen natural keys during the export session, filtering duplicates without database constraints. This demonstrates the framework's strategy pattern—adapt deduplication mechanics to the storage backend's capabilities.
Key Files and Their Roles
| File | Responsibility |
|---|---|
database/mongodb_store_base.py |
Core save_or_update upsert logic, create_index helper |
store/zhihu/_store_impl.py |
Zhihu-specific unique indexes on post_id |
store/douyin/_store_impl.py |
Douyin store with parallel dedup strategy |
config/db_config.py |
MongoDB connection parameters |
store/excel_store_base.py |
Non-database dedup via in-memory tracking |
Summary
- Upsert operations in
MongoDBStoreBase.save_or_updateprovide application-level deduplication using natural key queries - Unique indexes created via
create_index(unique=True)enforce constraints at the MongoDB level - Platform-specific stores (Zhihu, Douyin) define which fields constitute uniqueness for their content types
- Alternative backends like Excel use appropriate dedup strategies for their limitations
Frequently Asked Questions
How does MediaCrawler handle the same content being crawled multiple times?
MediaCrawler's upsert pattern automatically refreshes existing records. When save_or_update receives content with a matching post_id, the $set operator updates all fields rather than creating a duplicate document. This enables incremental crawling where newer data overwrites older versions.
What happens if two crawler instances try to insert the same post simultaneously?
The unique MongoDB index prevents the race condition. One instance's upsert succeeds; the other encounters a duplicate key error, which the store logs and handles gracefully. The atomicity of MongoDB's update_one with upsert=True minimizes this window.
Can I customize which fields determine uniqueness for my crawled data?
Yes. When implementing a custom store, override the index creation in your _store_impl.py file. Specify your natural key fields in create_index(keys=[...], unique=True), and ensure your save_or_update calls use matching query dictionaries.
Does deduplication work when exporting to Excel instead of MongoDB?
Partially. The Excel store uses runtime in-memory deduplication via a Python set tracking seen keys. This prevents duplicates within a single export session but does not persist across sessions or compare against historical data like MongoDB's indexes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →