MediaCrawler Data Deduplication Strategies for Database Storage Mode

MediaCrawler uses upsert-based deduplication and unique MongoDB indexes to prevent duplicate records when storing crawled media content.

When scraping media at scale, storing duplicates wastes storage and skews analytics. MediaCrawler addresses this through two layered mechanisms that work together in database storage mode. This article explains exactly how the framework prevents duplicate records based on its MongoDB-based architecture.

Upsert-Based Deduplication in MongoDBStoreBase

The primary deduplication strategy centers on upsert operations in database/mongodb_store_base.py. The save_or_update method implements this pattern:


# From mongodb_store_base.py lines 5-12 (conceptual structure)

async def save_or_update(
    self, 
    collection_suffix: str, 
    query: Dict, 
    data: Dict
) -> None:
    collection = self._get_collection(collection_suffix)
    await collection.update_one(
        query,           # natural key lookup

        {"$set": data},  # update existing or set new

        upsert=True      # insert if not exists

    )

How it works: The method accepts a query dictionary that defines the natural key (e.g., {"post_id": "12345"}). MongoDB's update_one with upsert=True either updates the matching document or inserts a new one if no match exists. This atomic operation guarantees single-record semantics without explicit existence checks.

This pattern appears across all platform-specific stores. For Zhihu content, the store passes {"post_id": post_id} as the query; for Douyin, it might use {"aweme_id": aweme_id}.

Unique Index Enforcement at the Database Layer

The second safeguard involves unique indexes created during store initialization. Each platform store defines which fields constitute uniqueness for its content types.

In store/zhihu/_store_impl.py (lines 45-48), the Zhihu store creates indexes like:


# Unique index creation example from _store_impl.py

await self.create_index(
    collection_suffix="article",
    keys=[("post_id", 1)],  # ascending index on natural key

    unique=True             # reject duplicate values

)

Why both mechanisms? The upsert handles the common case efficiently at the application layer. The unique index acts as a failsafe against race conditions, manual inserts, or bugs that might bypass the upsert logic. If a duplicate insertion somehow occurs, MongoDB raises DuplicateKeyError, which the store catches and logs rather than corrupting the dataset.

Practical Implementation Example

Here's the complete workflow for storing Zhihu articles without duplicates:

from store.zhihu._store_impl import ZhihuStoreImpl

store = ZhihuStoreImpl()

article = {
    "post_id": "67890",
    "title": "Reinforcement Learning Advances",
    "content": "...extracted text...",
    "crawl_time": "2024-01-15T10:30:00Z"
}

# Natural key query prevents duplicates

query = {"post_id": article["post_id"]}

# Upsert: updates if exists, inserts if new

await store.save_or_update(
    collection_suffix="article",
    query=query,
    data=article
)

For initial setup, ensure unique indexes exist:


# Run once per collection during application startup

await store.create_index(
    collection_suffix="article",
    keys=[("post_id", 1)],
    unique=True
)

Deduplication for Alternative Storage Backends

MediaCrawler also provides store/excel_store_base.py for Excel export scenarios. This implementation uses an in-memory set to track seen natural keys during the export session, filtering duplicates without database constraints. This demonstrates the framework's strategy pattern—adapt deduplication mechanics to the storage backend's capabilities.

Key Files and Their Roles

File Responsibility
database/mongodb_store_base.py Core save_or_update upsert logic, create_index helper
store/zhihu/_store_impl.py Zhihu-specific unique indexes on post_id
store/douyin/_store_impl.py Douyin store with parallel dedup strategy
config/db_config.py MongoDB connection parameters
store/excel_store_base.py Non-database dedup via in-memory tracking

Summary

  • Upsert operations in MongoDBStoreBase.save_or_update provide application-level deduplication using natural key queries
  • Unique indexes created via create_index(unique=True) enforce constraints at the MongoDB level
  • Platform-specific stores (Zhihu, Douyin) define which fields constitute uniqueness for their content types
  • Alternative backends like Excel use appropriate dedup strategies for their limitations

Frequently Asked Questions

How does MediaCrawler handle the same content being crawled multiple times?

MediaCrawler's upsert pattern automatically refreshes existing records. When save_or_update receives content with a matching post_id, the $set operator updates all fields rather than creating a duplicate document. This enables incremental crawling where newer data overwrites older versions.

What happens if two crawler instances try to insert the same post simultaneously?

The unique MongoDB index prevents the race condition. One instance's upsert succeeds; the other encounters a duplicate key error, which the store logs and handles gracefully. The atomicity of MongoDB's update_one with upsert=True minimizes this window.

Can I customize which fields determine uniqueness for my crawled data?

Yes. When implementing a custom store, override the index creation in your _store_impl.py file. Specify your natural key fields in create_index(keys=[...], unique=True), and ensure your save_or_update calls use matching query dictionaries.

Does deduplication work when exporting to Excel instead of MongoDB?

Partially. The Excel store uses runtime in-memory deduplication via a Python set tracking seen keys. This prevents duplicates within a single export session but does not persist across sessions or compare against historical data like MongoDB's indexes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →