# MediaCrawler Data Deduplication Strategies for Database Storage Mode

> Discover MediaCrawler's data deduplication strategies for database storage. Learn how upsert and unique MongoDB indexes prevent duplicate records effectively.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: best-practices
- Published: 2026-08-14

---

**MediaCrawler uses upsert-based deduplication and unique MongoDB indexes to prevent duplicate records when storing crawled media content.**

When scraping media at scale, storing duplicates wastes storage and skews analytics. MediaCrawler addresses this through two layered mechanisms that work together in database storage mode. This article explains exactly how the framework prevents duplicate records based on its MongoDB-based architecture.

## Upsert-Based Deduplication in MongoDBStoreBase

The primary deduplication strategy centers on **upsert operations** in [`database/mongodb_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/mongodb_store_base.py). The `save_or_update` method implements this pattern:

```python

# From mongodb_store_base.py lines 5-12 (conceptual structure)

async def save_or_update(
    self, 
    collection_suffix: str, 
    query: Dict, 
    data: Dict
) -> None:
    collection = self._get_collection(collection_suffix)
    await collection.update_one(
        query,           # natural key lookup

        {"$set": data},  # update existing or set new

        upsert=True      # insert if not exists

    )

```

**How it works:** The method accepts a `query` dictionary that defines the natural key (e.g., `{"post_id": "12345"}`). MongoDB's `update_one` with `upsert=True` either updates the matching document or inserts a new one if no match exists. This atomic operation guarantees single-record semantics without explicit existence checks.

This pattern appears across all platform-specific stores. For Zhihu content, the store passes `{"post_id": post_id}` as the query; for Douyin, it might use `{"aweme_id": aweme_id}`.

## Unique Index Enforcement at the Database Layer

The second safeguard involves **unique indexes** created during store initialization. Each platform store defines which fields constitute uniqueness for its content types.

In [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) (lines 45-48), the Zhihu store creates indexes like:

```python

# Unique index creation example from _store_impl.py

await self.create_index(
    collection_suffix="article",
    keys=[("post_id", 1)],  # ascending index on natural key

    unique=True             # reject duplicate values

)

```

**Why both mechanisms?** The upsert handles the common case efficiently at the application layer. The unique index acts as a failsafe against race conditions, manual inserts, or bugs that might bypass the upsert logic. If a duplicate insertion somehow occurs, MongoDB raises `DuplicateKeyError`, which the store catches and logs rather than corrupting the dataset.

## Practical Implementation Example

Here's the complete workflow for storing Zhihu articles without duplicates:

```python
from store.zhihu._store_impl import ZhihuStoreImpl

store = ZhihuStoreImpl()

article = {
    "post_id": "67890",
    "title": "Reinforcement Learning Advances",
    "content": "...extracted text...",
    "crawl_time": "2024-01-15T10:30:00Z"
}

# Natural key query prevents duplicates

query = {"post_id": article["post_id"]}

# Upsert: updates if exists, inserts if new

await store.save_or_update(
    collection_suffix="article",
    query=query,
    data=article
)

```

For initial setup, ensure unique indexes exist:

```python

# Run once per collection during application startup

await store.create_index(
    collection_suffix="article",
    keys=[("post_id", 1)],
    unique=True
)

```

## Deduplication for Alternative Storage Backends

MediaCrawler also provides [`store/excel_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/excel_store_base.py) for Excel export scenarios. This implementation uses an **in-memory set** to track seen natural keys during the export session, filtering duplicates without database constraints. This demonstrates the framework's strategy pattern—adapt deduplication mechanics to the storage backend's capabilities.

## Key Files and Their Roles

| File | Responsibility |
|------|---------------|
| [`database/mongodb_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/mongodb_store_base.py) | Core `save_or_update` upsert logic, `create_index` helper |
| [`store/zhihu/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/zhihu/_store_impl.py) | Zhihu-specific unique indexes on `post_id` |
| [`store/douyin/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/douyin/_store_impl.py) | Douyin store with parallel dedup strategy |
| [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) | MongoDB connection parameters |
| [`store/excel_store_base.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/excel_store_base.py) | Non-database dedup via in-memory tracking |

## Summary

- **Upsert operations** in `MongoDBStoreBase.save_or_update` provide application-level deduplication using natural key queries
- **Unique indexes** created via `create_index(unique=True)` enforce constraints at the MongoDB level
- **Platform-specific stores** (Zhihu, Douyin) define which fields constitute uniqueness for their content types
- **Alternative backends** like Excel use appropriate dedup strategies for their limitations

## Frequently Asked Questions

### How does MediaCrawler handle the same content being crawled multiple times?

MediaCrawler's upsert pattern automatically refreshes existing records. When `save_or_update` receives content with a matching `post_id`, the `$set` operator updates all fields rather than creating a duplicate document. This enables incremental crawling where newer data overwrites older versions.

### What happens if two crawler instances try to insert the same post simultaneously?

The unique MongoDB index prevents the race condition. One instance's upsert succeeds; the other encounters a duplicate key error, which the store logs and handles gracefully. The atomicity of MongoDB's `update_one` with `upsert=True` minimizes this window.

### Can I customize which fields determine uniqueness for my crawled data?

Yes. When implementing a custom store, override the index creation in your [`_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/_store_impl.py) file. Specify your natural key fields in `create_index(keys=[...], unique=True)`, and ensure your `save_or_update` calls use matching query dictionaries.

### Does deduplication work when exporting to Excel instead of MongoDB?

Partially. The Excel store uses runtime in-memory deduplication via a Python `set` tracking seen keys. This prevents duplicates within a single export session but does not persist across sessions or compare against historical data like MongoDB's indexes.