# How NanmiCoder/MediaCrawler Ensures the Integrity of Scraped Data from XHS

> Learn how NanmiCoder MediaCrawler guarantees XHS scraped data integrity with hashing, validated persistence, idempotent up-serts, and thread-safe writes. Prevent duplicates effortlessly.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-03

---

**NanmiCoder/MediaCrawler ensures the integrity of scraped data from XHS by implementing deterministic user ID hashing, schema-validated atomic persistence, idempotent up-serts, and thread-safe async writes that prevent duplication and data loss.**

The MediaCrawler project is a comprehensive open-source scraping framework designed for Chinese social media platforms, including Xiaohongshu (xhs). To ensure the integrity of scraped data from xhs, the codebase employs a layered defense strategy that combines cryptographic anonymization, strict database constraints, and atomic storage operations across multiple backends.

## Deterministic Hashing of User Identifiers

Raw user identifiers never persist in storage. The [[`tools/user_hash.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py) module provides **anonymize_user_id**, which generates a stable SHA-256 hash of the original `user_id`. This creates a consistent `creator_hash` that links content to authors without exposing personally identifiable information.

The [[`store/xhs/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) file imports these helpers and applies them through high-level functions like `update_xhs_note`. Additionally, **mask_nickname** sanitizes human-readable names by replacing sensitive characters, ensuring privacy compliance while maintaining debugging capability.

## Schema-Validated Atomic Persistence

Every note and comment passes through strict validation before reaching storage. The extraction logic in [[`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) prepares raw API responses, which are then validated against Pydantic-like ORM models (`XhsNote`, `XhsNoteComment`) defined in the database layer.

The [[`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) implementations—`XhsDbStoreImplement`, `XhsMongoStoreImplement`, and `XhsCsvStoreImplement`—verify that required fields like `note_id` and `comment_id` exist and are correctly typed. The `content_is_exist` and `comment_is_exist` methods query existing records before writing, preventing null value corruption and enforcing referential integrity.

## Idempotent Up-Serts for Data Consistency

To prevent duplicate entries during incremental crawls, MediaCrawler implements **idempotent up-serts**. In [[`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py), the `store_content` method checks for existing records:

- If the `note_id` exists, it calls `update_content` to overwrite mutable fields like `like_count` and `last_modify_ts` while preserving the original `creator_hash`.
- If absent, it calls `add_content` to insert the new record.

The MongoDB implementation delegates to `MongoDBStoreBase.save_or_update`, performing atomic upserts that guarantee documents are never partially written or duplicated.

## Unified Timestamping and Audit Trails

Every write operation includes precise temporal tracking using `get_current_timestamp` from the time utilities. The [[`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) implementations set `add_ts` at creation and update `last_modify_ts` on every subsequent change.

This dual-timestamp strategy enables detection of stale records and out-of-order updates during downstream analysis, creating a complete audit trail for each scraped item.

## Thread-Safe Async File Writing

For CSV, JSON, and JSONL outputs, the [[`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) provides **AsyncFileWriter**, which serializes concurrent writes to a single file per item type. This prevents race conditions that could corrupt file integrity when multiple coroutines process xhs data simultaneously.

## Automated Testing Safeguards

The integrity pipeline is enforced by automated tests in [[`tests/test_no_user_info.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py). These tests explicitly call helper functions like `_check_no_forbidden_keys` and `_check_nickname_masked` to assert that no raw `user_id` values appear in persisted payloads and that nicknames are properly redacted. The CI pipeline runs these checks on every commit, catching regressions automatically.

## Implementation Examples

The following patterns demonstrate the core integrity mechanisms in action:

```python

# Hash user identifiers before storage

from tools.user_hash import anonymize_user_id, mask_nickname

creator_hash = anonymize_user_id(user_info["user_id"])
safe_nickname = mask_nickname(user_info["nickname"])

payload = {
    "creator_hash": creator_hash,
    "nickname": safe_nickname,
    "note_id": note_data["id"],
    # additional fields...

}

```

```python

# Store note with idempotent upsert (SQL backend)

from store.xhs._store_impl import XhsDbStoreImplement

store = XhsDbStoreImplement()
await store.store_content(payload)  # Automatically handles update vs insert

```

```python

# Atomic MongoDB upsert

from store.xhs._store_impl import XhsMongoStoreImplement

mongo_store = XhsMongoStoreImplement()
await mongo_store.store_comment(comment_payload)  # Uses save_or_update internally

```

```python

# Verify integrity in tests

def test_no_raw_user_id_persisted():
    # After extraction and storage via update_xhs_note

    persisted = get_stored_note(test_note_id)
    assert "user_id" not in persisted
    assert "creator_hash" in persisted
    assert persisted["nickname"] != original_nickname

```

## Summary

- **Cryptographic anonymization** via `anonymize_user_id` ensures raw xhs user IDs never leak into storage.
- **Schema validation** through ORM models (`XhsNote`, `XhsNoteComment`) rejects malformed data before persistence.
- **Idempotent up-serts** in `store_content` and `store_comment` prevent duplication while preserving immutable hashes.
- **Existence checks** using `content_is_exist` and `comment_is_exist` enforce referential integrity before writes.
- **Atomic MongoDB operations** via `save_or_update` guarantee complete document writes without partial corruption.
- **Thread-safe async writing** through `AsyncFileWriter` prevents file corruption during concurrent crawls.
- **Automated testing** in [`test_no_user_info.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/test_no_user_info.py) enforces privacy and integrity constraints at the CI level.

## Frequently Asked Questions

### How does MediaCrawler prevent duplicate xhs notes during repeated crawls?

MediaCrawler checks for existing records using `content_is_exist` before writing. If the `note_id` already exists, it executes `update_content` to refresh mutable fields like engagement counts while keeping the original `creator_hash` and `add_ts` intact. This up-sert pattern ensures idempotency across incremental scraping runs.

### Where does the actual hashing of xhs user IDs occur?

The hashing logic resides in [[`tools/user_hash.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py), which exports `anonymize_user_id`. The [[`store/xhs/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) module imports this function and applies it to all creator payloads during the extraction-to-storage pipeline, ensuring no raw identifiers reach the persistence layer.

### What ensures that CSV files don't get corrupted when scraping xhs data concurrently?

The [[`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) implements `AsyncFileWriter`, which serializes write operations to a single file handle. By using async locks and ordered queues, it prevents race conditions that could interleave bytes or corrupt JSON structures when multiple xhs posts are written simultaneously.

### How does the system verify that raw user IDs are never stored?

The test suite in [[`tests/test_no_user_info.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py) explicitly validates stored payloads. It scans for forbidden keys like `user_id` and verifies that `creator_hash` is present instead. These tests run automatically in CI, ensuring that any code change that accidentally exposes raw identifiers fails the build before deployment.