How NanmiCoder/MediaCrawler Ensures the Integrity of Scraped Data from XHS

NanmiCoder/MediaCrawler ensures the integrity of scraped data from XHS by implementing deterministic user ID hashing, schema-validated atomic persistence, idempotent up-serts, and thread-safe async writes that prevent duplication and data loss.

The MediaCrawler project is a comprehensive open-source scraping framework designed for Chinese social media platforms, including Xiaohongshu (xhs). To ensure the integrity of scraped data from xhs, the codebase employs a layered defense strategy that combines cryptographic anonymization, strict database constraints, and atomic storage operations across multiple backends.

Deterministic Hashing of User Identifiers

Raw user identifiers never persist in storage. The [tools/user_hash.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py) module provides anonymize_user_id, which generates a stable SHA-256 hash of the original user_id. This creates a consistent creator_hash that links content to authors without exposing personally identifiable information.

The [store/xhs/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) file imports these helpers and applies them through high-level functions like update_xhs_note. Additionally, mask_nickname sanitizes human-readable names by replacing sensitive characters, ensuring privacy compliance while maintaining debugging capability.

Schema-Validated Atomic Persistence

Every note and comment passes through strict validation before reaching storage. The extraction logic in [media_platform/xhs/extractor.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) prepares raw API responses, which are then validated against Pydantic-like ORM models (XhsNote, XhsNoteComment) defined in the database layer.

The [store/xhs/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) implementations—XhsDbStoreImplement, XhsMongoStoreImplement, and XhsCsvStoreImplement—verify that required fields like note_id and comment_id exist and are correctly typed. The content_is_exist and comment_is_exist methods query existing records before writing, preventing null value corruption and enforcing referential integrity.

Idempotent Up-Serts for Data Consistency

To prevent duplicate entries during incremental crawls, MediaCrawler implements idempotent up-serts. In [store/xhs/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py), the store_content method checks for existing records:

  • If the note_id exists, it calls update_content to overwrite mutable fields like like_count and last_modify_ts while preserving the original creator_hash.
  • If absent, it calls add_content to insert the new record.

The MongoDB implementation delegates to MongoDBStoreBase.save_or_update, performing atomic upserts that guarantee documents are never partially written or duplicated.

Unified Timestamping and Audit Trails

Every write operation includes precise temporal tracking using get_current_timestamp from the time utilities. The [store/xhs/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) implementations set add_ts at creation and update last_modify_ts on every subsequent change.

This dual-timestamp strategy enables detection of stale records and out-of-order updates during downstream analysis, creating a complete audit trail for each scraped item.

Thread-Safe Async File Writing

For CSV, JSON, and JSONL outputs, the [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) provides AsyncFileWriter, which serializes concurrent writes to a single file per item type. This prevents race conditions that could corrupt file integrity when multiple coroutines process xhs data simultaneously.

Automated Testing Safeguards

The integrity pipeline is enforced by automated tests in [tests/test_no_user_info.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py). These tests explicitly call helper functions like _check_no_forbidden_keys and _check_nickname_masked to assert that no raw user_id values appear in persisted payloads and that nicknames are properly redacted. The CI pipeline runs these checks on every commit, catching regressions automatically.

Implementation Examples

The following patterns demonstrate the core integrity mechanisms in action:


# Hash user identifiers before storage

from tools.user_hash import anonymize_user_id, mask_nickname

creator_hash = anonymize_user_id(user_info["user_id"])
safe_nickname = mask_nickname(user_info["nickname"])

payload = {
    "creator_hash": creator_hash,
    "nickname": safe_nickname,
    "note_id": note_data["id"],
    # additional fields...

}

# Store note with idempotent upsert (SQL backend)

from store.xhs._store_impl import XhsDbStoreImplement

store = XhsDbStoreImplement()
await store.store_content(payload)  # Automatically handles update vs insert

# Atomic MongoDB upsert

from store.xhs._store_impl import XhsMongoStoreImplement

mongo_store = XhsMongoStoreImplement()
await mongo_store.store_comment(comment_payload)  # Uses save_or_update internally

# Verify integrity in tests

def test_no_raw_user_id_persisted():
    # After extraction and storage via update_xhs_note

    persisted = get_stored_note(test_note_id)
    assert "user_id" not in persisted
    assert "creator_hash" in persisted
    assert persisted["nickname"] != original_nickname

Summary

  • Cryptographic anonymization via anonymize_user_id ensures raw xhs user IDs never leak into storage.
  • Schema validation through ORM models (XhsNote, XhsNoteComment) rejects malformed data before persistence.
  • Idempotent up-serts in store_content and store_comment prevent duplication while preserving immutable hashes.
  • Existence checks using content_is_exist and comment_is_exist enforce referential integrity before writes.
  • Atomic MongoDB operations via save_or_update guarantee complete document writes without partial corruption.
  • Thread-safe async writing through AsyncFileWriter prevents file corruption during concurrent crawls.
  • Automated testing in test_no_user_info.py enforces privacy and integrity constraints at the CI level.

Frequently Asked Questions

How does MediaCrawler prevent duplicate xhs notes during repeated crawls?

MediaCrawler checks for existing records using content_is_exist before writing. If the note_id already exists, it executes update_content to refresh mutable fields like engagement counts while keeping the original creator_hash and add_ts intact. This up-sert pattern ensures idempotency across incremental scraping runs.

Where does the actual hashing of xhs user IDs occur?

The hashing logic resides in [tools/user_hash.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py), which exports anonymize_user_id. The [store/xhs/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) module imports this function and applies it to all creator payloads during the extraction-to-storage pipeline, ensuring no raw identifiers reach the persistence layer.

What ensures that CSV files don't get corrupted when scraping xhs data concurrently?

The [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) implements AsyncFileWriter, which serializes write operations to a single file handle. By using async locks and ordered queues, it prevents race conditions that could interleave bytes or corrupt JSON structures when multiple xhs posts are written simultaneously.

How does the system verify that raw user IDs are never stored?

The test suite in [tests/test_no_user_info.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py) explicitly validates stored payloads. It scans for forbidden keys like user_id and verifies that creator_hash is present instead. These tests run automatically in CI, ensuring that any code change that accidentally exposes raw identifiers fails the build before deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →