How NanmiCoder/MediaCrawler Ensures the Integrity of Scraped Data from XHS
NanmiCoder/MediaCrawler ensures the integrity of scraped data from XHS by implementing deterministic user ID hashing, schema-validated atomic persistence, idempotent up-serts, and thread-safe async writes that prevent duplication and data loss.
The MediaCrawler project is a comprehensive open-source scraping framework designed for Chinese social media platforms, including Xiaohongshu (xhs). To ensure the integrity of scraped data from xhs, the codebase employs a layered defense strategy that combines cryptographic anonymization, strict database constraints, and atomic storage operations across multiple backends.
Deterministic Hashing of User Identifiers
Raw user identifiers never persist in storage. The [tools/user_hash.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py) module provides anonymize_user_id, which generates a stable SHA-256 hash of the original user_id. This creates a consistent creator_hash that links content to authors without exposing personally identifiable information.
The [store/xhs/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) file imports these helpers and applies them through high-level functions like update_xhs_note. Additionally, mask_nickname sanitizes human-readable names by replacing sensitive characters, ensuring privacy compliance while maintaining debugging capability.
Schema-Validated Atomic Persistence
Every note and comment passes through strict validation before reaching storage. The extraction logic in [media_platform/xhs/extractor.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) prepares raw API responses, which are then validated against Pydantic-like ORM models (XhsNote, XhsNoteComment) defined in the database layer.
The [store/xhs/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) implementations—XhsDbStoreImplement, XhsMongoStoreImplement, and XhsCsvStoreImplement—verify that required fields like note_id and comment_id exist and are correctly typed. The content_is_exist and comment_is_exist methods query existing records before writing, preventing null value corruption and enforcing referential integrity.
Idempotent Up-Serts for Data Consistency
To prevent duplicate entries during incremental crawls, MediaCrawler implements idempotent up-serts. In [store/xhs/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py), the store_content method checks for existing records:
- If the
note_idexists, it callsupdate_contentto overwrite mutable fields likelike_countandlast_modify_tswhile preserving the originalcreator_hash. - If absent, it calls
add_contentto insert the new record.
The MongoDB implementation delegates to MongoDBStoreBase.save_or_update, performing atomic upserts that guarantee documents are never partially written or duplicated.
Unified Timestamping and Audit Trails
Every write operation includes precise temporal tracking using get_current_timestamp from the time utilities. The [store/xhs/_store_impl.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) implementations set add_ts at creation and update last_modify_ts on every subsequent change.
This dual-timestamp strategy enables detection of stale records and out-of-order updates during downstream analysis, creating a complete audit trail for each scraped item.
Thread-Safe Async File Writing
For CSV, JSON, and JSONL outputs, the [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) provides AsyncFileWriter, which serializes concurrent writes to a single file per item type. This prevents race conditions that could corrupt file integrity when multiple coroutines process xhs data simultaneously.
Automated Testing Safeguards
The integrity pipeline is enforced by automated tests in [tests/test_no_user_info.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py). These tests explicitly call helper functions like _check_no_forbidden_keys and _check_nickname_masked to assert that no raw user_id values appear in persisted payloads and that nicknames are properly redacted. The CI pipeline runs these checks on every commit, catching regressions automatically.
Implementation Examples
The following patterns demonstrate the core integrity mechanisms in action:
# Hash user identifiers before storage
from tools.user_hash import anonymize_user_id, mask_nickname
creator_hash = anonymize_user_id(user_info["user_id"])
safe_nickname = mask_nickname(user_info["nickname"])
payload = {
"creator_hash": creator_hash,
"nickname": safe_nickname,
"note_id": note_data["id"],
# additional fields...
}
# Store note with idempotent upsert (SQL backend)
from store.xhs._store_impl import XhsDbStoreImplement
store = XhsDbStoreImplement()
await store.store_content(payload) # Automatically handles update vs insert
# Atomic MongoDB upsert
from store.xhs._store_impl import XhsMongoStoreImplement
mongo_store = XhsMongoStoreImplement()
await mongo_store.store_comment(comment_payload) # Uses save_or_update internally
# Verify integrity in tests
def test_no_raw_user_id_persisted():
# After extraction and storage via update_xhs_note
persisted = get_stored_note(test_note_id)
assert "user_id" not in persisted
assert "creator_hash" in persisted
assert persisted["nickname"] != original_nickname
Summary
- Cryptographic anonymization via
anonymize_user_idensures raw xhs user IDs never leak into storage. - Schema validation through ORM models (
XhsNote,XhsNoteComment) rejects malformed data before persistence. - Idempotent up-serts in
store_contentandstore_commentprevent duplication while preserving immutable hashes. - Existence checks using
content_is_existandcomment_is_existenforce referential integrity before writes. - Atomic MongoDB operations via
save_or_updateguarantee complete document writes without partial corruption. - Thread-safe async writing through
AsyncFileWriterprevents file corruption during concurrent crawls. - Automated testing in
test_no_user_info.pyenforces privacy and integrity constraints at the CI level.
Frequently Asked Questions
How does MediaCrawler prevent duplicate xhs notes during repeated crawls?
MediaCrawler checks for existing records using content_is_exist before writing. If the note_id already exists, it executes update_content to refresh mutable fields like engagement counts while keeping the original creator_hash and add_ts intact. This up-sert pattern ensures idempotency across incremental scraping runs.
Where does the actual hashing of xhs user IDs occur?
The hashing logic resides in [tools/user_hash.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/user_hash.py), which exports anonymize_user_id. The [store/xhs/__init__.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) module imports this function and applies it to all creator payloads during the extraction-to-storage pipeline, ensuring no raw identifiers reach the persistence layer.
What ensures that CSV files don't get corrupted when scraping xhs data concurrently?
The [tools/async_file_writer.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py) implements AsyncFileWriter, which serializes write operations to a single file handle. By using async locks and ordered queues, it prevents race conditions that could interleave bytes or corrupt JSON structures when multiple xhs posts are written simultaneously.
How does the system verify that raw user IDs are never stored?
The test suite in [tests/test_no_user_info.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/tests/test_no_user_info.py) explicitly validates stored payloads. It scans for forbidden keys like user_id and verifies that creator_hash is present instead. These tests run automatically in CI, ensuring that any code change that accidentally exposes raw identifiers fails the build before deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →