Maintaining the XHS Crawler as the Platform Evolves: A Technical Guide to Long-Term Stability
To maintain the XHS crawler as Xiaohongshu evolves, you must monitor API contract changes, update authentication tokens like xsec_token, adjust rate-limiting configurations, and synchronize data models across the storage layer while keeping the modular architecture intact.
The Xiaohongshu (XHS) crawler in MediaCrawler relies on a clean separation between data acquisition and persistence. As the platform introduces new API endpoints, authentication schemes, or rate-limiting policies, understanding how to update the storage contracts and configuration layers becomes critical for maintaining data integrity.
Monitoring API Contract Changes
When Xiaohongshu modifies request parameters or response formats, the crawler's normalization layer must map new JSON fields to the expected storage contracts. In store/xhs/_store_impl.py, the XhsMongoStoreImplement.store_content method expects specific keys like note_id, title, and desc (lines 64-72). Any platform-side renaming requires updating these field mappings in the store implementations.
Handling Authentication Token Drift
The platform may refresh how it generates anti-scraping tokens such as xsec_token. The storage layer currently persists these tokens alongside content data (lines 58-60 in _store_impl.py). When token generation logic changes, update both the request-building code and the storage schema to capture new authentication metadata.
Adapting Rate Limits and Crawler Politeness
Xiaohongshu frequently adjusts per-IP rate limits and anti-bot detection. The AsyncFileWriter class receives the current crawler type via crawler_type_var.get() (lines 45-46 in _store_impl.py), allowing you to switch between aggressive and conservative modes. Modify the global crawler_type configuration or introduce throttling middleware to respect new platform constraints.
Managing Data Model Evolution
New fields like share_count or deprecated arrays like tag_list require synchronized updates across persistence layers. The AbstractStore hierarchy in base/base_crawler.py defines the contract, but each implementation must handle the actual data structures.
Relational Database Schema Updates
SQLAlchemy models in database/models.py define columns for XhsNote and XhsNoteComment. Adding a field requires database migrations and corresponding updates to add_content and update_content methods in the storage implementations (lines 40-57 in _store_impl.py).
Schema Flexibility in NoSQL and File Stores
MongoDB implementations tolerate schema drift better than relational stores, but the AbstractStore hierarchy still requires explicit field handling. For CSV, JSONL, and Excel outputs, the AsyncFileWriter manages file paths based on platform and crawler type, ensuring new fields propagate correctly to data/xhs/ subdirectories.
Extending Storage Backends
To support a new backend like PostgreSQL or a new MongoDB collection, inherit from AbstractStore in base/base_crawler.py and implement store_content, store_comment, and store_creator. The existing implementations in store/xhs/_store_impl.py (lines 42-106) provide templates for CSV, JSONL, SQLite, MongoDB, and Excel.
Maintaining Async I/O Compatibility
The crawler relies on aiofiles for non-blocking writes in XiaoHongShuImage.save_image and XiaoHongShuVideo.save_video (lines 78-82 in xhs_store_media.py). Verify compatibility when upgrading Python versions or aiofiles releases to prevent blocking operations that degrade throughput.
Configuration Drift Management
Platform-specific settings reside in config/xhs_config.py. When Xiaohongshu introduces new requirements like mandatory API keys, expose these via the global config module and ensure they reach store constructors such as AsyncFileWriter(platform="xhs", crawler_type=...).
Practical Code Examples
Storing Content in MongoDB
from store.xhs._store_impl import XhsMongoStoreImplement
note_data = {
"note_id": "12345",
"creator_hash": "abcde",
"nickname": "Alice",
"title": "My travel diary",
"desc": "A short description",
"video_url": "https://example.com/video.mp4",
"time": "2026-07-01T12:00:00Z",
"liked_count": 42,
}
mongo_store = XhsMongoStoreImplement()
await mongo_store.store_content(note_data)
This implementation maps the normalized dict to MongoDB documents according to the schema defined in lines 52-60 of _store_impl.py.
Persisting Binary Images
from store.xhs.xhs_store_media import XiaoHongShuImage
image_store = XiaoHongShuImage()
await image_store.store_image({
"notice_id": "12345",
"pic_content": b"...binary image data...",
"extension_file_name": "cover.jpg"
})
The helper builds a path like data/xhs/images/12345/cover.jpg and writes bytes asynchronously using aiofiles (lines 78-82 in xhs_store_media.py).
Exporting to CSV
from store.xhs._store_impl import XhsCsvStoreImplement
csv_store = XhsCsvStoreImplement()
await csv_store.store_content({
"note_id": "12345",
"title": "My travel diary",
"desc": "Description text",
})
The CSV implementation delegates to AsyncFileWriter.write_to_csv, which automatically creates daily-named files under data/xhs/ (lines 53-55 in _store_impl.py).
Summary
- Monitor API contracts in
store/xhs/_store_impl.pyto ensure field mappings match platform responses. - Update authentication handling when
xsec_tokengeneration changes, ensuring tokens are stored via the storage layer. - Adjust rate-limiting using the
crawler_typeconfiguration andAsyncFileWritercontext variables. - Synchronize data models across SQLAlchemy ORM classes, MongoDB schemas, and file-based stores when fields are added or removed.
- Verify async compatibility when upgrading Python or
aiofilesto prevent I/O blocking. - Extend the abstraction by inheriting from
AbstractStoreinbase/base_crawler.pyto add new backends without modifying core logic.
Frequently Asked Questions
How do I update the XHS crawler when Xiaohongshu changes API field names?
Update the field mappings in store/xhs/_store_impl.py where the store_content methods normalize API responses. For example, if the platform renames title to note_title, modify the dict key extraction in XhsMongoStoreImplement.store_content (lines 64-72) to maintain compatibility with the storage schema.
Where are authentication tokens like xsec_token stored?
The storage layer persists xsec_token values alongside content metadata (lines 58-60 in _store_impl.py). When the platform updates token generation logic, you must update the request-building code to extract the new tokens and ensure the storage schema includes columns or fields to persist them.
How do I add a new database backend to the XHS crawler?
Create a new class in store/xhs/_store_impl.py that inherits from AbstractStore (defined in base/base_crawler.py). Implement the three required async methods: store_content, store_comment, and store_creator. Follow the existing patterns for XhsDbStoreImplement or XhsMongoStoreImplement to ensure compatibility with the crawler's data flow.
What happens when rate limits change on the platform?
Modify the global crawler_type configuration in config/xhs_config.py to adjust request frequency. The AsyncFileWriter class uses crawler_type_var.get() (lines 45-46 in _store_impl.py) to determine file organization and can be extended to respect new throttling requirements. Consider adding exponential backoff logic to the request pipeline if the platform imposes stricter limits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →