Where to Find the Xiaohongshu (XHS) Scraping Implementation in MediaCrawler
The Xiaohongshu scraping logic in NanmiCoder/MediaCrawler is organized across modular components in media_platform/xhs/ for extraction and store/xhs/ for data persistence, with entry points exposed through store/xhs/__init__.py.
The MediaCrawler open-source project provides a robust framework for scraping content from multiple Chinese social media platforms. If you need to understand how the Xiaohongshu (XHS) scraping works, the implementation spans several specialized modules handling authentication, HTML extraction, and multi-format storage.
Crawl Entry Point and High-Level API
The primary interface for Xiaohongshu operations resides in store/xhs/__init__.py. This module exposes high-level helper functions that the generic crawler orchestrator (defined in base/base_crawler.py) invokes when processing XHS data.
Key functions include:
update_xhs_note()– Persists note metadata and contentupdate_xhs_note_comment()– Handles individual comment storagebatch_update_xhs_note_comments()– Processes comment batches- Media download utilities for images and videos
These functions abstract the underlying storage implementations, allowing the crawler to remain agnostic about whether data lands in CSV, JSON, SQLite, or MongoDB.
Data Extraction and HTML Parsing
Raw HTML processing occurs in media_platform/xhs/extractor.py, which contains the XiaoHongShuExtractor class. This component parses note pages and extracts structured data from embedded JavaScript variables.
The extractor provides two primary methods:
from media_platform.xhs.extractor import XiaoHongShuExtractor
extractor = XiaoHongShuExtractor()
note_detail = extractor.extract_note_detail_from_html(note_id, html)
creator_info = extractor.extract_creator_info_from_html(html)
The extract_note_detail_from_html() method returns a dictionary containing fields like title, desc, image_list, and engagement metrics, while extract_creator_info_from_html() extracts creator profile data from the same HTML source.
Authentication and Request Signatures
Xiaohongshu requires cryptographically signed requests to access note data. The repository handles this through two specialized modules:
media_platform/xhs/playwright_sign.py and media_platform/xhs/xhs_sign.py – These files generate required request signatures including X-Sec-Token and X-Trace-Id using the xhshow library. The Playwright-based implementation executes JavaScript in a headless browser to produce tokens that match XHS anti-bot expectations.
media_platform/xhs/login.py – Manages the phone-SMS verification flow and persists XHS-specific cookies for subsequent authenticated requests. This module ensures the crawler maintains valid session state without requiring manual intervention.
Data Models and Schema Validation
The Pydantic models defining URL structures and data schemas live in model/m_xiaohongshu.py. This file contains:
NoteUrlInfo– Validates note URL patterns and extracts note IDsCreatorUrlInfo– Handles creator profile URL parsing
These models enforce type safety and provide validation before data enters the storage pipeline.
Storage Implementations
Depending on your configuration, MediaCrawler supports multiple backend formats for XHS data. All implementations reside in store/xhs/_store_impl.py and inherit from the AbstractStore interface defined in base/base_crawler.py.
Available storage classes include:
XhsCsvStoreImplement– Writes to CSV files usingAsyncFileWriter(seetools/async_file_writer.py)XhsJsonStoreImplement/XhsJsonlStoreImplement– JSON and JSONL output formatsXhsSqliteStoreImplement– SQLite database via SQLAlchemyXhsMongoStoreImplement– MongoDB document storageXhsExcelStoreImplement– Excel file output (singleton pattern)
Each class implements the standard interface:
async def store_content(self, content_item: Dict): ...
async def store_comment(self, comment_item: Dict): ...
Data files are written to data/xhs/ directories by default, with subdirectories for images and videos.
Media Download Handling
Image and video assets referenced in notes are handled by store/xhs/xhs_store_media.py. This module manages asynchronous downloads and organizes files by note ID:
def download_image(self, url, note_id):
# Saves to data/xhs/images/<note_id>/
...
def download_video(self, url, note_id):
# Saves to data/xhs/videos/<note_id>/
...
The media downloader respects the platform-specific rate limits and directory structures defined in config/xhs_config.py.
Configuration
Platform-specific settings including default paths, request timeouts, and rate limits are centralized in config/xhs_config.py. This file allows you to adjust scraping behavior without modifying core logic.
End-to-End Scraping Flow
Understanding the complete data flow helps when debugging or extending the scraper:
- Initiation –
api/main.pyparses CLI arguments (--platform xhs) and instantiates the crawler - Authentication – The system loads cookies from
media_platform/xhs/login.pyor generates fresh signatures viaplaywright_sign.py - Fetching – HTTP requests execute through
httpx_utilor Playwright to retrieve HTML - Extraction –
XiaoHongShuExtractorparses the response into structured dictionaries - Storage – The factory in
store/xhs/__init__.pyroutes data to the appropriate store implementation in_store_impl.py - Media Handling –
xhs_store_media.pydownloads associated images and videos to local storage
Practical Code Examples
Storing a Note with the High-Level API
import asyncio
from store.xhs import update_xhs_note
note_item = {
"note_id": "1234567890",
"creator_hash": "abcde12345",
"title": "My XHS Note",
"liked_count": 42,
"desc": "Sample description"
}
asyncio.run(update_xhs_note(note_item))
Direct MongoDB Storage
from store.xhs._store_impl import XhsMongoStoreImplement
mongo_store = XhsMongoStoreImplement()
await mongo_store.store_content({
"note_id": "1234567890",
"title": "Sample Note",
"liked_count": 42,
"collected_count": 15
})
Extracting Creator Information
from media_platform.xhs.extractor import XiaoHongShuExtractor
html_content = "<html>...</html>" # Fetched page content
extractor = XiaoHongShuExtractor()
creator = extractor.extract_creator_info_from_html(html_content)
print(creator.get("nickname"))
Summary
- The entry point for XHS operations is
store/xhs/__init__.py, which exposesupdate_xhs_noteand related functions - HTML extraction logic lives in
media_platform/xhs/extractor.pyvia theXiaoHongShuExtractorclass - Authentication is handled by
media_platform/xhs/login.py(SMS flow) andplaywright_sign.py(signature generation) - Storage implementations in
store/xhs/_store_impl.pysupport CSV, JSON, SQLite, and MongoDB formats - Media downloads are managed by
store/xhs/xhs_store_media.py, saving files todata/xhs/ - All components wire together through the
AbstractStoreinterface defined inbase/base_crawler.py
Frequently Asked Questions
How do I extend the Xiaohongshu scraper to capture additional fields?
Modify the extract_note_detail_from_html() method in media_platform/xhs/extractor.py to parse additional JavaScript variables from the HTML, then update the corresponding storage methods in store/xhs/_store_impl.py to handle the new fields. Ensure you also update the Pydantic models in model/m_xiaohongshu.py if the changes affect URL validation.
Where is the X-Sec-Token signature generated for XHS requests?
The signature generation occurs in media_platform/xhs/playwright_sign.py and media_platform/xhs/xhs_sign.py, which utilize the xhshow library to produce cryptographic tokens. These files generate the X-Sec-Token and X-Trace-Id headers required by Xiaohongshu's anti-bot systems.
Can I store XHS data in multiple formats simultaneously?
Yes. The store factory in store/xhs/__init__.py can instantiate multiple implementations from _store_impl.py (such as XhsCsvStoreImplement and XhsMongoStoreImplement) and call their store_content() methods sequentially. Configure your desired backends in the main crawler configuration before initialization.
What controls the rate limiting for XHS scraping?
Rate limits and request delays are configured in config/xhs_config.py. This file contains platform-specific settings for concurrency limits, sleep intervals between requests, and retry policies that prevent the crawler from triggering IP bans or CAPTCHA challenges.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →