Where to Find the Xiaohongshu (XHS) Scraping Implementation in MediaCrawler

The Xiaohongshu scraping logic in NanmiCoder/MediaCrawler is organized across modular components in media_platform/xhs/ for extraction and store/xhs/ for data persistence, with entry points exposed through store/xhs/__init__.py.

The MediaCrawler open-source project provides a robust framework for scraping content from multiple Chinese social media platforms. If you need to understand how the Xiaohongshu (XHS) scraping works, the implementation spans several specialized modules handling authentication, HTML extraction, and multi-format storage.

Crawl Entry Point and High-Level API

The primary interface for Xiaohongshu operations resides in store/xhs/__init__.py. This module exposes high-level helper functions that the generic crawler orchestrator (defined in base/base_crawler.py) invokes when processing XHS data.

Key functions include:

  • update_xhs_note() – Persists note metadata and content
  • update_xhs_note_comment() – Handles individual comment storage
  • batch_update_xhs_note_comments() – Processes comment batches
  • Media download utilities for images and videos

These functions abstract the underlying storage implementations, allowing the crawler to remain agnostic about whether data lands in CSV, JSON, SQLite, or MongoDB.

Data Extraction and HTML Parsing

Raw HTML processing occurs in media_platform/xhs/extractor.py, which contains the XiaoHongShuExtractor class. This component parses note pages and extracts structured data from embedded JavaScript variables.

The extractor provides two primary methods:

from media_platform.xhs.extractor import XiaoHongShuExtractor

extractor = XiaoHongShuExtractor()
note_detail = extractor.extract_note_detail_from_html(note_id, html)
creator_info = extractor.extract_creator_info_from_html(html)

The extract_note_detail_from_html() method returns a dictionary containing fields like title, desc, image_list, and engagement metrics, while extract_creator_info_from_html() extracts creator profile data from the same HTML source.

Authentication and Request Signatures

Xiaohongshu requires cryptographically signed requests to access note data. The repository handles this through two specialized modules:

media_platform/xhs/playwright_sign.py and media_platform/xhs/xhs_sign.py – These files generate required request signatures including X-Sec-Token and X-Trace-Id using the xhshow library. The Playwright-based implementation executes JavaScript in a headless browser to produce tokens that match XHS anti-bot expectations.

media_platform/xhs/login.py – Manages the phone-SMS verification flow and persists XHS-specific cookies for subsequent authenticated requests. This module ensures the crawler maintains valid session state without requiring manual intervention.

Data Models and Schema Validation

The Pydantic models defining URL structures and data schemas live in model/m_xiaohongshu.py. This file contains:

  • NoteUrlInfo – Validates note URL patterns and extracts note IDs
  • CreatorUrlInfo – Handles creator profile URL parsing

These models enforce type safety and provide validation before data enters the storage pipeline.

Storage Implementations

Depending on your configuration, MediaCrawler supports multiple backend formats for XHS data. All implementations reside in store/xhs/_store_impl.py and inherit from the AbstractStore interface defined in base/base_crawler.py.

Available storage classes include:

  • XhsCsvStoreImplement – Writes to CSV files using AsyncFileWriter (see tools/async_file_writer.py)
  • XhsJsonStoreImplement / XhsJsonlStoreImplement – JSON and JSONL output formats
  • XhsSqliteStoreImplement – SQLite database via SQLAlchemy
  • XhsMongoStoreImplement – MongoDB document storage
  • XhsExcelStoreImplement – Excel file output (singleton pattern)

Each class implements the standard interface:

async def store_content(self, content_item: Dict): ...
async def store_comment(self, comment_item: Dict): ...

Data files are written to data/xhs/ directories by default, with subdirectories for images and videos.

Media Download Handling

Image and video assets referenced in notes are handled by store/xhs/xhs_store_media.py. This module manages asynchronous downloads and organizes files by note ID:

def download_image(self, url, note_id):
    # Saves to data/xhs/images/<note_id>/

    ...

def download_video(self, url, note_id):
    # Saves to data/xhs/videos/<note_id>/

    ...

The media downloader respects the platform-specific rate limits and directory structures defined in config/xhs_config.py.

Configuration

Platform-specific settings including default paths, request timeouts, and rate limits are centralized in config/xhs_config.py. This file allows you to adjust scraping behavior without modifying core logic.

End-to-End Scraping Flow

Understanding the complete data flow helps when debugging or extending the scraper:

  1. Initiation – api/main.py parses CLI arguments (--platform xhs) and instantiates the crawler
  2. Authentication – The system loads cookies from media_platform/xhs/login.py or generates fresh signatures via playwright_sign.py
  3. Fetching – HTTP requests execute through httpx_util or Playwright to retrieve HTML
  4. Extraction – XiaoHongShuExtractor parses the response into structured dictionaries
  5. Storage – The factory in store/xhs/__init__.py routes data to the appropriate store implementation in _store_impl.py
  6. Media Handling – xhs_store_media.py downloads associated images and videos to local storage

Practical Code Examples

Storing a Note with the High-Level API

import asyncio
from store.xhs import update_xhs_note

note_item = {
    "note_id": "1234567890",
    "creator_hash": "abcde12345",
    "title": "My XHS Note",
    "liked_count": 42,
    "desc": "Sample description"
}

asyncio.run(update_xhs_note(note_item))

Direct MongoDB Storage

from store.xhs._store_impl import XhsMongoStoreImplement

mongo_store = XhsMongoStoreImplement()
await mongo_store.store_content({
    "note_id": "1234567890",
    "title": "Sample Note",
    "liked_count": 42,
    "collected_count": 15
})

Extracting Creator Information

from media_platform.xhs.extractor import XiaoHongShuExtractor

html_content = "<html>...</html>"  # Fetched page content

extractor = XiaoHongShuExtractor()
creator = extractor.extract_creator_info_from_html(html_content)
print(creator.get("nickname"))

Summary

Frequently Asked Questions

How do I extend the Xiaohongshu scraper to capture additional fields?

Modify the extract_note_detail_from_html() method in media_platform/xhs/extractor.py to parse additional JavaScript variables from the HTML, then update the corresponding storage methods in store/xhs/_store_impl.py to handle the new fields. Ensure you also update the Pydantic models in model/m_xiaohongshu.py if the changes affect URL validation.

Where is the X-Sec-Token signature generated for XHS requests?

The signature generation occurs in media_platform/xhs/playwright_sign.py and media_platform/xhs/xhs_sign.py, which utilize the xhshow library to produce cryptographic tokens. These files generate the X-Sec-Token and X-Trace-Id headers required by Xiaohongshu's anti-bot systems.

Can I store XHS data in multiple formats simultaneously?

Yes. The store factory in store/xhs/__init__.py can instantiate multiple implementations from _store_impl.py (such as XhsCsvStoreImplement and XhsMongoStoreImplement) and call their store_content() methods sequentially. Configure your desired backends in the main crawler configuration before initialization.

What controls the rate limiting for XHS scraping?

Rate limits and request delays are configured in config/xhs_config.py. This file contains platform-specific settings for concurrency limits, sleep intervals between requests, and retry policies that prevent the crawler from triggering IP bans or CAPTCHA challenges.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →