# Where to Find the Xiaohongshu (XHS) Scraping Implementation in MediaCrawler

> Find the Xiaohongshu scraping implementation in NanmiCoder/MediaCrawler within the media_platform/xhs/ and store/xhs/ directories. Access entry points via store/xhs/__init__.py.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-03

---

**The Xiaohongshu scraping logic in NanmiCoder/MediaCrawler is organized across modular components in `media_platform/xhs/` for extraction and `store/xhs/` for data persistence, with entry points exposed through [`store/xhs/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py).**

The MediaCrawler open-source project provides a robust framework for scraping content from multiple Chinese social media platforms. If you need to understand how the Xiaohongshu (XHS) scraping works, the implementation spans several specialized modules handling authentication, HTML extraction, and multi-format storage.

## Crawl Entry Point and High-Level API

The primary interface for Xiaohongshu operations resides in [`store/xhs/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py). This module exposes high-level helper functions that the generic crawler orchestrator (defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)) invokes when processing XHS data.

Key functions include:

- `update_xhs_note()` – Persists note metadata and content
- `update_xhs_note_comment()` – Handles individual comment storage
- `batch_update_xhs_note_comments()` – Processes comment batches
- Media download utilities for images and videos

These functions abstract the underlying storage implementations, allowing the crawler to remain agnostic about whether data lands in CSV, JSON, SQLite, or MongoDB.

## Data Extraction and HTML Parsing

Raw HTML processing occurs in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py), which contains the `XiaoHongShuExtractor` class. This component parses note pages and extracts structured data from embedded JavaScript variables.

The extractor provides two primary methods:

```python
from media_platform.xhs.extractor import XiaoHongShuExtractor

extractor = XiaoHongShuExtractor()
note_detail = extractor.extract_note_detail_from_html(note_id, html)
creator_info = extractor.extract_creator_info_from_html(html)

```

The `extract_note_detail_from_html()` method returns a dictionary containing fields like `title`, `desc`, `image_list`, and engagement metrics, while `extract_creator_info_from_html()` extracts creator profile data from the same HTML source.

## Authentication and Request Signatures

Xiaohongshu requires cryptographically signed requests to access note data. The repository handles this through two specialized modules:

**[`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py)** and **[`media_platform/xhs/xhs_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/xhs_sign.py)** – These files generate required request signatures including `X-Sec-Token` and `X-Trace-Id` using the `xhshow` library. The Playwright-based implementation executes JavaScript in a headless browser to produce tokens that match XHS anti-bot expectations.

**[`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py)** – Manages the phone-SMS verification flow and persists XHS-specific cookies for subsequent authenticated requests. This module ensures the crawler maintains valid session state without requiring manual intervention.

## Data Models and Schema Validation

The Pydantic models defining URL structures and data schemas live in [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py). This file contains:

- `NoteUrlInfo` – Validates note URL patterns and extracts note IDs
- `CreatorUrlInfo` – Handles creator profile URL parsing

These models enforce type safety and provide validation before data enters the storage pipeline.

## Storage Implementations

Depending on your configuration, MediaCrawler supports multiple backend formats for XHS data. All implementations reside in [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) and inherit from the `AbstractStore` interface defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py).

Available storage classes include:

- **`XhsCsvStoreImplement`** – Writes to CSV files using `AsyncFileWriter` (see [`tools/async_file_writer.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/async_file_writer.py))
- **`XhsJsonStoreImplement`** / **`XhsJsonlStoreImplement`** – JSON and JSONL output formats
- **`XhsSqliteStoreImplement`** – SQLite database via SQLAlchemy
- **`XhsMongoStoreImplement`** – MongoDB document storage
- **`XhsExcelStoreImplement`** – Excel file output (singleton pattern)

Each class implements the standard interface:

```python
async def store_content(self, content_item: Dict): ...
async def store_comment(self, comment_item: Dict): ...

```

Data files are written to `data/xhs/` directories by default, with subdirectories for images and videos.

## Media Download Handling

Image and video assets referenced in notes are handled by [`store/xhs/xhs_store_media.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/xhs_store_media.py). This module manages asynchronous downloads and organizes files by note ID:

```python
def download_image(self, url, note_id):
    # Saves to data/xhs/images/<note_id>/

    ...

def download_video(self, url, note_id):
    # Saves to data/xhs/videos/<note_id>/

    ...

```

The media downloader respects the platform-specific rate limits and directory structures defined in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py).

## Configuration

Platform-specific settings including default paths, request timeouts, and rate limits are centralized in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py). This file allows you to adjust scraping behavior without modifying core logic.

## End-to-End Scraping Flow

Understanding the complete data flow helps when debugging or extending the scraper:

1. **Initiation** – [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py) parses CLI arguments (`--platform xhs`) and instantiates the crawler
2. **Authentication** – The system loads cookies from [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) or generates fresh signatures via [`playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/playwright_sign.py)
3. **Fetching** – HTTP requests execute through `httpx_util` or Playwright to retrieve HTML
4. **Extraction** – `XiaoHongShuExtractor` parses the response into structured dictionaries
5. **Storage** – The factory in [`store/xhs/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) routes data to the appropriate store implementation in [`_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/_store_impl.py)
6. **Media Handling** – [`xhs_store_media.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/xhs_store_media.py) downloads associated images and videos to local storage

## Practical Code Examples

### Storing a Note with the High-Level API

```python
import asyncio
from store.xhs import update_xhs_note

note_item = {
    "note_id": "1234567890",
    "creator_hash": "abcde12345",
    "title": "My XHS Note",
    "liked_count": 42,
    "desc": "Sample description"
}

asyncio.run(update_xhs_note(note_item))

```

### Direct MongoDB Storage

```python
from store.xhs._store_impl import XhsMongoStoreImplement

mongo_store = XhsMongoStoreImplement()
await mongo_store.store_content({
    "note_id": "1234567890",
    "title": "Sample Note",
    "liked_count": 42,
    "collected_count": 15
})

```

### Extracting Creator Information

```python
from media_platform.xhs.extractor import XiaoHongShuExtractor

html_content = "<html>...</html>"  # Fetched page content

extractor = XiaoHongShuExtractor()
creator = extractor.extract_creator_info_from_html(html_content)
print(creator.get("nickname"))

```

## Summary

- The **entry point** for XHS operations is [`store/xhs/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py), which exposes `update_xhs_note` and related functions
- **HTML extraction** logic lives in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) via the `XiaoHongShuExtractor` class
- **Authentication** is handled by [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) (SMS flow) and [`playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/playwright_sign.py) (signature generation)
- **Storage implementations** in [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) support CSV, JSON, SQLite, and MongoDB formats
- **Media downloads** are managed by [`store/xhs/xhs_store_media.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/xhs_store_media.py), saving files to `data/xhs/`
- All components wire together through the `AbstractStore` interface defined in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)

## Frequently Asked Questions

### How do I extend the Xiaohongshu scraper to capture additional fields?

Modify the `extract_note_detail_from_html()` method in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) to parse additional JavaScript variables from the HTML, then update the corresponding storage methods in [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) to handle the new fields. Ensure you also update the Pydantic models in [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py) if the changes affect URL validation.

### Where is the X-Sec-Token signature generated for XHS requests?

The signature generation occurs in [`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py) and [`media_platform/xhs/xhs_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/xhs_sign.py), which utilize the `xhshow` library to produce cryptographic tokens. These files generate the `X-Sec-Token` and `X-Trace-Id` headers required by Xiaohongshu's anti-bot systems.

### Can I store XHS data in multiple formats simultaneously?

Yes. The store factory in [`store/xhs/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/__init__.py) can instantiate multiple implementations from [`_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/_store_impl.py) (such as `XhsCsvStoreImplement` and `XhsMongoStoreImplement`) and call their `store_content()` methods sequentially. Configure your desired backends in the main crawler configuration before initialization.

### What controls the rate limiting for XHS scraping?

Rate limits and request delays are configured in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py). This file contains platform-specific settings for concurrency limits, sleep intervals between requests, and retry policies that prevent the crawler from triggering IP bans or CAPTCHA challenges.