# What Research and Tools Power the XHS Scraping Logic in MediaCrawler

> Learn how MediaCrawler's XHS scraping logic uses xhshow, Playwright, and custom tools to reverse-engineer XiaoHongShu's private API for efficient data extraction.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-03

---

**The XHS scraping logic in MediaCrawler combines the open-source xhshow signature algorithm, Playwright browser automation, and custom HTML extraction utilities to reverse-engineer XiaoHongShu's private API.**

The MediaCrawler repository by NanmiCoder implements a robust XHS (XiaoHongShu) scraping system that relies on community research into the platform's signing protocol. This article examines the specific open-source tools, cryptographic implementations, and browser automation techniques that enable the extraction of note data, creator information, and media from XiaoHongShu's protected endpoints.

## Core Signature Algorithm: The xhshow Implementation

The foundation of the XHS scraping logic rests on **xhshow**, an open-source pure-algorithm implementation of XiaoHongShu's signing protocol. This library generates the critical `x-s-common`, `x-t`, and `x-b3-traceid` headers required to authenticate requests against the platform's private API.

### How the Signing Protocol Works

According to the source code in [`media_platform/xhs/xhs_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/xhs_sign.py), the signature generation involves several cryptographic steps:

- **CRC32-based field computation** using custom functions like `mrc` and `encode_utf8`
- **Custom Base64 variant encoding** via `b64_encode` that matches XiaoHongShu's specific expectations
- **UTF-8 string encoding** handled by the `encode_utf8` utility to ensure proper byte alignment

The `sign_with_xhshow` function in [`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py) wraps this logic, injecting it into a Playwright browser context to obtain valid signed headers.

### Patching the a3_hash Bug

The MediaCrawler implementation includes a custom patch for a known bug in xhshow's `a3_hash` calculation. The `_patch_xhshow_a3_hash` function modifies the algorithm to produce correct hash values that match the platform's server-side validation, ensuring the generated signatures are accepted by XiaoHongShu's API.

## Browser Automation with Playwright

To execute the signature code and obtain valid session cookies, the scraper uses **Playwright** (Python) to launch a headless Chromium instance. This approach bridges the gap between pure algorithmic signature generation and the dynamic cookie requirements of the platform.

### Headless Execution for Dynamic Headers

The [`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py) file orchestrates the browser injection process. It loads the patched xhshow implementation, executes the signing logic within the browser context, and harvests the resulting headers along with the dynamic `trace_id` required for request tracing.

The [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) module leverages this Playwright-based sign-helper to perform the initial login flow, retrieving and persisting session cookies that are subsequently used for authenticated data requests.

## HTML Parsing and Data Extraction

Once authenticated requests retrieve XHS pages, the scraper extracts structured data from the HTML responses using a combination of regular expressions and key normalization libraries.

### Extracting window.__INITIAL_STATE__

The `XiaoHongShuExtractor` class in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) implements the core parsing logic. It uses regular expressions to locate and extract the `window.__INITIAL_STATE__` JSON payload embedded within the page HTML. This payload contains comprehensive note details, creator information, and comment threads.

The extractor provides specific methods for different data types:
- `extract_note_detail_from_html` – Parses note content, images, and metadata
- `extract_creator_info_from_html` – Extracts user profiles and statistics

### Normalizing CamelCase with humps

Raw XHS data uses CamelCase keys that are inconvenient for Python consumption. The scraper utilizes the **humps** library, specifically `humps.decamelize`, to transform JSON keys into snake_case format. This normalization occurs during the extraction phase, ensuring downstream storage operations receive consistently formatted data.

## Data Persistence and Storage

The scraping pipeline concludes with asynchronous storage operations managed through platform-specific store modules.

### MongoDB and File System Integration

The [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) module implements async file writers and MongoDB persistence layers for XHS data. It handles the write operations for notes, comments, and media metadata. Complementing this, [`store/xhs/xhs_store_media.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/xhs_store_media.py) manages image and video path resolution, coordinating downloads and local file storage.

Configuration centralization resides in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py), which defines platform-specific paths, default timeouts, and storage parameters used across the extraction pipeline.

## Complete Scraping Pipeline Example

The following example demonstrates how the components integrate to fetch a note and persist its media:

```python

# Example: Fetch a note and store its media

import asyncio
from store import xhs as xs
from utils import logger

async def fetch_and_save(note_id: str, html: str):
    # 1️⃣ Extract note details from raw HTML

    extractor = xs.XiaoHongShuExtractor()
    note_data = extractor.extract_note_detail_from_html(note_id, html)
    if not note_data:
        logger.error("Failed to extract note data")
        return

    # 2️⃣ Update note record in the DB

    await xs.update_xhs_note(note_data)

    # 3️⃣ Download and store images / videos (if any)

    for img_url in note_data.get("image_list", []):
        await xs.update_xhs_note_image(note_id, await xs.download(img_url), "jpg")
    for video in note_data.get("video_list", []):
        await xs.update_xhs_note_video(note_id, await xs.download(video), "mp4")

# Run the coroutine

asyncio.run(fetch_and_save("1234567890", raw_html))

```

To generate the signed headers required for the initial request, use the Playwright wrapper:

```python

# Example: Generating signed headers for a request

from media_platform.xhs.playwright_sign import sign_with_xhshow

def get_signed_headers(cookie: str, payload: str):
    # The function internally uses the patched xhshow implementation

    # and returns a dict with the required X‑HS headers.

    return sign_with_xhshow(cookie, payload)

signed = get_signed_headers("a1=xxxxx; webId=yyy;", '{"note_id":"123"}')
print(signed["x-s-common"], signed["x-t"], signed["x-b3-traceid"])

```

## Summary

- **xhshow algorithm**: The open-source signing protocol implementation generates required authentication headers (`x-s-common`, `x-t`, `x-b3-traceid`) with a custom patch for the `a3_hash` calculation.
- **Playwright integration**: Headless browser automation in [`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py) executes signature logic and harvests dynamic cookies and trace IDs.
- **HTML extraction**: The `XiaoHongShuExtractor` class in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) parses `window.__INITIAL_STATE__` JSON and uses the **hums** library to normalize CamelCase keys.
- **Cryptographic utilities**: [`media_platform/xhs/xhs_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/xhs_sign.py) provides CRC32-based encoding, custom Base64 variants, and UTF-8 handling specific to XHS requirements.
- **Async storage**: [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) and [`store/xhs/xhs_store_media.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/xhs_store_media.py) manage MongoDB persistence and media file downloads.

## Frequently Asked Questions

### What is the xhshow library used in the XHS scraping logic?

**xhshow** is an open-source pure-algorithm implementation of XiaoHongShu's signing protocol that produces the cryptographic headers required to authenticate API requests. MediaCrawler wraps this library in [`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py) and applies a custom patch to fix an `a3_hash` calculation bug, ensuring valid signature generation.

### Why does MediaCrawler use Playwright instead of pure Python requests?

The XHS scraping logic requires dynamic elements including session cookies and the `x-b3-traceid` header that are difficult to generate statically. Playwright launches a headless Chromium browser to execute the xhshow JavaScript signing code, harvest valid cookies from [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py), and obtain real-time trace identifiers that satisfy the platform's anti-bot measures.

### How does the scraper extract data from XHS HTML pages?

After fetching pages with signed headers, the `XiaoHongShuExtractor` class in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) uses regular expressions to locate the `window.__INITIAL_STATE__` JSON payload embedded in the HTML. The extractor then uses the **humps** library to convert CamelCase keys to snake_case, producing normalized Python dictionaries for storage.

### Where is scraped XHS data stored in the MediaCrawler project?

Scraped data persists through asynchronous operations defined in [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py), which handles MongoDB connections and file system writes. Media files (images and videos) are managed by [`store/xhs/xhs_store_media.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/xhs_store_media.py), which resolves storage paths based on configuration in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) and performs async downloads.