What Research and Tools Power the XHS Scraping Logic in MediaCrawler

The XHS scraping logic in MediaCrawler combines the open-source xhshow signature algorithm, Playwright browser automation, and custom HTML extraction utilities to reverse-engineer XiaoHongShu's private API.

The MediaCrawler repository by NanmiCoder implements a robust XHS (XiaoHongShu) scraping system that relies on community research into the platform's signing protocol. This article examines the specific open-source tools, cryptographic implementations, and browser automation techniques that enable the extraction of note data, creator information, and media from XiaoHongShu's protected endpoints.

Core Signature Algorithm: The xhshow Implementation

The foundation of the XHS scraping logic rests on xhshow, an open-source pure-algorithm implementation of XiaoHongShu's signing protocol. This library generates the critical x-s-common, x-t, and x-b3-traceid headers required to authenticate requests against the platform's private API.

How the Signing Protocol Works

According to the source code in media_platform/xhs/xhs_sign.py, the signature generation involves several cryptographic steps:

  • CRC32-based field computation using custom functions like mrc and encode_utf8
  • Custom Base64 variant encoding via b64_encode that matches XiaoHongShu's specific expectations
  • UTF-8 string encoding handled by the encode_utf8 utility to ensure proper byte alignment

The sign_with_xhshow function in media_platform/xhs/playwright_sign.py wraps this logic, injecting it into a Playwright browser context to obtain valid signed headers.

Patching the a3_hash Bug

The MediaCrawler implementation includes a custom patch for a known bug in xhshow's a3_hash calculation. The _patch_xhshow_a3_hash function modifies the algorithm to produce correct hash values that match the platform's server-side validation, ensuring the generated signatures are accepted by XiaoHongShu's API.

Browser Automation with Playwright

To execute the signature code and obtain valid session cookies, the scraper uses Playwright (Python) to launch a headless Chromium instance. This approach bridges the gap between pure algorithmic signature generation and the dynamic cookie requirements of the platform.

Headless Execution for Dynamic Headers

The media_platform/xhs/playwright_sign.py file orchestrates the browser injection process. It loads the patched xhshow implementation, executes the signing logic within the browser context, and harvests the resulting headers along with the dynamic trace_id required for request tracing.

The media_platform/xhs/login.py module leverages this Playwright-based sign-helper to perform the initial login flow, retrieving and persisting session cookies that are subsequently used for authenticated data requests.

HTML Parsing and Data Extraction

Once authenticated requests retrieve XHS pages, the scraper extracts structured data from the HTML responses using a combination of regular expressions and key normalization libraries.

Extracting window.INITIAL_STATE

The XiaoHongShuExtractor class in media_platform/xhs/extractor.py implements the core parsing logic. It uses regular expressions to locate and extract the window.__INITIAL_STATE__ JSON payload embedded within the page HTML. This payload contains comprehensive note details, creator information, and comment threads.

The extractor provides specific methods for different data types:

  • extract_note_detail_from_html – Parses note content, images, and metadata
  • extract_creator_info_from_html – Extracts user profiles and statistics

Normalizing CamelCase with humps

Raw XHS data uses CamelCase keys that are inconvenient for Python consumption. The scraper utilizes the humps library, specifically humps.decamelize, to transform JSON keys into snake_case format. This normalization occurs during the extraction phase, ensuring downstream storage operations receive consistently formatted data.

Data Persistence and Storage

The scraping pipeline concludes with asynchronous storage operations managed through platform-specific store modules.

MongoDB and File System Integration

The store/xhs/_store_impl.py module implements async file writers and MongoDB persistence layers for XHS data. It handles the write operations for notes, comments, and media metadata. Complementing this, store/xhs/xhs_store_media.py manages image and video path resolution, coordinating downloads and local file storage.

Configuration centralization resides in config/xhs_config.py, which defines platform-specific paths, default timeouts, and storage parameters used across the extraction pipeline.

Complete Scraping Pipeline Example

The following example demonstrates how the components integrate to fetch a note and persist its media:


# Example: Fetch a note and store its media

import asyncio
from store import xhs as xs
from utils import logger

async def fetch_and_save(note_id: str, html: str):
    # 1️⃣ Extract note details from raw HTML

    extractor = xs.XiaoHongShuExtractor()
    note_data = extractor.extract_note_detail_from_html(note_id, html)
    if not note_data:
        logger.error("Failed to extract note data")
        return

    # 2️⃣ Update note record in the DB

    await xs.update_xhs_note(note_data)

    # 3️⃣ Download and store images / videos (if any)

    for img_url in note_data.get("image_list", []):
        await xs.update_xhs_note_image(note_id, await xs.download(img_url), "jpg")
    for video in note_data.get("video_list", []):
        await xs.update_xhs_note_video(note_id, await xs.download(video), "mp4")

# Run the coroutine

asyncio.run(fetch_and_save("1234567890", raw_html))

To generate the signed headers required for the initial request, use the Playwright wrapper:


# Example: Generating signed headers for a request

from media_platform.xhs.playwright_sign import sign_with_xhshow

def get_signed_headers(cookie: str, payload: str):
    # The function internally uses the patched xhshow implementation

    # and returns a dict with the required X‑HS headers.

    return sign_with_xhshow(cookie, payload)

signed = get_signed_headers("a1=xxxxx; webId=yyy;", '{"note_id":"123"}')
print(signed["x-s-common"], signed["x-t"], signed["x-b3-traceid"])

Summary

  • xhshow algorithm: The open-source signing protocol implementation generates required authentication headers (x-s-common, x-t, x-b3-traceid) with a custom patch for the a3_hash calculation.
  • Playwright integration: Headless browser automation in media_platform/xhs/playwright_sign.py executes signature logic and harvests dynamic cookies and trace IDs.
  • HTML extraction: The XiaoHongShuExtractor class in media_platform/xhs/extractor.py parses window.__INITIAL_STATE__ JSON and uses the hums library to normalize CamelCase keys.
  • Cryptographic utilities: media_platform/xhs/xhs_sign.py provides CRC32-based encoding, custom Base64 variants, and UTF-8 handling specific to XHS requirements.
  • Async storage: store/xhs/_store_impl.py and store/xhs/xhs_store_media.py manage MongoDB persistence and media file downloads.

Frequently Asked Questions

What is the xhshow library used in the XHS scraping logic?

xhshow is an open-source pure-algorithm implementation of XiaoHongShu's signing protocol that produces the cryptographic headers required to authenticate API requests. MediaCrawler wraps this library in media_platform/xhs/playwright_sign.py and applies a custom patch to fix an a3_hash calculation bug, ensuring valid signature generation.

Why does MediaCrawler use Playwright instead of pure Python requests?

The XHS scraping logic requires dynamic elements including session cookies and the x-b3-traceid header that are difficult to generate statically. Playwright launches a headless Chromium browser to execute the xhshow JavaScript signing code, harvest valid cookies from media_platform/xhs/login.py, and obtain real-time trace identifiers that satisfy the platform's anti-bot measures.

How does the scraper extract data from XHS HTML pages?

After fetching pages with signed headers, the XiaoHongShuExtractor class in media_platform/xhs/extractor.py uses regular expressions to locate the window.__INITIAL_STATE__ JSON payload embedded in the HTML. The extractor then uses the humps library to convert CamelCase keys to snake_case, producing normalized Python dictionaries for storage.

Where is scraped XHS data stored in the MediaCrawler project?

Scraped data persists through asynchronous operations defined in store/xhs/_store_impl.py, which handles MongoDB connections and file system writes. Media files (images and videos) are managed by store/xhs/xhs_store_media.py, which resolves storage paths based on configuration in config/xhs_config.py and performs async downloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →