What Research and Tools Power the XHS Scraping Logic in MediaCrawler
The XHS scraping logic in MediaCrawler combines the open-source xhshow signature algorithm, Playwright browser automation, and custom HTML extraction utilities to reverse-engineer XiaoHongShu's private API.
The MediaCrawler repository by NanmiCoder implements a robust XHS (XiaoHongShu) scraping system that relies on community research into the platform's signing protocol. This article examines the specific open-source tools, cryptographic implementations, and browser automation techniques that enable the extraction of note data, creator information, and media from XiaoHongShu's protected endpoints.
Core Signature Algorithm: The xhshow Implementation
The foundation of the XHS scraping logic rests on xhshow, an open-source pure-algorithm implementation of XiaoHongShu's signing protocol. This library generates the critical x-s-common, x-t, and x-b3-traceid headers required to authenticate requests against the platform's private API.
How the Signing Protocol Works
According to the source code in media_platform/xhs/xhs_sign.py, the signature generation involves several cryptographic steps:
- CRC32-based field computation using custom functions like
mrcandencode_utf8 - Custom Base64 variant encoding via
b64_encodethat matches XiaoHongShu's specific expectations - UTF-8 string encoding handled by the
encode_utf8utility to ensure proper byte alignment
The sign_with_xhshow function in media_platform/xhs/playwright_sign.py wraps this logic, injecting it into a Playwright browser context to obtain valid signed headers.
Patching the a3_hash Bug
The MediaCrawler implementation includes a custom patch for a known bug in xhshow's a3_hash calculation. The _patch_xhshow_a3_hash function modifies the algorithm to produce correct hash values that match the platform's server-side validation, ensuring the generated signatures are accepted by XiaoHongShu's API.
Browser Automation with Playwright
To execute the signature code and obtain valid session cookies, the scraper uses Playwright (Python) to launch a headless Chromium instance. This approach bridges the gap between pure algorithmic signature generation and the dynamic cookie requirements of the platform.
Headless Execution for Dynamic Headers
The media_platform/xhs/playwright_sign.py file orchestrates the browser injection process. It loads the patched xhshow implementation, executes the signing logic within the browser context, and harvests the resulting headers along with the dynamic trace_id required for request tracing.
The media_platform/xhs/login.py module leverages this Playwright-based sign-helper to perform the initial login flow, retrieving and persisting session cookies that are subsequently used for authenticated data requests.
HTML Parsing and Data Extraction
Once authenticated requests retrieve XHS pages, the scraper extracts structured data from the HTML responses using a combination of regular expressions and key normalization libraries.
Extracting window.INITIAL_STATE
The XiaoHongShuExtractor class in media_platform/xhs/extractor.py implements the core parsing logic. It uses regular expressions to locate and extract the window.__INITIAL_STATE__ JSON payload embedded within the page HTML. This payload contains comprehensive note details, creator information, and comment threads.
The extractor provides specific methods for different data types:
extract_note_detail_from_html– Parses note content, images, and metadataextract_creator_info_from_html– Extracts user profiles and statistics
Normalizing CamelCase with humps
Raw XHS data uses CamelCase keys that are inconvenient for Python consumption. The scraper utilizes the humps library, specifically humps.decamelize, to transform JSON keys into snake_case format. This normalization occurs during the extraction phase, ensuring downstream storage operations receive consistently formatted data.
Data Persistence and Storage
The scraping pipeline concludes with asynchronous storage operations managed through platform-specific store modules.
MongoDB and File System Integration
The store/xhs/_store_impl.py module implements async file writers and MongoDB persistence layers for XHS data. It handles the write operations for notes, comments, and media metadata. Complementing this, store/xhs/xhs_store_media.py manages image and video path resolution, coordinating downloads and local file storage.
Configuration centralization resides in config/xhs_config.py, which defines platform-specific paths, default timeouts, and storage parameters used across the extraction pipeline.
Complete Scraping Pipeline Example
The following example demonstrates how the components integrate to fetch a note and persist its media:
# Example: Fetch a note and store its media
import asyncio
from store import xhs as xs
from utils import logger
async def fetch_and_save(note_id: str, html: str):
# 1️⃣ Extract note details from raw HTML
extractor = xs.XiaoHongShuExtractor()
note_data = extractor.extract_note_detail_from_html(note_id, html)
if not note_data:
logger.error("Failed to extract note data")
return
# 2️⃣ Update note record in the DB
await xs.update_xhs_note(note_data)
# 3️⃣ Download and store images / videos (if any)
for img_url in note_data.get("image_list", []):
await xs.update_xhs_note_image(note_id, await xs.download(img_url), "jpg")
for video in note_data.get("video_list", []):
await xs.update_xhs_note_video(note_id, await xs.download(video), "mp4")
# Run the coroutine
asyncio.run(fetch_and_save("1234567890", raw_html))
To generate the signed headers required for the initial request, use the Playwright wrapper:
# Example: Generating signed headers for a request
from media_platform.xhs.playwright_sign import sign_with_xhshow
def get_signed_headers(cookie: str, payload: str):
# The function internally uses the patched xhshow implementation
# and returns a dict with the required X‑HS headers.
return sign_with_xhshow(cookie, payload)
signed = get_signed_headers("a1=xxxxx; webId=yyy;", '{"note_id":"123"}')
print(signed["x-s-common"], signed["x-t"], signed["x-b3-traceid"])
Summary
- xhshow algorithm: The open-source signing protocol implementation generates required authentication headers (
x-s-common,x-t,x-b3-traceid) with a custom patch for thea3_hashcalculation. - Playwright integration: Headless browser automation in
media_platform/xhs/playwright_sign.pyexecutes signature logic and harvests dynamic cookies and trace IDs. - HTML extraction: The
XiaoHongShuExtractorclass inmedia_platform/xhs/extractor.pyparseswindow.__INITIAL_STATE__JSON and uses the hums library to normalize CamelCase keys. - Cryptographic utilities:
media_platform/xhs/xhs_sign.pyprovides CRC32-based encoding, custom Base64 variants, and UTF-8 handling specific to XHS requirements. - Async storage:
store/xhs/_store_impl.pyandstore/xhs/xhs_store_media.pymanage MongoDB persistence and media file downloads.
Frequently Asked Questions
What is the xhshow library used in the XHS scraping logic?
xhshow is an open-source pure-algorithm implementation of XiaoHongShu's signing protocol that produces the cryptographic headers required to authenticate API requests. MediaCrawler wraps this library in media_platform/xhs/playwright_sign.py and applies a custom patch to fix an a3_hash calculation bug, ensuring valid signature generation.
Why does MediaCrawler use Playwright instead of pure Python requests?
The XHS scraping logic requires dynamic elements including session cookies and the x-b3-traceid header that are difficult to generate statically. Playwright launches a headless Chromium browser to execute the xhshow JavaScript signing code, harvest valid cookies from media_platform/xhs/login.py, and obtain real-time trace identifiers that satisfy the platform's anti-bot measures.
How does the scraper extract data from XHS HTML pages?
After fetching pages with signed headers, the XiaoHongShuExtractor class in media_platform/xhs/extractor.py uses regular expressions to locate the window.__INITIAL_STATE__ JSON payload embedded in the HTML. The extractor then uses the humps library to convert CamelCase keys to snake_case, producing normalized Python dictionaries for storage.
Where is scraped XHS data stored in the MediaCrawler project?
Scraped data persists through asynchronous operations defined in store/xhs/_store_impl.py, which handles MongoDB connections and file system writes. Media files (images and videos) are managed by store/xhs/xhs_store_media.py, which resolves storage paths based on configuration in config/xhs_config.py and performs async downloads.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →