XHS Data Models in MediaCrawler: How Scraped Xiaohongshu Content Is Structured

MediaCrawler utilizes Pydantic-based models NoteUrlInfo and CreatorUrlInfo to parse Xiaohongshu URLs, while the XhsNote ORM class handles persistent storage of scraped content.

The MediaCrawler repository provides a robust framework for scraping content from Chinese social media platforms. When targeting Xiaohongshu (XHS), the codebase employs specific data models to transform raw URLs into strongly-typed objects and store extracted content. These XHS data models form the backbone of the extraction pipeline, ensuring type safety and consistent data handling throughout the crawling process.

Core Pydantic Models for XHS URL Parsing

The primary data structures for XHS content scraping reside in model/m_xiaohongshu.py. These models validate and encapsulate authentication tokens and identifiers extracted from URLs before the crawler makes API requests.

NoteUrlInfo

The NoteUrlInfo model captures the essential identifiers required to fetch a specific XHS note. It extracts and validates the note_id, xsec_token, and xsec_source from raw URLs encountered during the crawling process.

Key fields include:

  • note_id: str – The unique identifier for the XHS note
  • xsec_token: str – Authentication token required for API calls
  • xsec_source: str – Token source indicator (e.g., pc_search)

This model is instantiated by the parse_note_info_from_note_url function in media_platform/xhs/help.py, which handles the regex extraction and validation logic.

CreatorUrlInfo

The CreatorUrlInfo model handles user profile URLs, storing the creator's unique identifier along with optional authentication tokens. It mirrors the structure of NoteUrlInfo but focuses on user-centric rather than content-centric data.

Key fields include:

  • user_id: str – The creator's unique identifier
  • xsec_token: str – Optional authentication token (defaults to empty)
  • xsec_source: str – Optional source indicator (defaults to empty)

The parse_creator_info_from_url function in media_platform/xhs/help.py produces these instances when processing creator homepage URLs.

Database Persistence with the XhsNote ORM

While the Pydantic models handle URL parsing and API request preparation, the actual scraped content persistence relies on the XhsNote class defined in database/models.py. This ORM model maps to a relational database table and stores the enriched note data returned by the XHS API.

The XhsNote model captures comprehensive content including:

  • Content metadata (note_id, title, desc)
  • Media assets (image_list)
  • Creator information (creator_hash, nickname)
  • Authentication context (xsec_token)

The storage layer in store/xhs/_store_impl.py receives a plain dictionary representation of the scraped data and handles the transition from the extraction models to the persistent ORM layer.

The XHS Data Flow: From URL to Storage

Understanding how these models interact reveals the architecture of the XHS scraping pipeline. The process flows through three distinct stages:

  1. URL Parsing – The help.py module converts raw XHS URLs into NoteUrlInfo or CreatorUrlInfo instances
  2. Data Extraction – The extractor.py module uses these models to make authenticated API requests and enrich the data
  3. Storage – The store implementation converts the extracted data into dictionaries and persists via the XhsNote ORM

This separation of concerns ensures that URL validation occurs early in the pipeline, while the heavy lifting of content storage happens through a clean abstraction layer.

Practical Implementation Examples

The following examples demonstrate how to instantiate these models and use them in the XHS crawling workflow.

Parsing a Note URL

To extract structured data from a raw XHS note URL:

from model.m_xiaohongshu import NoteUrlInfo
from media_platform.xhs.help import parse_note_info_from_note_url

raw_url = "https://www.xiaohongshu.com/explore/66fad51c000000001b0224b8?xsec_token=AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=&xsec_source=pc_search"
note_info: NoteUrlInfo = parse_note_info_from_note_url(raw_url)

print(note_info.note_id)      # → 66fad51c000000001b0224b8

print(note_info.xsec_token)   # → AB3rO‑QopW5sgrJ41GwN01...

print(note_info.xsec_source)  # → pc_search

Parsing a Creator URL

Similarly, extract creator information from profile URLs:

from model.m_xiaohongshu import CreatorUrlInfo
from media_platform.xhs.help import parse_creator_info_from_url

creator_url = "https://www.xiaohongshu.com/user/profile/5eb8e1d400000000010075ae?xsec_token=AB1nWBKCo1vE2HEkfoJUOi5B6BE5n7wVrbdpHoWIj5xHw=&xsec_source=pc_feed"
creator_info: CreatorUrlInfo = parse_creator_info_from_url(creator_url)

print(creator_info.user_id)   # → 5eb8e1d400000000010075ae

print(creator_info.xsec_token)   # token string (may be empty)

print(creator_info.xsec_source)  # → pc_feed

Storing Scraped Content

Once extracted, store the content using the database implementation:

from store.xhs._store_impl import XhsDbStoreImplement
import asyncio

async def store_note():
    db_store = XhsDbStoreImplement()
    note_dict = {
        "note_id": note_info.note_id,
        "creator_hash": "hashed_creator_id",
        "nickname": "匿名用户",
        "title": "示例笔记标题",
        "desc": "笔记正文内容",
        "image_list": ["https://.../image1.png", "https://.../image2.png"],
        "tag_list": ["#标签1", "#标签2"],
        "xsec_token": note_info.xsec_token,
        # ... other fields required by XhsNote model

    }
    await db_store.store_content(note_dict)

asyncio.run(store_note())

Summary

The MediaCrawler repository implements a layered approach to XHS data modeling:

This architecture ensures type safety during the extraction phase while maintaining flexibility for downstream storage operations.

Frequently Asked Questions

What is the difference between NoteUrlInfo and CreatorUrlInfo?

NoteUrlInfo extracts data from individual post URLs containing note_id parameters, while CreatorUrlInfo handles user profile URLs containing user_id parameters. Both models validate xsec_token and xsec_source fields, but target different XHS endpoints. The NoteUrlInfo model feeds the content extraction pipeline, whereas CreatorUrlInfo supports user-specific crawling operations.

Where does MediaCrawler store the actual scraped XHS content?

Scraped content persists through the XhsNote ORM class defined in database/models.py. This model maps to a relational database table and stores comprehensive note data including titles, descriptions, image lists, and creator metadata. The store/xhs/_store_impl.py module handles the actual insertion logic, supporting both database and file-based (CSV/JSON) storage backends.

How does the extractor use these data models?

The extractor in media_platform/xhs/extractor.py consumes NoteUrlInfo instances to make authenticated API requests. It uses the note_id and xsec_token fields from these models to construct proper request headers and parameters. After fetching raw data from the XHS API, the extractor transforms the response into dictionaries that match the schema expected by the XhsNote ORM.

Are these data models specific to Xiaohongshu only?

Yes, the NoteUrlInfo and CreatorUrlInfo models in model/m_xiaohongshu.py are specifically designed for XHS URL patterns and authentication token structures. MediaCrawler maintains separate model files for different platforms (e.g., Douyin, Weibo) in the model/ directory, each with platform-specific validation logic and field requirements tailored to their respective APIs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →