# XHS Data Models in MediaCrawler: How Scraped Xiaohongshu Content Is Structured

> Explore XHS data models in MediaCrawler. Learn how Pydantic models structure scraped Xiaohongshu content for efficient parsing and storage with the XhsNote ORM.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: architecture
- Published: 2026-07-03

---

**MediaCrawler utilizes Pydantic-based models `NoteUrlInfo` and `CreatorUrlInfo` to parse Xiaohongshu URLs, while the `XhsNote` ORM class handles persistent storage of scraped content.**

The MediaCrawler repository provides a robust framework for scraping content from Chinese social media platforms. When targeting Xiaohongshu (XHS), the codebase employs specific data models to transform raw URLs into strongly-typed objects and store extracted content. These XHS data models form the backbone of the extraction pipeline, ensuring type safety and consistent data handling throughout the crawling process.

## Core Pydantic Models for XHS URL Parsing

The primary data structures for XHS content scraping reside in **[`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py)**. These models validate and encapsulate authentication tokens and identifiers extracted from URLs before the crawler makes API requests.

### NoteUrlInfo

The **`NoteUrlInfo`** model captures the essential identifiers required to fetch a specific XHS note. It extracts and validates the `note_id`, `xsec_token`, and `xsec_source` from raw URLs encountered during the crawling process.

Key fields include:

- **`note_id: str`** – The unique identifier for the XHS note
- **`xsec_token: str`** – Authentication token required for API calls
- **`xsec_source: str`** – Token source indicator (e.g., `pc_search`)

This model is instantiated by the **`parse_note_info_from_note_url`** function in **[`media_platform/xhs/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/help.py)**, which handles the regex extraction and validation logic.

### CreatorUrlInfo

The **`CreatorUrlInfo`** model handles user profile URLs, storing the creator's unique identifier along with optional authentication tokens. It mirrors the structure of `NoteUrlInfo` but focuses on user-centric rather than content-centric data.

Key fields include:

- **`user_id: str`** – The creator's unique identifier
- **`xsec_token: str`** – Optional authentication token (defaults to empty)
- **`xsec_source: str`** – Optional source indicator (defaults to empty)

The **`parse_creator_info_from_url`** function in **[`media_platform/xhs/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/help.py)** produces these instances when processing creator homepage URLs.

## Database Persistence with the XhsNote ORM

While the Pydantic models handle URL parsing and API request preparation, the actual scraped content persistence relies on the **`XhsNote`** class defined in **[`database/models.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/models.py)**. This ORM model maps to a relational database table and stores the enriched note data returned by the XHS API.

The `XhsNote` model captures comprehensive content including:

- Content metadata (`note_id`, `title`, `desc`)
- Media assets (`image_list`)
- Creator information (`creator_hash`, `nickname`)
- Authentication context (`xsec_token`)

The storage layer in **[`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py)** receives a plain dictionary representation of the scraped data and handles the transition from the extraction models to the persistent ORM layer.

## The XHS Data Flow: From URL to Storage

Understanding how these models interact reveals the architecture of the XHS scraping pipeline. The process flows through three distinct stages:

1. **URL Parsing** – The [`help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/help.py) module converts raw XHS URLs into `NoteUrlInfo` or `CreatorUrlInfo` instances
2. **Data Extraction** – The [`extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/extractor.py) module uses these models to make authenticated API requests and enrich the data
3. **Storage** – The store implementation converts the extracted data into dictionaries and persists via the `XhsNote` ORM

This separation of concerns ensures that URL validation occurs early in the pipeline, while the heavy lifting of content storage happens through a clean abstraction layer.

## Practical Implementation Examples

The following examples demonstrate how to instantiate these models and use them in the XHS crawling workflow.

### Parsing a Note URL

To extract structured data from a raw XHS note URL:

```python
from model.m_xiaohongshu import NoteUrlInfo
from media_platform.xhs.help import parse_note_info_from_note_url

raw_url = "https://www.xiaohongshu.com/explore/66fad51c000000001b0224b8?xsec_token=AB3rO-QopW5sgrJ41GwN01WCXh6yWPxjSoFI9D5JIMgKw=&xsec_source=pc_search"
note_info: NoteUrlInfo = parse_note_info_from_note_url(raw_url)

print(note_info.note_id)      # → 66fad51c000000001b0224b8

print(note_info.xsec_token)   # → AB3rO‑QopW5sgrJ41GwN01...

print(note_info.xsec_source)  # → pc_search

```

### Parsing a Creator URL

Similarly, extract creator information from profile URLs:

```python
from model.m_xiaohongshu import CreatorUrlInfo
from media_platform.xhs.help import parse_creator_info_from_url

creator_url = "https://www.xiaohongshu.com/user/profile/5eb8e1d400000000010075ae?xsec_token=AB1nWBKCo1vE2HEkfoJUOi5B6BE5n7wVrbdpHoWIj5xHw=&xsec_source=pc_feed"
creator_info: CreatorUrlInfo = parse_creator_info_from_url(creator_url)

print(creator_info.user_id)   # → 5eb8e1d400000000010075ae

print(creator_info.xsec_token)   # token string (may be empty)

print(creator_info.xsec_source)  # → pc_feed

```

### Storing Scraped Content

Once extracted, store the content using the database implementation:

```python
from store.xhs._store_impl import XhsDbStoreImplement
import asyncio

async def store_note():
    db_store = XhsDbStoreImplement()
    note_dict = {
        "note_id": note_info.note_id,
        "creator_hash": "hashed_creator_id",
        "nickname": "匿名用户",
        "title": "示例笔记标题",
        "desc": "笔记正文内容",
        "image_list": ["https://.../image1.png", "https://.../image2.png"],
        "tag_list": ["#标签1", "#标签2"],
        "xsec_token": note_info.xsec_token,
        # ... other fields required by XhsNote model

    }
    await db_store.store_content(note_dict)

asyncio.run(store_note())

```

## Summary

The MediaCrawler repository implements a layered approach to XHS data modeling:

- **Pydantic models** (`NoteUrlInfo`, `CreatorUrlInfo`) in [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py) handle URL parsing and token validation
- **Helper functions** in [`media_platform/xhs/help.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/help.py) instantiate these models from raw URLs
- **ORM layer** (`XhsNote` in [`database/models.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/models.py)) provides persistent storage for scraped content
- **Storage implementations** in [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) bridge the gap between extracted dictionaries and the database

This architecture ensures type safety during the extraction phase while maintaining flexibility for downstream storage operations.

## Frequently Asked Questions

### What is the difference between NoteUrlInfo and CreatorUrlInfo?

**NoteUrlInfo** extracts data from individual post URLs containing `note_id` parameters, while **CreatorUrlInfo** handles user profile URLs containing `user_id` parameters. Both models validate `xsec_token` and `xsec_source` fields, but target different XHS endpoints. The `NoteUrlInfo` model feeds the content extraction pipeline, whereas `CreatorUrlInfo` supports user-specific crawling operations.

### Where does MediaCrawler store the actual scraped XHS content?

Scraped content persists through the **XhsNote** ORM class defined in [`database/models.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/models.py). This model maps to a relational database table and stores comprehensive note data including titles, descriptions, image lists, and creator metadata. The [`store/xhs/_store_impl.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/store/xhs/_store_impl.py) module handles the actual insertion logic, supporting both database and file-based (CSV/JSON) storage backends.

### How does the extractor use these data models?

The extractor in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) consumes `NoteUrlInfo` instances to make authenticated API requests. It uses the `note_id` and `xsec_token` fields from these models to construct proper request headers and parameters. After fetching raw data from the XHS API, the extractor transforms the response into dictionaries that match the schema expected by the `XhsNote` ORM.

### Are these data models specific to Xiaohongshu only?

Yes, the `NoteUrlInfo` and `CreatorUrlInfo` models in [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py) are specifically designed for XHS URL patterns and authentication token structures. MediaCrawler maintains separate model files for different platforms (e.g., Douyin, Weibo) in the `model/` directory, each with platform-specific validation logic and field requirements tailored to their respective APIs.