# How NanmiCoder/MediaCrawler Handles Xiaohongshu (XHS) Platform-Specific Implementation

> Discover how NanmiCoder/MediaCrawler manages Xiaohongshu XHS platform specifics. Explore its modular design featuring typed data models authentication, crypto signing, and three crawl modes.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: deep-dive
- Published: 2026-07-03

---

**NanmiCoder/MediaCrawler isolates all Xiaohongshu-specific crawling logic in a self-contained module under `media_platform/xhs/`, implementing a layered architecture that covers typed data models, multiple authentication strategies, cryptographic request signing, and three distinct crawl modes.**

The NanmiCoder/MediaCrawler repository treats each social media platform as a standalone package. For Xiaohongshu (XHS), the implementation resides entirely within `media_platform/xhs/` and encapsulates everything from QR-code login flows to cryptographically signed HTTP requests, ensuring clean separation from other platform implementations like TikTok or Weibo.

## Architecture Overview

The XHS module follows a strict layered architecture where each layer has a single responsibility. This design allows developers to modify authentication logic or update API endpoints without affecting the core crawling engine.

**Data Layer**: Typed containers for URLs and creator metadata live in [`model/m_xiaohongshu.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/model/m_xiaohongshu.py), defining `NoteUrlInfo` and `CreatorUrlInfo` dataclasses that validate input URLs before processing.

**Platform Package**: The `media_platform/xhs/` directory contains all XHS-specific logic including login handlers, the API client, crawler orchestration, and HTML extractors.

**Persistence**: Retrieved notes, comments, and media are stored via the `store/xhs` module, invoked directly from the crawler core.

## Authentication and Login Strategies

The XHS implementation supports three distinct login methods with robust retry logic, all implemented in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py).

**`XiaoHongShuLogin.begin`** serves as the entry point that delegates to specific strategies based on `config.LOGIN_TYPE`:

- **`login_by_qrcode`**: Generates a QR code for mobile app scanning
- **`login_by_mobile`**: Uses SMS verification for phone-based authentication  
- **`login_by_cookies`**: Resumes an existing session from stored cookies

The login class updates browser cookies upon successful authentication, which are then imported into the API client via `client.update_cookies`.

## API Client with Request Signing

Every request to Xiaohongshu’s backend requires cryptographic headers (`X-S`, `X-T`, etc.) to prevent unauthorized scraping. The `XiaoHongShuClient` class in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) handles this automatically.

**`XiaoHongShuClient._pre_headers`** invokes the `sign_with_xhshow` function (located in [`media_platform/xhs/xhs_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/xhs_sign.py)) to generate signed headers before each HTTP request. The client uses `httpx` for async HTTP communication and maintains a cookie dictionary synchronized with the browser context.

The client also provides high-level methods for data retrieval:

- `get_note_by_keyword`: Searches notes by keyword with pagination
- `get_note_info`: Fetches detailed metadata for a specific note ID
- `get_creator_info`: Retrieves creator profile data
- `get_all_notes_by_creator`: Paginates through all notes from a specific creator

## Proxy Rotation and Resilience

To avoid IP-based rate limiting, the client integrates `ProxyRefreshMixin` which checks proxy expiration before each request via `await self._refresh_proxy_if_expired()`.

If a proxy fails, the client automatically fetches a new IP from the pool initialized by `self.init_proxy_pool(proxy_ip_pool)`. This mechanism ensures continuous crawling even when individual IPs encounter blocks or bans.

## Crawler Core and Execution Modes

The `XiaoHongShuCrawler` class in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) orchestrates the entire crawling process. It launches a Playwright browser (standard or CDP mode), manages login fallback, and executes one of three crawl modes:

**Search Mode**: The `search()` method iterates over `config.KEYWORDS`, builds a `search_id`, and calls `client.get_note_by_keyword`. Returned note IDs are processed concurrently using `asyncio.Semaphore` to fetch full details, comments, and media while respecting `config.CRAWLER_MAX_SLEEP_SEC` throttling.

**Detail Mode**: `get_specified_notes()` parses URLs from `config.XHS_SPECIFIED_NOTE_URL_LIST` (using `NoteUrlInfo`), then fetches comprehensive note details and associated comments.

**Creator Mode**: `get_creators_and_notes()` processes creator URLs (`CreatorUrlInfo`), retrieves creator profiles via `client.get_creator_info`, then walks through all creator content using `client.get_all_notes_by_creator`.

## Practical Code Examples

### Running a Search Crawl

```python
import asyncio
from media_platform.xhs.core import XiaoHongShuCrawler

async def main():
    # Configuration read from .env (LOGIN_TYPE, KEYWORDS, etc.)

    crawler = XiaoHongShuCrawler()
    await crawler.start()  # Launches browser, authenticates, executes search

    await crawler.close()  # Cleanup Playwright resources

if __name__ == "__main__":
    asyncio.run(main())

```

### Manual API Client Usage

```python
import asyncio
from media_platform.xhs.client import XiaoHongShuClient
from tools import utils

async def demo():
    headers = {
        "Cookie": "web_session=abcdef12345;",
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)..."
    }
    
    client = XiaoHongShuClient(
        proxy=None,
        headers=headers,
        playwright_page=None,
        cookie_dict=utils.convert_str_cookie_to_dict(headers["Cookie"]),
    )
    
    # Search for travel-related content

    resp = await client.get_note_by_keyword(keyword="travel", page=1)
    print(resp["items"][0]["note_card"]["note_id"])

asyncio.run(demo())

```

### Extending with Custom Login

```python

# In media_platform/xhs/login.py

class XiaoHongShuLogin(AbstractLogin):
    async def login_by_api_token(self):
        """Experimental token-based authentication."""
        token = config.XHS_API_TOKEN
        self.headers["Authorization"] = f"Bearer {token}"
        client = XiaoHongShuClient(
            headers=self.headers,
            playwright_page=self.context_page,
            cookie_dict={}
        )
        if await client.pong():
            utils.logger.info("[XiaoHongShuLogin] API token login successful")
        else:
            raise RuntimeError("Invalid API token")

```

## Summary

- **Modular Design**: All XHS logic is isolated in `media_platform/xhs/`, preventing cross-platform contamination.
- **Robust Authentication**: Three login strategies (QR-code, SMS, cookies) with automatic session validation via `client.pong()`.
- **Cryptographic Security**: Automatic header signing via `sign_with_xhshow` in `XiaoHongShuClient._pre_headers`.
- **Flexible Crawling**: Three execution modes (search, detail, creator) support diverse data collection requirements.
- **Production Resilience**: Built-in proxy rotation via `ProxyRefreshMixin` and concurrency control through async semaphores.

## Frequently Asked Questions

### How does MediaCrawler handle Xiaohongshu authentication?

MediaCrawler implements three authentication strategies in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py): QR-code scanning, mobile SMS verification, and cookie-based session resumption. The `XiaoHongShuLogin.begin` method selects the appropriate strategy based on the `LOGIN_TYPE` configuration value, automatically retrying failed attempts and updating the API client cookies upon success.

### What are the X-S and X-T headers required for XHS requests?

Xiaohongshu requires cryptographic signatures in HTTP headers to prevent unauthorized API access. The `X-S` and `X-T` headers are generated by the `sign_with_xhshow` function within `XiaoHongShuClient._pre_headers` (located in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py)). These headers are computed from request parameters and timestamp data to validate the legitimacy of each API call.

### How does the crawler manage proxy rotation for XHS?

The `XiaoHongShuClient` class inherits from `ProxyRefreshMixin`, which checks proxy expiration before each request via `_refresh_proxy_if_expired()`. If the current proxy is banned or expired, the client automatically retrieves a new IP from the pool initialized through `init_proxy_pool(proxy_ip_pool)`, ensuring continuous operation without manual intervention.

### What crawl modes are available for Xiaohongshu data collection?

The implementation supports three distinct modes controlled by `XiaoHongShuCrawler`: **Search mode** (`search()`) discovers content by keywords; **Detail mode** (`get_specified_notes()`) fetches specific URLs provided in configuration; and **Creator mode** (`get_creators_and_notes()`) extracts all content from specified creator profiles. Each mode uses the same underlying API client but applies different orchestration logic for data retrieval.