# Main Dependencies for the Xiaohongshu (XHS) Crawler Module in MediaCrawler

> Discover the essential dependencies for the Xiaohongshu crawler in MediaCrawler. Learn how httpx, playwright, tenacity, and xhshow power its functionality.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: deep-dive
- Published: 2026-07-03

---

**The Xiaohongshu crawler relies on four core third-party libraries—httpx, playwright, tenacity, and xhshow—to handle async HTTP requests, browser automation, retry logic, and API signature generation.**

The MediaCrawler repository implements a dedicated module for scraping Xiaohongshu (Xiaohongshu/RED) content under `media_platform/xhs/`. Understanding the main dependencies for the Xiaohongshu crawler module is essential for maintenance, troubleshooting, and extending the scraper's capabilities. These dependencies are explicitly pinned in the repository's [`requirements.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/requirements.txt) and imported throughout the module's client, login, and core orchestration files.

## Core Third-Party Dependencies

### httpx for Async HTTP Requests

**httpx** (version `0.28.1`) serves as the primary async HTTP client for the XHS module. In [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) (lines 25-27), the library is wrapped by `tools.httpx_util.make_async_client` to create a configured client instance that supports proxy rotation and timeout handling. The `XiaoHongShuClient` class uses this client for all API calls to Xiaohongshu's endpoints, including fetching note details and search results.

### Playwright for Browser Automation

**playwright** (version `>=1.61.0`) drives the headless Chromium instance required for the authentication flow and dynamic data extraction. The module imports `BrowserContext` and `Page` objects in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) (lines 26-28) and [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) (lines 26-32). Playwright handles QR-code scanning, phone number login, and cookie-based session persistence that pure HTTP requests cannot achieve alone.

### Tenacity for Resilient Retries

**tenacity** (version `8.2.2`) supplies the `@retry` decorator used throughout the module to handle transient network failures. The login polling mechanism in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) (line 51) and the API request wrapper in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) (line 15) both implement tenacity's exponential backoff strategies. This ensures the crawler automatically retries failed requests before raising exceptions.

### xhshow for API Signature Generation

**xhshow** (version `>=0.1.9`) is a specialized "pure-algorithm" library that reproduces the signature generation used by Xiaohongshu's official web API. The `sign_with_xhshow` helper function in [`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py) (lines 19-21 and 113-119) generates the mandatory `X-S`, `X-T`, `x-s-common`, and `X-B3-Traceid` headers. Without these cryptographically signed headers, the Xiaohongshu API rejects requests with authentication errors.

## Standard Library and Internal Utilities

Beyond external packages, the module relies on Python's standard library and internal project utilities:

- **asyncio**: Drives the entire asynchronous workflow, from concurrent page fetches to semaphore-protected comment scraping loops in [`client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/client.py) and [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py).
- **urllib.parse**: Handles query string construction in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) (lines 22-24), ensuring commas remain unescaped to match browser encoding rules.
- **Internal modules**: The crawler imports `config` for runtime options (e.g., `XHS_INTERNATIONAL`, `LOGIN_TYPE`), `base.base_crawler` for the `AbstractApiClient` interface, `proxy.proxy_mixin` for automatic proxy refresh, and `store.xhs` for database persistence.

## Dependency Installation

All dependencies are declared in the top-level [`requirements.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/requirements.txt) file:

```text
httpx==0.28.1
playwright>=1.61.0
tenacity==8.2.2
xhshow>=0.1.9   # signature algorithm for Xiaohongshu

```

Install these specifically for the Xiaohongshu module using:

```bash
pip install httpx==0.28.1 playwright>=1.61.0 tenacity==8.2.2 xhshow>=0.1.9
playwright install chromium

```

## Code Implementation Examples

### Creating a Signed API Client

This example demonstrates initializing the browser context, performing login, and constructing a signed client:

```python
from media_platform.xhs.client import XiaoHongShuClient
from media_platform.xhs.login import XiaoHongShuLogin
from playwright.async_api import async_playwright

async def init_xhs():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        context = await browser.new_context()
        page = await context.new_page()
        await page.goto("https://www.xiaohongshu.com")

        # Perform login (qrcode, phone, or cookie)

        login = XiaoHongShuLogin(
            login_type="qrcode",
            browser_context=context,
            context_page=page,
            cookie_str="...",
        )
        await login.begin()

        # Build a client with signed headers

        client = XiaoHongShuClient(
            headers={"User-Agent": "my-crawler"},
            playwright_page=page,
            cookie_dict={"web_session": "…"},
        )
        # Example API call – get note details

        note = await client.get_note_by_id("66fad51c000000001b0224b8")
        print(note)

```

### Running a Full Search Crawl

Execute the high-level crawler orchestration defined in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py):

```python
from media_platform.xhs.core import XiaoHongShuCrawler
import asyncio

async def run_search():
    crawler = XiaoHongShuCrawler()
    await crawler.start()   # respects config.CRAWLER_TYPE="search"

asyncio.run(run_search())

```

## Summary

- **httpx** provides the async HTTP foundation with proxy support in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py).
- **playwright** enables browser automation for login and dynamic content in [`login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/login.py) and [`core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/core.py).
- **tenacity** implements retry logic with decorators on network operations throughout the module.
- **xhshow** generates required cryptographic signatures (`X-S`, `X-T` headers) via [`playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/playwright_sign.py).
- All dependencies are pinned in [`requirements.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/requirements.txt) and must be installed alongside Playwright's Chromium browser binary.

## Frequently Asked Questions

### What version of Playwright is required for the XHS crawler?

MediaCrawler specifies `playwright>=1.61.0` in [`requirements.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/requirements.txt). This version provides the `BrowserContext` and `Page` APIs used in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) (lines 26-28) and [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) (lines 26-32) to handle authentication flows and dynamic page interactions.

### Why does the Xiaohongshu crawler need the xhshow library?

Xiaohongshu's API requires cryptographically signed headers (`X-S`, `X-T`, `x-s-common`, `X-B3-Traceid`) to verify request legitimacy. The `xhshow` library (version `>=0.1.9`) reproduces the official signing algorithm. Without it, the `sign_with_xhshow` function in `media_platform/xhs/playwright_sign.py