# Xiaohongshu (XHS) Crawler Limitations and Known Issues in MediaCrawler

> Explore Xiaohongshu crawler limitations in MediaCrawler including authentication, pagination, and rate-limiting issues. Learn how these restrict large-scale XHS data scraping.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: known-issues
- Published: 2026-07-03

---

**The MediaCrawler Xiaohongshu implementation has nine documented limitations including hard-coded pagination, QR-code-only authentication, two-level comment depth, and minimal rate-limit handling that restrict its use for large-scale production scraping.**

The Xiaohongshu (小红书) crawler in the [NanmiCoder/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler) repository provides functional public-data harvesting, but several architectural constraints and implementation bugs limit its reliability. Understanding these **Xiaohongshu crawler limitations** is essential before deploying the tool for any serious data collection workflow.

## Hard-Coded Pagination and Result Limits

### Fixed Page Size of 20 Items

The crawler assumes every note-list page returns exactly **20 items**, a value hard-coded in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) at line 132:

```python

# From media_platform/xhs/core.py (line 132)

xhs_limit_count = 20

```

If Xiaohongshu changes its API page size, the crawler will either miss items or request redundant pages, potentially triggering platform rate limits. This constraint affects all search and user-profile crawls.

### Maximum Notes Capped by Configuration

The crawler enforces a hard stop after reaching `CRAWLER_MAX_NOTES_COUNT`, checked at lines 133-134 in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py). The default limit is **200 notes** as defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).

To scrape more data, you must manually increase the configuration value:

```python

# In config/base_config.py

CRAWLER_MAX_NOTES_COUNT = 500  # Raise limit if you need more items

```

Increasing this value raises the risk of IP throttling or account suspension, as the platform may flag high-volume requests as automated traffic.

## Authentication Constraints

### QR-Code Login Only

The current implementation **only supports QR-code authentication** via the `--lt qrcode` flag. Password and SMS login flows are not exposed in the CLI interface, as seen in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) usage examples.

Running the crawler requires:

```bash
uv run main.py --platform xhs --lt qrcode --type search

```

This design means automation scripts must:
- Maintain a Chrome instance with remote debugging enabled
- Allow manual QR-code scanning whenever sessions expire

### Cookie Expiration Handling

Cookies are stored in the Chrome profile directory managed by [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py). When these cookies become stale—typically after a few days of inactivity—the crawler fails to sign requests and requires **complete re-authentication**. There is no automatic token refresh mechanism.

## Data Extraction Limitations

### Two-Level Comment Depth Only

The extractor in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) retrieves only the primary comment list and a single level of replies. Deeper nested comment threads are ignored entirely.

Users requiring full conversation threads must implement recursive parsing themselves by extending the extractor logic. This limitation significantly impacts social network analysis and sentiment research use cases.

### No Access to Private or Paid Content

The crawler filters out any content behind authentication walls, including **"付费笔记" (paid notes)**. According to the URL filtering logic in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py), attempting to scrape paid content returns empty results without error warnings. Only publicly available notes are accessible.

## Rate Limiting and Infrastructure Issues

### Minimal Rate-Limit Handling

The [`media_platform/xhs/api_limits.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/api_limits.py) module implements basic retry logic with fixed sleep intervals. There is **no exponential back-off**, adaptive throttling, or intelligent retry queues.

Heavy or parallel crawling operations quickly exhaust the platform's hidden request quotas, resulting in HTTP 429 errors. The current implementation simply sleeps and retries, making it unsuitable for high-throughput extraction.

### Limited Proxy Rotation Support

Although [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) defines an `ENABLE_PROXY_POOL` flag, the **XHS crawler does not currently utilize it**. Proxy rotation must be implemented externally by the user, as the core XHS request logic bypasses the IP-pool integration present in other platform crawlers.

This limitation makes the crawler vulnerable to IP-based blocking during extended scraping sessions.

## Signing and Cryptographic Bugs

### Fragile xhshow Library Patch

The `xhshow` algorithm used for request signing contains a bug that miscalculates the `a3_hash` value. MediaCrawler patches this in [`media_platform/xhs/playwright_sign.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/playwright_sign.py) at lines 36-46:

```python

# Patch implementation in media_platform/xhs/playwright_sign.py (lines 36-46)

# Fixes a3_hash calculation from upstream xhshow library

```

While this patch currently works, any upstream changes to the `xhshow` library or Xiaohongshu's signing algorithm could reintroduce the bug, breaking the crawler's ability to generate valid requests.

## Summary

- **Hard-coded pagination** (`xhs_limit_count = 20`) in [`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py) assumes fixed page sizes that may change without warning.
- **QR-code-only authentication** requires manual intervention and frequent re-authentication when cookies expire.
- **Two-level comment depth** in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) limits social graph analysis capabilities.
- **Minimal rate-limit handling** lacks exponential back-off, risking quick HTTP 429 blocks.
- **No proxy pool integration** despite the configuration flag existing in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).
- **Fragile cryptographic patching** of the `xhshow` library may break with platform updates.

## Frequently Asked Questions

### Why does the Xiaohongshu crawler stop at 200 notes?

The crawler respects the `CRAWLER_MAX_NOTES_COUNT` setting defined in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py). To scrape more content, increase this value in the configuration file, though higher values increase the risk of rate limiting.

### Can I use password login instead of QR codes?

No. As implemented in [`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py) and reflected in [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) CLI examples, the crawler only supports the `--lt qrcode` login type. Password and SMS authentication flows are not exposed.

### How do I handle comment replies in Xiaohongshu posts?

The current extractor in [`media_platform/xhs/extractor.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/extractor.py) only retrieves primary comments and one level of replies. Full thread extraction requires extending the extractor with recursive parsing logic.

### What happens when my cookies expire?

The crawler stores session cookies in the Chrome profile. When they become stale (typically after a few days), you must re-run the QR-code authentication flow. There is no automatic session renewal mechanism.

### Is there any way to rotate proxies for the XHS crawler?

Not natively. While `ENABLE_PROXY_PROXY` exists in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), the XHS crawler implementation does not currently utilize the IP pool. Users must implement their own proxy rotation at the network level.