Xiaohongshu (XHS) Crawler Limitations and Known Issues in MediaCrawler

The MediaCrawler Xiaohongshu implementation has nine documented limitations including hard-coded pagination, QR-code-only authentication, two-level comment depth, and minimal rate-limit handling that restrict its use for large-scale production scraping.

The Xiaohongshu (小红书) crawler in the NanmiCoder/MediaCrawler repository provides functional public-data harvesting, but several architectural constraints and implementation bugs limit its reliability. Understanding these Xiaohongshu crawler limitations is essential before deploying the tool for any serious data collection workflow.

Hard-Coded Pagination and Result Limits

Fixed Page Size of 20 Items

The crawler assumes every note-list page returns exactly 20 items, a value hard-coded in media_platform/xhs/core.py at line 132:


# From media_platform/xhs/core.py (line 132)

xhs_limit_count = 20

If Xiaohongshu changes its API page size, the crawler will either miss items or request redundant pages, potentially triggering platform rate limits. This constraint affects all search and user-profile crawls.

Maximum Notes Capped by Configuration

The crawler enforces a hard stop after reaching CRAWLER_MAX_NOTES_COUNT, checked at lines 133-134 in media_platform/xhs/core.py. The default limit is 200 notes as defined in config/base_config.py.

To scrape more data, you must manually increase the configuration value:


# In config/base_config.py

CRAWLER_MAX_NOTES_COUNT = 500  # Raise limit if you need more items

Increasing this value raises the risk of IP throttling or account suspension, as the platform may flag high-volume requests as automated traffic.

Authentication Constraints

QR-Code Login Only

The current implementation only supports QR-code authentication via the --lt qrcode flag. Password and SMS login flows are not exposed in the CLI interface, as seen in main.py usage examples.

Running the crawler requires:

uv run main.py --platform xhs --lt qrcode --type search

This design means automation scripts must:

  • Maintain a Chrome instance with remote debugging enabled
  • Allow manual QR-code scanning whenever sessions expire

Cookies are stored in the Chrome profile directory managed by media_platform/xhs/login.py. When these cookies become stale—typically after a few days of inactivity—the crawler fails to sign requests and requires complete re-authentication. There is no automatic token refresh mechanism.

Data Extraction Limitations

Two-Level Comment Depth Only

The extractor in media_platform/xhs/extractor.py retrieves only the primary comment list and a single level of replies. Deeper nested comment threads are ignored entirely.

Users requiring full conversation threads must implement recursive parsing themselves by extending the extractor logic. This limitation significantly impacts social network analysis and sentiment research use cases.

No Access to Private or Paid Content

The crawler filters out any content behind authentication walls, including "付费笔记" (paid notes). According to the URL filtering logic in media_platform/xhs/extractor.py, attempting to scrape paid content returns empty results without error warnings. Only publicly available notes are accessible.

Rate Limiting and Infrastructure Issues

Minimal Rate-Limit Handling

The media_platform/xhs/api_limits.py module implements basic retry logic with fixed sleep intervals. There is no exponential back-off, adaptive throttling, or intelligent retry queues.

Heavy or parallel crawling operations quickly exhaust the platform's hidden request quotas, resulting in HTTP 429 errors. The current implementation simply sleeps and retries, making it unsuitable for high-throughput extraction.

Limited Proxy Rotation Support

Although config/base_config.py defines an ENABLE_PROXY_POOL flag, the XHS crawler does not currently utilize it. Proxy rotation must be implemented externally by the user, as the core XHS request logic bypasses the IP-pool integration present in other platform crawlers.

This limitation makes the crawler vulnerable to IP-based blocking during extended scraping sessions.

Signing and Cryptographic Bugs

Fragile xhshow Library Patch

The xhshow algorithm used for request signing contains a bug that miscalculates the a3_hash value. MediaCrawler patches this in media_platform/xhs/playwright_sign.py at lines 36-46:


# Patch implementation in media_platform/xhs/playwright_sign.py (lines 36-46)

# Fixes a3_hash calculation from upstream xhshow library

While this patch currently works, any upstream changes to the xhshow library or Xiaohongshu's signing algorithm could reintroduce the bug, breaking the crawler's ability to generate valid requests.

Summary

  • Hard-coded pagination (xhs_limit_count = 20) in media_platform/xhs/core.py assumes fixed page sizes that may change without warning.
  • QR-code-only authentication requires manual intervention and frequent re-authentication when cookies expire.
  • Two-level comment depth in media_platform/xhs/extractor.py limits social graph analysis capabilities.
  • Minimal rate-limit handling lacks exponential back-off, risking quick HTTP 429 blocks.
  • No proxy pool integration despite the configuration flag existing in config/base_config.py.
  • Fragile cryptographic patching of the xhshow library may break with platform updates.

Frequently Asked Questions

Why does the Xiaohongshu crawler stop at 200 notes?

The crawler respects the CRAWLER_MAX_NOTES_COUNT setting defined in config/base_config.py. To scrape more content, increase this value in the configuration file, though higher values increase the risk of rate limiting.

Can I use password login instead of QR codes?

No. As implemented in media_platform/xhs/login.py and reflected in main.py CLI examples, the crawler only supports the --lt qrcode login type. Password and SMS authentication flows are not exposed.

How do I handle comment replies in Xiaohongshu posts?

The current extractor in media_platform/xhs/extractor.py only retrieves primary comments and one level of replies. Full thread extraction requires extending the extractor with recursive parsing logic.

What happens when my cookies expire?

The crawler stores session cookies in the Chrome profile. When they become stale (typically after a few days), you must re-run the QR-code authentication flow. There is no automatic session renewal mechanism.

Is there any way to rotate proxies for the XHS crawler?

Not natively. While ENABLE_PROXY_PROXY exists in config/base_config.py, the XHS crawler implementation does not currently utilize the IP pool. Users must implement their own proxy rotation at the network level.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →