Xiaohongshu (XHS) Crawler Limitations and Known Issues in MediaCrawler
The MediaCrawler Xiaohongshu implementation has nine documented limitations including hard-coded pagination, QR-code-only authentication, two-level comment depth, and minimal rate-limit handling that restrict its use for large-scale production scraping.
The Xiaohongshu (小红书) crawler in the NanmiCoder/MediaCrawler repository provides functional public-data harvesting, but several architectural constraints and implementation bugs limit its reliability. Understanding these Xiaohongshu crawler limitations is essential before deploying the tool for any serious data collection workflow.
Hard-Coded Pagination and Result Limits
Fixed Page Size of 20 Items
The crawler assumes every note-list page returns exactly 20 items, a value hard-coded in media_platform/xhs/core.py at line 132:
# From media_platform/xhs/core.py (line 132)
xhs_limit_count = 20
If Xiaohongshu changes its API page size, the crawler will either miss items or request redundant pages, potentially triggering platform rate limits. This constraint affects all search and user-profile crawls.
Maximum Notes Capped by Configuration
The crawler enforces a hard stop after reaching CRAWLER_MAX_NOTES_COUNT, checked at lines 133-134 in media_platform/xhs/core.py. The default limit is 200 notes as defined in config/base_config.py.
To scrape more data, you must manually increase the configuration value:
# In config/base_config.py
CRAWLER_MAX_NOTES_COUNT = 500 # Raise limit if you need more items
Increasing this value raises the risk of IP throttling or account suspension, as the platform may flag high-volume requests as automated traffic.
Authentication Constraints
QR-Code Login Only
The current implementation only supports QR-code authentication via the --lt qrcode flag. Password and SMS login flows are not exposed in the CLI interface, as seen in main.py usage examples.
Running the crawler requires:
uv run main.py --platform xhs --lt qrcode --type search
This design means automation scripts must:
- Maintain a Chrome instance with remote debugging enabled
- Allow manual QR-code scanning whenever sessions expire
Cookie Expiration Handling
Cookies are stored in the Chrome profile directory managed by media_platform/xhs/login.py. When these cookies become stale—typically after a few days of inactivity—the crawler fails to sign requests and requires complete re-authentication. There is no automatic token refresh mechanism.
Data Extraction Limitations
Two-Level Comment Depth Only
The extractor in media_platform/xhs/extractor.py retrieves only the primary comment list and a single level of replies. Deeper nested comment threads are ignored entirely.
Users requiring full conversation threads must implement recursive parsing themselves by extending the extractor logic. This limitation significantly impacts social network analysis and sentiment research use cases.
No Access to Private or Paid Content
The crawler filters out any content behind authentication walls, including "付费笔记" (paid notes). According to the URL filtering logic in media_platform/xhs/extractor.py, attempting to scrape paid content returns empty results without error warnings. Only publicly available notes are accessible.
Rate Limiting and Infrastructure Issues
Minimal Rate-Limit Handling
The media_platform/xhs/api_limits.py module implements basic retry logic with fixed sleep intervals. There is no exponential back-off, adaptive throttling, or intelligent retry queues.
Heavy or parallel crawling operations quickly exhaust the platform's hidden request quotas, resulting in HTTP 429 errors. The current implementation simply sleeps and retries, making it unsuitable for high-throughput extraction.
Limited Proxy Rotation Support
Although config/base_config.py defines an ENABLE_PROXY_POOL flag, the XHS crawler does not currently utilize it. Proxy rotation must be implemented externally by the user, as the core XHS request logic bypasses the IP-pool integration present in other platform crawlers.
This limitation makes the crawler vulnerable to IP-based blocking during extended scraping sessions.
Signing and Cryptographic Bugs
Fragile xhshow Library Patch
The xhshow algorithm used for request signing contains a bug that miscalculates the a3_hash value. MediaCrawler patches this in media_platform/xhs/playwright_sign.py at lines 36-46:
# Patch implementation in media_platform/xhs/playwright_sign.py (lines 36-46)
# Fixes a3_hash calculation from upstream xhshow library
While this patch currently works, any upstream changes to the xhshow library or Xiaohongshu's signing algorithm could reintroduce the bug, breaking the crawler's ability to generate valid requests.
Summary
- Hard-coded pagination (
xhs_limit_count = 20) inmedia_platform/xhs/core.pyassumes fixed page sizes that may change without warning. - QR-code-only authentication requires manual intervention and frequent re-authentication when cookies expire.
- Two-level comment depth in
media_platform/xhs/extractor.pylimits social graph analysis capabilities. - Minimal rate-limit handling lacks exponential back-off, risking quick HTTP 429 blocks.
- No proxy pool integration despite the configuration flag existing in
config/base_config.py. - Fragile cryptographic patching of the
xhshowlibrary may break with platform updates.
Frequently Asked Questions
Why does the Xiaohongshu crawler stop at 200 notes?
The crawler respects the CRAWLER_MAX_NOTES_COUNT setting defined in config/base_config.py. To scrape more content, increase this value in the configuration file, though higher values increase the risk of rate limiting.
Can I use password login instead of QR codes?
No. As implemented in media_platform/xhs/login.py and reflected in main.py CLI examples, the crawler only supports the --lt qrcode login type. Password and SMS authentication flows are not exposed.
How do I handle comment replies in Xiaohongshu posts?
The current extractor in media_platform/xhs/extractor.py only retrieves primary comments and one level of replies. Full thread extraction requires extending the extractor with recursive parsing logic.
What happens when my cookies expire?
The crawler stores session cookies in the Chrome profile. When they become stale (typically after a few days), you must re-run the QR-code authentication flow. There is no automatic session renewal mechanism.
Is there any way to rotate proxies for the XHS crawler?
Not natively. While ENABLE_PROXY_PROXY exists in config/base_config.py, the XHS crawler implementation does not currently utilize the IP pool. Users must implement their own proxy rotation at the network level.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →