MediaCrawler Ethical Guidelines for Xiaohongshu (XHS): Responsible Web Scraping Practices
MediaCrawler explicitly prohibits commercial use and mandates compliance with Xiaohongshu’s Terms of Service through configuration files like config/xhs_config.py and rate-limiting settings in config/base_config.py.
MediaCrawler is an open-source web crawling framework designed for research and educational purposes across multiple self-media platforms. When targeting Xiaohongshu (XHS), the repository enforces strict ethical guidelines encoded directly in its Python configuration files to prevent abuse and ensure respectful data collection. Understanding these constraints is essential before executing any crawl operations against the platform.
Core Ethical Guidelines for XHS Crawling
The repository ships a dedicated ethical-use policy in config/xhs_config.py that governs all XHS interactions. These rules are not merely suggestions—they are architectural requirements built into the codebase.
Non-Commercial Use Only
The project license strictly limits usage to learning, research, and personal projects. According to config/xhs_config.py (lines 10-12), any commercial redistribution or profit-generating services derived from scraped data are explicitly prohibited. This restriction applies specifically to XHS crawling operations and is reinforced by the repository’s disclaimer in README.md.
Platform Terms and robots.txt Compliance
Before initiating any crawl, you must review Xiaohongshu’s Terms of Service and robots.txt directives. The ethical policy in config/xhs_config.py (lines 11-13) requires users to respect these platform rules, ensuring the crawler does not access restricted endpoints or violate service agreements.
No Large-Scale Scraping
The framework is designed for modest, targeted data collection rather than bulk harvesting. As noted in config/xhs_config.py (lines 12-14), users should avoid massive request volumes that could overwhelm XHS infrastructure. Target only specific content using whitelisted URLs in XHS_SPECIFIED_NOTE_URL_LIST or XHS_CREATOR_ID_LIST.
Request Throttling and Rate Limiting
To minimize server load and avoid IP blocking, MediaCrawler implements configurable delays between HTTP calls. The REQUEST_INTERVAL parameter in config/base_config.py defaults to 1 second, allowing you to insert realistic delays that reduce detection risk. Adjusting this to 2.0 seconds or higher provides additional safety margins for sensitive crawling operations.
Privacy and Legal Compliance
The crawler must never harvest private user data, harass individuals, or support unlawful activities. The ethical guidelines in config/xhs_config.py (lines 14-16) explicitly forbid using scraped content for illegal purposes or redistributing it without creator permission.
Key Configuration Files for Ethical Compliance
Several source files define the ethical boundaries for XHS operations:
-
config/xhs_config.py: Contains XHS-specific ethical notices and URL whitelists (XHS_SPECIFIED_NOTE_URL_LIST,XHS_CREATOR_ID_LIST). This is where you declare specific allowed targets before crawling. -
config/base_config.py: Houses global throttling controls includingREQUEST_INTERVALandENABLE_GET_COMMENTS. These settings prevent aggressive crawling that could trigger platform rate limits. -
media_platform/xhs/core.py: Implements the core HTTP client that respects the rate-limiting configurations defined in base settings. -
media_platform/xhs/login.py: Manages QR-code and cookie-based authentication, reducing repeated login attempts that might flag automated behavior.
Implementing Ethical XHS Crawling
The following examples demonstrate compliant usage patterns that respect MediaCrawler’s ethical framework.
Search-Based Crawling with Rate Limiting
Execute keyword-based searches using the default 1-second interval to maintain respectful request frequency:
# Uses QR-code login with session caching to avoid repeated authentication
uv run main.py --platform xhs --lt qrcode --type search
Targeted Detail Crawling
Crawl only specific notes defined in your configuration whitelist:
# config/xhs_config.py - Define allowed targets explicitly
XHS_SPECIFIED_NOTE_URL_LIST = [
"https://www.xiaohongshu.com/explore/64b95d01000000000c034587"
]
# Execute with detailed type parameter
# uv run main.py --platform xhs --lt qrcode --type detail
Adjusting Throttling Parameters
Increase delays and disable comment paging to reduce server load:
# config/base_config.py
ENABLE_GET_COMMENTS = False # Disable if not needed
REQUEST_INTERVAL = 2.0 # 2 seconds between requests
Summary
- Commercial use is prohibited by the license terms in
config/xhs_config.py(lines 10-12) and the projectREADME.md. - Respect platform terms by reviewing Xiaohongshu’s
robots.txtand Terms of Service before crawling. - Implement rate limiting using
REQUEST_INTERVALinconfig/base_config.pyto prevent server overload. - Limit scope to specific URLs listed in
XHS_SPECIFIED_NOTE_URL_LISTrather than bulk scraping. - Never redistribute content without explicit permission from original creators.
Frequently Asked Questions
Can I use MediaCrawler for XHS data to train commercial AI models?
No. The repository explicitly prohibits commercial use in config/xhs_config.py (lines 10-12) and the README disclaimer. Training commercial AI models would violate the "learning and research only" restriction, regardless of whether you purchase the data or scrape it yourself.
What happens if I set REQUEST_INTERVAL to 0.1 seconds?
While technically possible, setting REQUEST_INTERVAL below 1.0 seconds in config/base_config.py violates the ethical guideline against overwhelming servers. Such aggressive crawling increases risk of IP bans and violates the "no large-scale scraping" rule in config/xhs_config.py (lines 12-14).
Is it legal to scrape public XHS posts for academic research?
MediaCrawler permits academic research usage, provided you comply with Xiaohongshu’s Terms of Service and configure XHS_SPECIFIED_NOTE_URL_LIST to target specific content. However, you must verify that your research institution’s data collection policies align with platform terms and local privacy laws.
How do I ensure I'm not violating the robots.txt policy?
Review Xiaohongshu’s robots.txt file before crawling and configure XHS_SPECIFIED_NOTE_URL_LIST to avoid restricted paths. The ethical guidelines in config/xhs_config.py (lines 11-13) require manual verification of platform rules, as the crawler does not automatically parse robots.txt directives.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →