# MediaCrawler Ethical Guidelines for Xiaohongshu (XHS): Responsible Web Scraping Practices

> Discover MediaCrawler ethical guidelines for Xiaohongshu scraping. Learn responsible practices and ensure compliance with XHS Terms of Service. Download MediaCrawler for ethical data collection.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: best-practices
- Published: 2026-07-03

---

**MediaCrawler explicitly prohibits commercial use and mandates compliance with Xiaohongshu’s Terms of Service through configuration files like [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) and rate-limiting settings in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).**

MediaCrawler is an open-source web crawling framework designed for research and educational purposes across multiple self-media platforms. When targeting **Xiaohongshu (XHS)**, the repository enforces strict ethical guidelines encoded directly in its Python configuration files to prevent abuse and ensure respectful data collection. Understanding these constraints is essential before executing any crawl operations against the platform.

## Core Ethical Guidelines for XHS Crawling

The repository ships a dedicated ethical-use policy in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) that governs all XHS interactions. These rules are not merely suggestions—they are architectural requirements built into the codebase.

### Non-Commercial Use Only

The project license strictly limits usage to **learning, research, and personal projects**. According to [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 10-12), any commercial redistribution or profit-generating services derived from scraped data are explicitly prohibited. This restriction applies specifically to XHS crawling operations and is reinforced by the repository’s disclaimer in [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md).

### Platform Terms and robots.txt Compliance

Before initiating any crawl, you must review Xiaohongshu’s Terms of Service and [`robots.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/robots.txt) directives. The ethical policy in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 11-13) requires users to respect these platform rules, ensuring the crawler does not access restricted endpoints or violate service agreements.

### No Large-Scale Scraping

The framework is designed for **modest, targeted data collection** rather than bulk harvesting. As noted in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 12-14), users should avoid massive request volumes that could overwhelm XHS infrastructure. Target only specific content using whitelisted URLs in `XHS_SPECIFIED_NOTE_URL_LIST` or `XHS_CREATOR_ID_LIST`.

### Request Throttling and Rate Limiting

To minimize server load and avoid IP blocking, MediaCrawler implements configurable delays between HTTP calls. The **`REQUEST_INTERVAL`** parameter in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) defaults to 1 second, allowing you to insert realistic delays that reduce detection risk. Adjusting this to 2.0 seconds or higher provides additional safety margins for sensitive crawling operations.

### Privacy and Legal Compliance

The crawler must never harvest private user data, harass individuals, or support unlawful activities. The ethical guidelines in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 14-16) explicitly forbid using scraped content for illegal purposes or redistributing it without creator permission.

## Key Configuration Files for Ethical Compliance

Several source files define the ethical boundaries for XHS operations:

- **[`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py)**: Contains XHS-specific ethical notices and URL whitelists (`XHS_SPECIFIED_NOTE_URL_LIST`, `XHS_CREATOR_ID_LIST`). This is where you declare specific allowed targets before crawling.

- **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)**: Houses global throttling controls including `REQUEST_INTERVAL` and `ENABLE_GET_COMMENTS`. These settings prevent aggressive crawling that could trigger platform rate limits.

- **[`media_platform/xhs/core.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/core.py)**: Implements the core HTTP client that respects the rate-limiting configurations defined in base settings.

- **[`media_platform/xhs/login.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/login.py)**: Manages QR-code and cookie-based authentication, reducing repeated login attempts that might flag automated behavior.

## Implementing Ethical XHS Crawling

The following examples demonstrate compliant usage patterns that respect MediaCrawler’s ethical framework.

### Search-Based Crawling with Rate Limiting

Execute keyword-based searches using the default 1-second interval to maintain respectful request frequency:

```bash

# Uses QR-code login with session caching to avoid repeated authentication

uv run main.py --platform xhs --lt qrcode --type search

```

### Targeted Detail Crawling

Crawl only specific notes defined in your configuration whitelist:

```python

# config/xhs_config.py - Define allowed targets explicitly

XHS_SPECIFIED_NOTE_URL_LIST = [
    "https://www.xiaohongshu.com/explore/64b95d01000000000c034587"
]

# Execute with detailed type parameter

# uv run main.py --platform xhs --lt qrcode --type detail

```

### Adjusting Throttling Parameters

Increase delays and disable comment paging to reduce server load:

```python

# config/base_config.py

ENABLE_GET_COMMENTS = False    # Disable if not needed

REQUEST_INTERVAL = 2.0         # 2 seconds between requests

```

## Summary

- **Commercial use is prohibited** by the license terms in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 10-12) and the project [`README.md`](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md).
- **Respect platform terms** by reviewing Xiaohongshu’s [`robots.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/robots.txt) and Terms of Service before crawling.
- **Implement rate limiting** using `REQUEST_INTERVAL` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) to prevent server overload.
- **Limit scope** to specific URLs listed in `XHS_SPECIFIED_NOTE_URL_LIST` rather than bulk scraping.
- **Never redistribute content** without explicit permission from original creators.

## Frequently Asked Questions

### Can I use MediaCrawler for XHS data to train commercial AI models?

No. The repository explicitly prohibits commercial use in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 10-12) and the README disclaimer. Training commercial AI models would violate the "learning and research only" restriction, regardless of whether you purchase the data or scrape it yourself.

### What happens if I set REQUEST_INTERVAL to 0.1 seconds?

While technically possible, setting `REQUEST_INTERVAL` below 1.0 seconds in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) violates the ethical guideline against overwhelming servers. Such aggressive crawling increases risk of IP bans and violates the "no large-scale scraping" rule in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 12-14).

### Is it legal to scrape public XHS posts for academic research?

MediaCrawler permits academic research usage, provided you comply with Xiaohongshu’s Terms of Service and configure `XHS_SPECIFIED_NOTE_URL_LIST` to target specific content. However, you must verify that your research institution’s data collection policies align with platform terms and local privacy laws.

### How do I ensure I'm not violating the robots.txt policy?

Review Xiaohongshu’s [`robots.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/robots.txt) file before crawling and configure `XHS_SPECIFIED_NOTE_URL_LIST` to avoid restricted paths. The ethical guidelines in [`config/xhs_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/xhs_config.py) (lines 11-13) require manual verification of platform rules, as the crawler does not automatically parse [`robots.txt`](https://github.com/NanmiCoder/MediaCrawler/blob/main/robots.txt) directives.