How to Configure MediaCrawler for International Xiaohongshu (rednote.com) vs the Domestic Version
Configure MediaCrawler for international or domestic Xiaohongshu crawling by editing two constants in config/xhs_config.py — BASE_URL and COOKIE_DOMAIN — which control whether requests target www.xiaohongshu.com or www.rednote.com.
MediaCrawler treats Xiaohongshu (小红书) as a pluggable platform whose request routing, authentication handling, and cookie management are driven entirely by configuration values. Whether you need to crawl the mainland China service (xiaohongshu.com) or the international Rednote service (rednote.com), the switch requires only a single file change with no code modifications elsewhere.
Where Xiaohongshu Configuration Lives
The platform-specific configuration is isolated in config/xhs_config.py. This file is imported by the request builder in model/m_xiaohongshu.py and used at runtime to construct URLs, validate cookies, and handle authentication tokens.
| Setting | Purpose | Default Value |
|---|---|---|
BASE_URL |
Root hostname for all API and web requests | https://www.xiaohongshu.com |
COOKIE_DOMAIN |
Domain scope for session cookies | xiaohongshu.com |
These constants propagate through the codebase:
model/m_xiaohongshu.py— Builds note URLs, search endpoints, and user profile links usingBASE_URLstore/xhs/_store_impl.py— Persists crawled data; storage paths remain unchanged regardless of target hostwebui/src/components/config/CrawlerConfigPanel.tsx— Validates that submitted URLs contain requiredxsec_tokenparameters
Step-by-Step: Switch Between Domestic and International
1. Open the Configuration File
# Repository root
nano config/xhs_config.py
2. Modify BASE_URL for Your Target Service
# config/xhs_config.py
# --- Domestic (Mainland China) ---
BASE_URL = "https://www.xiaohongshu.com"
# --- International (Rednote) ---
# Uncomment the line below and comment the domestic line above
# BASE_URL = "https://www.rednote.com"
3. Update COOKIE_DOMAIN to Match
# config/xhs_config.py
# --- Domestic ---
COOKIE_DOMAIN = "xiaohongshu.com"
# --- International ---
# COOKIE_DOMAIN = "rednote.com"
4. Save and Verify
No further changes are required. The model/m_xiaohongshu.py module constructs URLs dynamically:
# model/m_xiaohongshu.py — runtime URL construction
from config.xhs_config import BASE_URL
def build_note_url(note_id: str) -> str:
"""Construct full note URL using configured base host."""
return f"{BASE_URL}/note/{note_id}"
def build_explore_url(note_id: str, xsec_token: str) -> str:
"""Build explore page with required security token."""
return f"{BASE_URL}/explore/{note_id}?xsec_token={xsec_token}"
Authentication Requirements for Both Versions
Both xiaohongshu.com and rednote.com enforce identical authentication mechanisms:
xsec_token— Required query parameter on all note URLs- Cookie-based session — Must match the configured
COOKIE_DOMAIN
The web UI enforces token presence before starting crawls. As implemented in webui/src/components/config/CrawlerConfigPanel.tsx, submissions lacking xsec_token trigger a validation warning regardless of which host is configured.
Storage Behavior Across Versions
Downloaded media is organized consistently regardless of target host:
# store/xhs/_store_impl.py — path construction
self.image_store_path = f"{config.SAVE_DATA_PATH}/xhs/images"
self.video_store_path = f"{config.SAVE_DATA_PATH}/xhs/videos"
Data lands in data/xhs/ subdirectories by content type, not by origin domain. If you crawl both services, implement separate SAVE_DATA_PATH values in config/base_config.py to isolate datasets.
Running the Crawler with International Configuration
# Target a public note on rednote.com
python -m media_crawler run \
--platform xhs \
--url "https://www.rednote.com/explore/123456?xsec_token=abcdef"
Verify correct host resolution in log output — successful initialization shows the active BASE_URL during platform setup.
Summary
- Single-file configuration — All host-specific behavior lives in
config/xhs_config.py - Two constants control everything —
BASE_URLandCOOKIE_DOMAIN - No code changes needed —
model/m_xiaohongshu.pyconsumes configuration at runtime - Authentication unchanged —
xsec_tokenrequirement applies to both hosts - Storage paths consistent — Separate data sets via
base_config.pyif crawling both services
Frequently Asked Questions
What happens if I forget to update COOKIE_DOMAIN when switching to rednote.com?
Session cookies will fail to attach to requests. The crawler authenticates against rednote.com endpoints but sends cookies scoped to xiaohongshu.com, causing immediate authentication errors on protected endpoints. Both constants must match your target service.
Does the international version use different API endpoints or response formats?
No. Both services expose identical URL paths and JSON structures. The model/m_xiaohongshu.py request builder prefixes the same relative paths (/note/, /explore/, /api/sns/web/) with whichever BASE_URL is configured, and response parsing logic remains unchanged.
Can I crawl both domestic and international versions simultaneously?
Not from a single process instance. BASE_URL is loaded once at startup. To crawl both services concurrently, run separate crawler instances with isolated configuration files, or modify base_config.py to use distinct SAVE_DATA_PATH directories and launch independent processes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →