How NanmiCoder/MediaCrawler Handles Xiaohongshu (XHS) Platform-Specific Implementation
NanmiCoder/MediaCrawler isolates all Xiaohongshu-specific crawling logic in a self-contained module under media_platform/xhs/, implementing a layered architecture that covers typed data models, multiple authentication strategies, cryptographic request signing, and three distinct crawl modes.
The NanmiCoder/MediaCrawler repository treats each social media platform as a standalone package. For Xiaohongshu (XHS), the implementation resides entirely within media_platform/xhs/ and encapsulates everything from QR-code login flows to cryptographically signed HTTP requests, ensuring clean separation from other platform implementations like TikTok or Weibo.
Architecture Overview
The XHS module follows a strict layered architecture where each layer has a single responsibility. This design allows developers to modify authentication logic or update API endpoints without affecting the core crawling engine.
Data Layer: Typed containers for URLs and creator metadata live in model/m_xiaohongshu.py, defining NoteUrlInfo and CreatorUrlInfo dataclasses that validate input URLs before processing.
Platform Package: The media_platform/xhs/ directory contains all XHS-specific logic including login handlers, the API client, crawler orchestration, and HTML extractors.
Persistence: Retrieved notes, comments, and media are stored via the store/xhs module, invoked directly from the crawler core.
Authentication and Login Strategies
The XHS implementation supports three distinct login methods with robust retry logic, all implemented in media_platform/xhs/login.py.
XiaoHongShuLogin.begin serves as the entry point that delegates to specific strategies based on config.LOGIN_TYPE:
login_by_qrcode: Generates a QR code for mobile app scanninglogin_by_mobile: Uses SMS verification for phone-based authenticationlogin_by_cookies: Resumes an existing session from stored cookies
The login class updates browser cookies upon successful authentication, which are then imported into the API client via client.update_cookies.
API Client with Request Signing
Every request to Xiaohongshu’s backend requires cryptographic headers (X-S, X-T, etc.) to prevent unauthorized scraping. The XiaoHongShuClient class in media_platform/xhs/client.py handles this automatically.
XiaoHongShuClient._pre_headers invokes the sign_with_xhshow function (located in media_platform/xhs/xhs_sign.py) to generate signed headers before each HTTP request. The client uses httpx for async HTTP communication and maintains a cookie dictionary synchronized with the browser context.
The client also provides high-level methods for data retrieval:
get_note_by_keyword: Searches notes by keyword with paginationget_note_info: Fetches detailed metadata for a specific note IDget_creator_info: Retrieves creator profile dataget_all_notes_by_creator: Paginates through all notes from a specific creator
Proxy Rotation and Resilience
To avoid IP-based rate limiting, the client integrates ProxyRefreshMixin which checks proxy expiration before each request via await self._refresh_proxy_if_expired().
If a proxy fails, the client automatically fetches a new IP from the pool initialized by self.init_proxy_pool(proxy_ip_pool). This mechanism ensures continuous crawling even when individual IPs encounter blocks or bans.
Crawler Core and Execution Modes
The XiaoHongShuCrawler class in media_platform/xhs/core.py orchestrates the entire crawling process. It launches a Playwright browser (standard or CDP mode), manages login fallback, and executes one of three crawl modes:
Search Mode: The search() method iterates over config.KEYWORDS, builds a search_id, and calls client.get_note_by_keyword. Returned note IDs are processed concurrently using asyncio.Semaphore to fetch full details, comments, and media while respecting config.CRAWLER_MAX_SLEEP_SEC throttling.
Detail Mode: get_specified_notes() parses URLs from config.XHS_SPECIFIED_NOTE_URL_LIST (using NoteUrlInfo), then fetches comprehensive note details and associated comments.
Creator Mode: get_creators_and_notes() processes creator URLs (CreatorUrlInfo), retrieves creator profiles via client.get_creator_info, then walks through all creator content using client.get_all_notes_by_creator.
Practical Code Examples
Running a Search Crawl
import asyncio
from media_platform.xhs.core import XiaoHongShuCrawler
async def main():
# Configuration read from .env (LOGIN_TYPE, KEYWORDS, etc.)
crawler = XiaoHongShuCrawler()
await crawler.start() # Launches browser, authenticates, executes search
await crawler.close() # Cleanup Playwright resources
if __name__ == "__main__":
asyncio.run(main())
Manual API Client Usage
import asyncio
from media_platform.xhs.client import XiaoHongShuClient
from tools import utils
async def demo():
headers = {
"Cookie": "web_session=abcdef12345;",
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)..."
}
client = XiaoHongShuClient(
proxy=None,
headers=headers,
playwright_page=None,
cookie_dict=utils.convert_str_cookie_to_dict(headers["Cookie"]),
)
# Search for travel-related content
resp = await client.get_note_by_keyword(keyword="travel", page=1)
print(resp["items"][0]["note_card"]["note_id"])
asyncio.run(demo())
Extending with Custom Login
# In media_platform/xhs/login.py
class XiaoHongShuLogin(AbstractLogin):
async def login_by_api_token(self):
"""Experimental token-based authentication."""
token = config.XHS_API_TOKEN
self.headers["Authorization"] = f"Bearer {token}"
client = XiaoHongShuClient(
headers=self.headers,
playwright_page=self.context_page,
cookie_dict={}
)
if await client.pong():
utils.logger.info("[XiaoHongShuLogin] API token login successful")
else:
raise RuntimeError("Invalid API token")
Summary
- Modular Design: All XHS logic is isolated in
media_platform/xhs/, preventing cross-platform contamination. - Robust Authentication: Three login strategies (QR-code, SMS, cookies) with automatic session validation via
client.pong(). - Cryptographic Security: Automatic header signing via
sign_with_xhshowinXiaoHongShuClient._pre_headers. - Flexible Crawling: Three execution modes (search, detail, creator) support diverse data collection requirements.
- Production Resilience: Built-in proxy rotation via
ProxyRefreshMixinand concurrency control through async semaphores.
Frequently Asked Questions
How does MediaCrawler handle Xiaohongshu authentication?
MediaCrawler implements three authentication strategies in media_platform/xhs/login.py: QR-code scanning, mobile SMS verification, and cookie-based session resumption. The XiaoHongShuLogin.begin method selects the appropriate strategy based on the LOGIN_TYPE configuration value, automatically retrying failed attempts and updating the API client cookies upon success.
What are the X-S and X-T headers required for XHS requests?
Xiaohongshu requires cryptographic signatures in HTTP headers to prevent unauthorized API access. The X-S and X-T headers are generated by the sign_with_xhshow function within XiaoHongShuClient._pre_headers (located in media_platform/xhs/client.py). These headers are computed from request parameters and timestamp data to validate the legitimacy of each API call.
How does the crawler manage proxy rotation for XHS?
The XiaoHongShuClient class inherits from ProxyRefreshMixin, which checks proxy expiration before each request via _refresh_proxy_if_expired(). If the current proxy is banned or expired, the client automatically retrieves a new IP from the pool initialized through init_proxy_pool(proxy_ip_pool), ensuring continuous operation without manual intervention.
What crawl modes are available for Xiaohongshu data collection?
The implementation supports three distinct modes controlled by XiaoHongShuCrawler: Search mode (search()) discovers content by keywords; Detail mode (get_specified_notes()) fetches specific URLs provided in configuration; and Creator mode (get_creators_and_notes()) extracts all content from specified creator profiles. Each mode uses the same underlying API client but applies different orchestration logic for data retrieval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →