How the XHS Crawler Handles Dynamic Content and JavaScript-Rendered Elements

The XHS crawler handles dynamic content by executing JavaScript in a headless Playwright browser while maintaining an HTML fallback parser that extracts data from server-rendered initial state when API calls fail.

The MediaCrawler repository provides a robust Xiaohongshu (XHS) scraper that can retrieve content from JavaScript-heavy pages. Unlike simple HTTP clients that only fetch static HTML, this implementation uses browser automation to render dynamic content exactly as a real user would see it.

Playwright-Driven Browser Context

The primary mechanism for handling JavaScript-rendered elements relies on Playwright to launch and control a real Chromium browser instance.

Stealth-Enabled Browser Launch

In media_platform/xhs/core.py, the XiaoHongShuCrawler class initializes a browser context that executes all page scripts. The crawler creates a BrowserContext using async_playwright() and injects libs/stealth.min.js to mask automation fingerprints.


# From media_platform/xhs/core.py (lines 72-94)

async def launch_browser(self):
    self.browser_context = await self.launcher.chromium.launch_persistent_context(
        user_data_dir=self.user_data_dir,
        headless=config.HEADLESS,
        args=browser_args
    )
    # Inject stealth script to avoid detection

    await self.browser_context.add_init_script(
        path="libs/stealth.min.js"
    )

This approach ensures that Single Page Application (SPA) content fully renders before extraction, including dynamically loaded images and text that only appear after JavaScript execution.

Optional CDP Mode

When config.ENABLE_CDP_MODE is set to True, the crawler connects to an existing Chrome instance via Chrome DevTools Protocol instead of launching a new browser. The CDPBrowserManager class in tools/cdp_browser.py handles this connection, reducing resource overhead while maintaining full JavaScript execution capabilities.

HTML Fallback Extraction

When the XHS public API returns errors or rate limits, the crawler falls back to fetching the raw HTML page and parsing the embedded application state.

Parsing the Initial State

The get_note_by_id_from_html method in media_platform/xhs/client.py performs a signed GET request to the note URL, then delegates parsing to XiaoHongShuExtractor.extract_note_detail_from_html in media_platform/xhs/extractor.py. This extractor pulls the window.__INITIAL_STATE__ JSON object directly from the rendered HTML source.


# From media_platform/xhs/extractor.py (lines 31-50)

def extract_note_detail_from_html(self, html_content: str) -> dict:
    # Extract JSON state embedded by the SPA

    initial_state_match = re.search(
        r'window\.__INITIAL_STATE__\s*=\s*({.+?});', 
        html_content
    )
    if initial_state_match:
        return json.loads(initial_state_match.group(1))
    return None

This technique captures the same data available to the JavaScript runtime without requiring API endpoints.

Request Signing and Anti-Detection

Both API and HTML requests require cryptographic signatures to appear legitimate. The playwright_sign.py module implements sign_with_xhshow, which generates X-S, X-T, and other headers using the pure-Python xhshow library.

In media_platform/xhs/client.py (lines 98-111), these headers are injected into every request:

async def sign_request(self, url: str, data: dict = None):
    sign_headers = await sign_with_xhshow(url, data)
    self.headers.update(sign_headers)
    return self.headers

Implementation Workflow

The crawler follows this deterministic flow when handling dynamic content:

  1. Browser Initialization – Launches Chromium via Playwright or connects via CDP mode
  2. Context Creation – Opens a new page and navigates to XHS homepage to obtain fresh cookies
  3. API Attempt – First tries the official API via get_note_by_id
  4. HTML Fallback – If the API fails, calls get_note_by_id_from_html to fetch the rendered page
  5. Data Extraction – Parses either the API response or the HTML initial state using XiaoHongShuExtractor
  6. Storage – Stores results via xhs_store.update_xhs_note

Code Examples

Using the Crawler for Dynamic Content

This example demonstrates automatic handling of JavaScript-rendered notes:

from media_platform.xhs.core import XiaoHongShuCrawler
import asyncio

async def fetch_dynamic_note():
    # Initialize crawler (configuration from config/xhs_config.py)

    crawler = XiaoHongShuCrawler()
    
    # Specify target URLs containing xsec_token parameters

    from config import XHS_SPECIFIED_NOTE_URL_LIST
    XHS_SPECIFIED_NOTE_URL_LIST = [
        "https://www.xiaohongshu.com/explore/abc123?xsec_token=XXXX"
    ]
    
    # Launch browser, render page, and extract data

    await crawler.start()

asyncio.run(fetch_dynamic_note())

Under the hood, crawler.start() triggers launch_browser, creates the Playwright context, and routes through get_note_detail_async_task, which handles the API-to-HTML fallback chain automatically.

Direct HTML Fallback Invocation

For debugging or specific use cases, you can invoke the HTML parser directly:

from media_platform.xhs.client import XiaoHongShuClient
from media_platform.xhs.extractor import XiaoHongShuExtractor
import asyncio

async def html_fallback_demo():
    client = XiaoHongShuClient(
        proxy=None,
        headers={"User-Agent": "Mozilla/5.0"},
        playwright_page=None,
        cookie_dict={}
    )
    
    # Fetch raw HTML when API is unavailable

    html_content = await client.get_note_by_id_from_html(
        note_id="abc123",
        xsec_source="source_xyz", 
        xsec_token="token_123",
        enable_cookie=False  # Strips cookies to avoid bot detection

    )
    
    # Extract structured data from rendered HTML

    extractor = XiaoHongShuExtractor()
    note_data = extractor.extract_note_detail_from_html(html_content)
    print(note_data)

asyncio.run(html_fallback_demo())

Summary

  • Playwright integration in media_platform/xhs/core.py executes JavaScript by launching a real Chromium browser with stealth injection
  • HTML fallback via get_note_by_id_from_html extracts window.__INITIAL_STATE__ from rendered pages when APIs fail
  • CDP mode in tools/cdp_browser.py allows connection to existing Chrome instances for resource efficiency
  • XHShow signing generates required headers to mimic legitimate browser requests and avoid blocking
  • Hybrid extraction ensures data retrieval regardless of whether content comes from API endpoints or client-side JavaScript rendering

Frequently Asked Questions

How does the XHS crawler avoid detection when using Playwright?

The crawler injects libs/stealth.min.js into every browser context using add_init_script() to mask the Playwright automation fingerprints. Additionally, it generates signed headers (X-S, X-T) via the xhshow library to make requests appear identical to those from genuine XHS mobile or web applications.

Can the crawler handle pages that require scrolling to load content?

Yes. Because the implementation uses a full Playwright browser context stored in self.browser_context, you can extend the base crawler to execute scroll actions or wait for specific selectors before extraction. The existing architecture provides the foundation for infinite scroll handling through the context_page object.

What happens when the XHS API rate-limits the crawler?

When get_note_by_id returns an error or empty result, the get_note_detail_async_task method automatically falls back to get_note_by_id_from_html. This method fetches the public note URL as a rendered HTML document and extracts the JSON state embedded in the page source, ensuring data collection continues even during API restrictions.

Is CDP mode required for JavaScript rendering?

No. CDP mode is optional and controlled by config.ENABLE_CDP_MODE. The standard Playwright launcher (launch_browser) fully supports JavaScript execution independently. CDP mode simply offers a resource-efficient alternative when connecting to an already-running Chrome instance is preferable to launching a new browser process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →