# How NanmiCoder/MediaCrawler Handles Proxy Rotation for XHS Scraping

> Learn how NanmiCoder/MediaCrawler handles XHS scraping with automatic proxy rotation. Discover its three-layer architecture: ProxyIpPool, ProxyRefreshMixin, and XiaoHongShuClient.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-03

---

**NanmiCoder/MediaCrawler implements automatic proxy rotation for XHS scraping through a three-layer architecture: a ProxyIpPool managing proxy lifecycles, a ProxyRefreshMixin providing pre-request validation hooks, and the XiaoHongShuClient orchestrating both to ensure every HTTP request routes through a fresh, validated proxy.**

The MediaCrawler repository provides a robust Python framework for scraping Xiaohongshu (XHS) content while mitigating IP-based rate limiting through intelligent proxy management. This open-source solution decouples proxy rotation logic from platform-specific implementations, creating a reusable pipeline that validates, rotates, and refreshes proxy connections automatically. Understanding how NanmiCoder/MediaCrawler handles proxy rotation for scraping XHS reveals a design pattern that scales across all supported platforms including Weibo, Douyin, and Bilibili.

## Core Components of the Proxy Rotation System

The proxy rotation mechanism relies on three interconnected components that handle proxy storage, validation, and automatic refreshing.

### ProxyIpPool: The Proxy Lifecycle Manager

The `ProxyIpPool` class in [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py) serves as the central repository for proxy management. It maintains a collection of `IpInfoModel` descriptors and handles loading, validation, and expiration tracking.

Key methods include:

- **`get_proxy()`** – Selects a random proxy from the pool, optionally validates it against a test URL, and marks it as the current active proxy.
- **`is_current_proxy_expired()`** – Checks the `expired_time_ts` field against the current timestamp with a configurable buffer to determine if rotation is needed.
- **`get_or_refresh_proxy()`** – Returns the current proxy if still valid; otherwise triggers acquisition of a new proxy via `get_proxy()`.

The pool is typically instantiated via `create_ip_pool()`, which accepts parameters for pool size and validation settings.

### ProxyRefreshMixin: Automatic Pre-Request Validation

The `ProxyRefreshMixin` class in [`proxy/proxy_mixin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_mixin.py) provides the automatic refresh logic that platform clients inherit. It stores a reference to a `ProxyIpPool` instance and intercepts HTTP requests to ensure proxy freshness.

The critical method is **`_refresh_proxy_if_expired()`**, which:
1. Checks if the current proxy is expired via `self._proxy_ip_pool.is_current_proxy_expired()`
2. If expired, calls `get_or_refresh_proxy()` to obtain a new `IpInfoModel`
3. Constructs the proper proxy URL format (`http://user:pass@ip:port` or `http://ip:port`)
4. Updates `self.proxy` with the new connection string

This mix-in ensures that proxy rotation happens transparently before any network request executes.

### XiaoHongShuClient: XHS-Specific Implementation

The `XiaoHongShuClient` in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) inherits from both `AbstractApiClient` and `ProxyRefreshMixin`, integrating the proxy rotation pipeline into XHS-specific API calls.

During initialization, the client accepts an optional `proxy_ip_pool` argument and forwards it to `ProxyRefreshMixin.init_proxy_pool()`. Every public request method (`request`, `get`, `post`) invokes `await self._refresh_proxy_if_expired()` at the start of execution, guaranteeing fresh proxies for each XHS API interaction.

## Step-by-Step Proxy Rotation Flow

The complete rotation workflow operates as follows:

1. **Pool Initialization** – At startup, `create_ip_pool()` instantiates a `ProxyIpPool` with a configured number of proxies and validation settings.

2. **Client Configuration** – When `XiaoHongShuClient` initializes, it receives the pool reference and stores it via the mix-in's initialization method.

3. **Pre-Request Validation** – Before executing any HTTP call, `XiaoHongShuClient.request()` triggers `_refresh_proxy_if_expired()`.

4. **Expiration Check** – The mix-in queries `is_current_proxy_expired()` to determine if the current proxy is still valid.

5. **Proxy Refresh** – If expired, `get_or_refresh_proxy()` fetches a new proxy from the pool, which may invoke `get_proxy()` to retrieve and validate a fresh endpoint.

6. **HTTP Execution** – The request proceeds through `make_async_client(proxy=self.proxy)` in [`tools/httpx_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/httpx_util.py), which creates an `httpx.AsyncClient` routing traffic through the selected proxy.

7. **Continuous Rotation** – Subsequent requests repeat this validation cycle, enabling automatic rotation without manual intervention.

## Code Implementation Examples

### Initializing the Proxy Pool and XHS Client

```python
from proxy.proxy_ip_pool import create_ip_pool
from media_platform.xhs.client import XiaoHongShuClient
from tools.httpx_util import make_async_client

# Build a pool maintaining 5 proxies with validation enabled

proxy_pool = await create_ip_pool(ip_pool_count=5, enable_validate_ip=True)

# Initialize the XHS client with proxy support

client = XiaoHongShuClient(
    timeout=30,
    headers={"User-Agent": "MediaCrawler/1.0"},
    playwright_page=page,          # Playwright Page instance

    cookie_dict={"session": "abc123"},
    proxy_ip_pool=proxy_pool,
)

# Automatic proxy rotation occurs on every request

note_data = await client.get("/api/sns/v1/note/detail", params={"note_id": "123456"})

```

### Request Method with Auto-Refresh Hook

```python

# Inside media_platform/xhs/client.py

async def request(self, method, url, **kwargs):
    # Automatic proxy refresh before every request

    await self._refresh_proxy_if_expired()
    
    async with make_async_client(proxy=self.proxy) as client:
        response = await client.request(
            method, 
            url, 
            timeout=self.timeout, 
            **kwargs
        )
    return response

```

### ProxyRefreshMixin Internal Logic

```python

# Inside proxy/proxy_mixin.py

async def _refresh_proxy_if_expired(self):
    if self._proxy_ip_pool.is_current_proxy_expired():
        new_proxy = await self._proxy_ip_pool.get_or_refresh_proxy()
        
        # Construct proxy URL with optional authentication

        if new_proxy.user and new_proxy.password:
            self.proxy = f"http://{new_proxy.user}:{new_proxy.password}@{new_proxy.ip}:{new_proxy.port}"
        else:
            self.proxy = f"http://{new_proxy.ip}:{new_proxy.port}"

```

## Summary

- **NanmiCoder/MediaCrawler** implements proxy rotation through a decoupled three-component architecture reusable across all platform clients.
- The **`ProxyIpPool`** in [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py) handles proxy acquisition, validation, and expiration tracking via `get_proxy()` and `is_current_proxy_expired()`.
- **`ProxyRefreshMixin`** in [`proxy/proxy_mixin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_mixin.py) provides the `_refresh_proxy_if_expired()` hook that automatically updates proxy connections before HTTP requests.
- **`XiaoHongShuClient`** in [`media_platform/xhs/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/xhs/client.py) inherits this mix-in and triggers validation at the start of every request method.
- Proxy URLs are formatted with optional authentication and passed to `httpx.AsyncClient` via `make_async_client()` in [`tools/httpx_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/httpx_util.py).

## Frequently Asked Questions

### How does the proxy pool validate that an IP address is working before using it?

The `ProxyIpPool.get_proxy()` method optionally validates proxies when `enable_validate_ip` is set to `True` during pool creation. It tests the proxy against a target URL to ensure connectivity before marking it as the current active proxy, preventing the use of dead or blocked IPs in the rotation cycle.

### Can I use the same proxy rotation system for platforms other than XHS?

Yes, the architecture is designed for reusability. Any platform client can inherit from `ProxyRefreshMixin` and initialize with a `ProxyIpPool` instance. The repository uses this same pattern for Weibo, Douyin, and Bilibili clients, requiring only the specific API implementation while sharing the proxy management logic.

### What happens if all proxies in the pool expire simultaneously?

If `is_current_proxy_expired()` returns `True` and the pool has no valid proxies remaining, `get_or_refresh_proxy()` will attempt to fetch a new proxy via `get_proxy()`. If the provider cannot supply fresh IPs, the method may raise an exception or return `None`, depending on the specific provider implementation in files like [`proxy/providers/wandou_http_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/providers/wandou_http_proxy.py).

### How do I configure the proxy pool size and validation settings?

Pass parameters to `create_ip_pool()` at application startup. Set `ip_pool_count` to define how many proxies to maintain simultaneously, and set `enable_validate_ip=True` to force validation against test URLs before proxies enter the rotation cycle. These settings apply to all subsequent XHS scraping operations using that pool instance.