# Implementing Multi-Account Crawling with IP Rotation in MediaCrawler

> Learn how to implement multi-account crawling with IP rotation in MediaCrawler. Utilize a shared proxy pool and automatic IP rotation to avoid rate limits and boost your crawling efficiency.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**MediaCrawler enables scalable multi-account crawling by combining a shared proxy pool with automatic IP rotation, ensuring each account request uses a fresh, validated proxy to avoid rate limits.**

This architecture allows you to crawl multiple user accounts on platforms like Zhihu or Douyin concurrently while distributing requests across rotating IP addresses. By leveraging a centralized `ProxyIpPool` and a mixin-based refresh mechanism, the system maintains high availability without exhausting single-provider quotas or triggering platform bans.

## Core Architecture Components

The multi-account crawling system relies on three integrated components that handle proxy lifecycle management and request coordination.

### ProxyIpPool: Centralized IP Management

The `ProxyIpPool` class in [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py) acts as the single source of truth for all proxy operations. It fetches proxy lists from configured providers, validates connectivity, tracks expiration timestamps, and maintains a rotating buffer of healthy IPs.

When initialized via `create_ip_pool()`, the pool loads proxies based on `config.IP_PROXY_PROVIDER_NAME` (supporting providers like Kuai Daili or Wandou) and optionally validates each endpoint. The pool exposes `get_or_refresh_proxy()` to serve fresh IPs and `is_current_proxy_expired()` to check rotation timing.

### ProxyRefreshMixin: Automatic Rotation Logic

The `ProxyRefreshMixin` in [`proxy/proxy_mixin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_mixin.py) injects automatic proxy refresh capabilities into platform-specific clients. Any client inheriting this mixin gains the `_refresh_proxy_if_expired()` method, which intercepts requests to ensure IP freshness.

During client initialization, calling `init_proxy_pool(pool)` binds the shared pool instance. Before each HTTP request, the mixin checks `ProxyIpPool.is_current_proxy_expired()` and automatically rewrites `self.proxy` via `await self._proxy_ip_pool.get_or_refresh_proxy()` when necessary.

### BaseCrawler: Multi-Account Orchestration

The `BaseCrawler` in [`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py) manages the concurrent execution layer. It instantiates separate client objects for each account credential, injects the shared `ProxyIpPool` into each instance, and runs them using `asyncio.gather()`.

Because all clients reference the same pool instance, IP rotation operates globally across accounts rather than per-client, maximizing provider quota efficiency.

## Step-by-Step Implementation Guide

Follow these steps to configure and deploy multi-account crawling with automatic IP rotation.

### Configure the Proxy Provider

Set your proxy source in the configuration files (e.g., [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) or platform-specific configs like [`zhihu_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/zhihu_config.py)):

```python

# config/base_config.py

IP_PROXY_PROVIDER_NAME = "kuai_daili"  # Options: kuai_daili, wandou, or static

STATIC_PROXY_URL = ""                  # Used when provider is "static"

```

Available providers implement fetching logic in `proxy/providers/`, including `new_kuai_daili_proxy()` and `new_wandou_http_proxy()`.

### Initialize the Shared Proxy Pool

Create the pool at application startup to preload and validate proxies:

```python
import asyncio
from proxy.proxy_ip_pool import create_ip_pool

async def init_proxy():
    pool = await create_ip_pool(
        ip_pool_count=10,          # Maintain 10 proxies in rotation

        enable_validate_ip=True    # Verify connectivity before adding to pool

    )
    return pool

# Initialize once for all accounts

proxy_pool = asyncio.run(init_proxy())

```

The `create_ip_pool` function (lines 98-115 in [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py)) handles provider selection and optional validation based on your configuration.

### Create Mixin-Enabled Platform Clients

Modify existing platform clients to inherit from `ProxyRefreshMixin`:

```python
from proxy.proxy_mixin import ProxyRefreshMixin
from media_platform.zhihu.client import ZhihuClient

class RotatingZhihuClient(ProxyRefreshMixin, ZhihuClient):
    def __init__(self, account_cfg, proxy_pool):
        super().__init__(**account_cfg)      # Initialize base ZhihuClient

        self.init_proxy_pool(proxy_pool)     # Register pool with mixin (lines 34-46)

        self.proxy = ""                      # Populated automatically by mixin

```

The `init_proxy_pool()` call (implemented in [`proxy/proxy_mixin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_mixin.py)) stores the pool reference, enabling the automatic refresh logic to fetch new IPs when the current proxy expires.

### Orchestrate Multi-Account Execution

Use `BaseCrawler` patterns to run multiple accounts concurrently with shared IP rotation:

```python
import asyncio
from base.base_crawler import BaseCrawler

async def run_multi_account():
    proxy_pool = await init_proxy()
    
    accounts = [
        {"username": "user_a", "password": "pwd_a"},
        {"username": "user_b", "password": "pwd_b"},
        {"username": "user_c", "password": "pwd_c"}
    ]
    
    crawlers = []
    for cfg in accounts:
        client = RotatingZhihuClient(cfg, proxy_pool)
        crawlers.append(client.crawl())   # Each crawl() implements platform logic

    
    # Execute all clients concurrently

    await asyncio.gather(*crawlers)

asyncio.run(run_multi_account())

```

Each `client.crawl()` invocation internally triggers `await self._refresh_proxy_if_expired()` (lines 57-75 in [`proxy/proxy_mixin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_mixin.py)) before network requests, guaranteeing fresh IPs without manual intervention.

## Key Files and Implementation Details

Understanding these source files helps customize rotation behavior or add new platforms:

- **[`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py)** – Core pool implementation handling proxy fetching, validation, and expiration tracking. Contains `create_ip_pool()` factory and `ProxyIpPool` class methods.
  
- **[`proxy/proxy_mixin.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_mixin.py)** – Inheritance mixin providing `_refresh_proxy_if_expired()` and `init_proxy_pool()` for automatic proxy management.
  
- **`proxy/providers/`** – Provider-specific factories that fetch raw proxy lists from external services like Kuai Daili or Wandou.
  
- **[`base/base_crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base/base_crawler.py)** – Orchestration layer that creates per-account client instances and manages concurrent execution.
  
- **`media_platform/*/client.py`** – Platform-specific implementations (e.g., [`media_platform/zhihu/client.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/media_platform/zhihu/client.py)) that perform actual HTTP requests and inherit the rotation mixin.

## Summary

- **Centralized pool architecture** – `ProxyIpPool` in [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py) manages fetching, validation, and expiration of proxy IPs across all accounts.
- **Mixin-based injection** – `ProxyRefreshMixin` automatically refreshes IPs before requests without modifying platform-specific logic.
- **Concurrent execution** – `BaseCrawler` patterns enable running multiple account clients simultaneously while sharing the same proxy pool.
- **Provider flexibility** – Support for static proxies or dynamic providers (Kuai Daili, Wandou) via configuration in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).
- **Automatic rotation** – Each HTTP request triggers expiration checks via `_refresh_proxy_if_expired()`, ensuring continuous IP freshness.

## Frequently Asked Questions

### How does MediaCrawler prevent IP bans during multi-account crawling?

MediaCrawler prevents IP bans by routing each request through a rotating proxy pool. The `ProxyRefreshMixin` checks `is_current_proxy_expired()` before every request and automatically calls `get_or_refresh_proxy()` to swap IPs, ensuring no single address generates excessive requests across multiple accounts.

### Can I use static proxies instead of dynamic providers?

Yes. Set `IP_PROXY_PROVIDER_NAME` to `"static"` and provide your proxy URL in `STATIC_PROXY_URL` within the configuration files. The `ProxyIpPool` treats static proxies as a single-item pool, though dynamic providers are recommended for multi-account crawling to maximize IP diversity.

### What happens when the proxy pool runs out of valid IPs?

If `enable_validate_ip` is enabled during `create_ip_pool()` initialization, only validated proxies enter the pool. When the pool exhausts valid IPs, `get_or_refresh_proxy()` attempts to fetch new proxies from the configured provider. If all providers fail, the request uses the last known proxy or raises a connection error depending on the client implementation.

### Is each account guaranteed a unique IP for every request?

Not necessarily unique per request, but each request uses the current proxy from the shared pool. If the current proxy expires between requests (based on provider TTL or rotation policy), `ProxyRefreshMixin` assigns a fresh IP. For strict per-request rotation, configure shorter expiration intervals in your proxy provider settings or call `_refresh_proxy_if_expired()` with custom logic.