How to Implement Multi-Account Rotation with IP Proxy Pools in MediaCrawler
MediaCrawler provides built-in proxy pooling infrastructure and platform-specific exception handling that you can combine with a lightweight account manager to rotate both IP addresses and user credentials automatically.
The NanmiCoder/MediaCrawler repository separates two critical concerns for large-scale social media scraping: credential management and network anonymity. This separation lets you build a resilient crawler that switches accounts when one gets blocked, while simultaneously routing traffic through rotating IP addresses to avoid detection.
Understanding the Core Components
MediaCrawler's architecture gives you two extensible systems to leverage:
| Component | Purpose | Key Location |
|---|---|---|
| Proxy Pool | Load, validate, and rotate IP proxies on-demand | proxy/proxy_ip_pool.py |
| Account Credentials | Per-platform authentication (phone, password, cookies) | config/xhs_config.py, config/douyin_config.py, etc. |
| Access Exceptions | Detect blocks, bans, and rate-limits | media_platform/xhs/exception.py |
The ProxyIpPool class handles the networking layer. Your custom AccountManager handles the identity layer. When combined inside a retry loop, they create a self-healing crawler.
Step 1: Initialize the Proxy Pool
The proxy/proxy_ip_pool.py module exports create_ip_pool—a factory that returns a configured ProxyIpPool instance.
from proxy.proxy_ip_pool import create_ip_pool
proxy_pool = await create_ip_pool(
ip_pool_count=10, # Keep 10 proxies in rotation
enable_validate_ip=True # Verify each proxy before use
)
Under the hood, get_or_refresh_proxy() checks is_current_proxy_expired and draws from self.proxy_list when needed. When validation is enabled, each candidate is tested against self.valid_ip_url before being returned.
Step 2: Build a Round-Robin Account Manager
MediaCrawler does not include a multi-account manager, but you can add one in approximately 20 lines. Create utils/account_manager.py:
from typing import List, Dict
class AccountManager:
"""Round-robin credential rotation for any platform."""
def __init__(self, accounts: List[Dict]):
if not accounts:
raise ValueError("At least one account must be supplied")
self._accounts = accounts
self._cur = 0
def current(self) -> Dict:
"""Return credentials for the active account."""
return self._accounts[self._cur]
def next(self) -> Dict:
"""Advance index and return new credentials."""
self._cur = (self._cur + 1) % len(self._accounts)
return self.current()
Load accounts from environment variables, JSON files, or a secrets manager—extending the pattern already used in config/xhs_config.py.
Step 3: Wire Proxy and Account Rotation Together
The integration happens in your request wrapper. Here's a complete pattern using the XHS client:
import asyncio
from proxy.proxy_ip_pool import create_ip_pool
from tools.httpx_util import make_async_client
from utils.account_manager import AccountManager
from media_platform.xhs.client import XHSClient, PlatformAccessError
async def resilient_crawl():
# Initialize infrastructure
proxy_pool = await create_ip_pool(10, enable_validate_ip=True)
accounts = [
{"phone": "13800138000", "password": "account_1_pass"},
{"phone": "13900139000", "password": "account_2_pass"},
{"phone": "13700137000", "password": "account_3_pass"},
]
acc_mgr = AccountManager(accounts)
# Create client with first account
client = XHSClient(**acc_mgr.current())
for target_url in crawl_queue:
success = False
attempts = 0
max_attempts = len(accounts) * 2 # Allow full rotation twice
while not success and attempts < max_attempts:
proxy = await proxy_pool.get_or_refresh_proxy()
proxy_url = f"http://{proxy.ip}:{proxy.port}"
async with make_async_client(proxy=proxy_url) as http:
try:
data = await client.fetch_post_detail(http, target_url)
process(data)
success = True
except PlatformAccessError as e:
# Block detected: rotate both identity and network
attempts += 1
print(f"Block detected on account {acc_mgr.current()['phone']}")
# Switch account
client.update_credentials(**acc_mgr.next())
# Force fresh proxy for clean slate
await proxy_pool.get_or_refresh_proxy()
if not success:
print(f"Failed to fetch {target_url} after exhausting all accounts")
Key Integration Points
make_async_client(tools/httpx_util.py): Creates anhttpx.AsyncClientpreconfigured with your proxy URLPlatformAccessError(media_platform/xhs/exception.py): Caught to trigger rotation; equivalent exceptions exist for Douyin, Tieba, and Weibo clientsclient.update_credentials(): Platform clients store credentials inself.account; replacing this object changes all subsequent request headers
Step 4: Persist Rotation State (Optional)
For long-running crawlers, persist the account index to avoid restarting with account 0 after every crash. The repository includes test_expiring_local_cache.py demonstrating file-based caching:
import json
import os
class PersistentAccountManager(AccountManager):
def __init__(self, accounts: List[Dict], state_file: str = ".account_idx"):
super().__init__(accounts)
self._state_file = state_file
self._load_state()
def _load_state(self):
if os.path.exists(self._state_file):
with open(self._state_file) as f:
saved = json.load(f)
self._cur = saved.get("index", 0) % len(self._accounts)
def next(self) -> Dict:
creds = super().next()
with open(self._state_file, "w") as f:
json.dump({"index": self._cur}, f)
return creds
Proxy Provider Integration
proxy/providers/ contains ready-made factories for commercial proxy services:
| Provider | Factory Function | File |
|---|---|---|
| 快代理 (Kuai Daili) | new_kuai_daili_proxy |
proxy/providers/kuai_daili.py |
| 豌豆HTTP | new_wandou_http_proxy |
proxy/providers/wandou_http.py |
These return proxy dictionaries that ProxyIpPool consumes. Add your own provider by matching this output format:
{"ip": "1.2.3.4", "port": 8080, "protocol": "http"}
Platform-Specific Considerations
Different platforms trigger blocks differently. Adjust your exception handling per client:
- XiaoHongShu (XHS): Raises
PlatformAccessErrorfor rate-limits and IP blocks - Douyin: Raises authentication errors when session cookies expire
- Tieba/Weibo: Similar patterns in their respective
client.pyfiles
Inspect media_platform/<platform>/client.py to identify the exact exception class for each target.
Summary
proxy/proxy_ip_pool.pyprovidescreate_ip_poolandProxyIpPool.get_or_refresh_proxy()for IP rotation- Build a lightweight
AccountManagerto cycle through credential sets using round-robin logic - Wrap requests with
make_async_client(proxy=...)and catchPlatformAccessErrorto trigger rotation - Combine both rotations—switching accounts and refreshing proxies—whenever a security exception occurs
- Optionally persist account index using the caching pattern from
test_expiring_local_cache.py
Frequently Asked Questions
How many accounts and proxies do I need for reliable crawling?
Start with 3-5 accounts and 10 proxies per account. MediaCrawler's ProxyIpPool validates IPs before use, so your effective pool size depends on provider quality. Monitor block rates and scale horizontally—more accounts reduce per-account request volume, while more proxies distribute IP-based risk.
Can I use the same proxy pool across multiple platforms?
Yes. ProxyIpPool is platform-agnostic. Initialize one pool and share it across XHS, Douyin, and Weibo clients. Each make_async_client call requests a fresh proxy independently, so concurrent platform scraping naturally distributes traffic across your IP pool.
What happens if all accounts get blocked?
The example code includes a max_attempts limit. When exhausted, the specific URL fails and logs for manual review. Implement exponential backoff or CAPTCHA-solving hooks in the failure path. For critical pipelines, integrate with notification systems (email, Slack, PagerDuty) when account exhaustion occurs.
Does MediaCrawler support cookie-based account rotation instead of password login?
Yes. Platform clients accept cookie dictionaries in their account parameters. Store multiple cookie sets in your accounts list and rotate them identically. This avoids expensive re-login requests and reduces detection fingerprints compared to repeated password authentication flows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →