Integrating Third-Party Proxy Provider APIs with MediaCrawler: A Complete Implementation Guide

MediaCrawler abstracts proxy handling behind a modular, async-first architecture that lets you plug in any HTTP-based proxy service by implementing a single abstract method, without touching core crawling logic.

MediaCrawler is an open-source social media crawling framework that delegates proxy management to a pluggable provider system. When you need to integrate a new third-party proxy provider API, you create a concrete implementation of the ProxyProvider abstract base class and leverage the built-in IpCache for Redis-backed storage. This design keeps vendor-specific HTTP logic isolated while the crawler automatically handles rotation, expiration, and fallback to remote APIs when local caches deplete.

Understanding the Proxy Architecture

MediaCrawler organizes proxy functionality into three distinct layers that separate concerns between abstract contracts, caching infrastructure, and vendor-specific implementations.

The ProxyProvider Abstract Base Class

The foundation resides in proxy/base_proxy.py, where the ProxyProvider ABC defines the contract every provider must fulfill. The critical method signature appears at lines 42-50:

class ProxyProvider(ABC):
    @abstractmethod
    async def get_proxy(self, num: int) -> List[IpInfoModel]:
        """Fetch proxies from the provider."""
        pass

Any concrete provider must implement this async method to return a list of IpInfoModel objects, which are Pydantic models defined in proxy/types.py containing fields for ip, port, user, password, and expired_time_ts.

The IpCache Layer

Directly beneath the provider interface sits IpCache, implemented in proxy/base_proxy.py lines 54-84. This helper stores fetched proxies in Redis with TTL management derived from the provider's expire_time. The cache logic ensures that:

  • Expired entries are purged automatically via Redis TTL
  • Valid proxies serve from local memory before hitting external APIs
  • Rate limits imposed by proxy vendors are respected

Reference Implementation: WanDou Provider

The proxy/providers/wandou_http_proxy.py file demonstrates a production-ready integration. It shows how to translate vendor-specific JSON responses into standardized IpInfoModel instances while handling environment-based credential injection. The factory helper at lines 108-124 (new_wandou_http_proxy) reads the WANDOU_APP_KEY environment variable and returns a configured provider instance ready for immediate use.

Step-by-Step Integration Guide

Adding a new third-party proxy provider requires four discrete steps that keep your changes isolated to a single new file under proxy/providers/.

Step 1: Create the Provider Module

Create a new Python file in proxy/providers/ (e.g., fastproxy_http.py) and inherit from ProxyProvider:

import os
from typing import List
from proxy.base_proxy import ProxyProvider, IpCache
from proxy.types import IpInfoModel

class FastProxyHttp(ProxyProvider):
    def __init__(self, api_key: str, num: int = 50):
        self.proxy_brand_name = "FASTPROXY"
        self.api_uri = "https://api.fastproxy.com/get"
        self.params = {"key": api_key, "count": num}
        self.ip_cache = IpCache()

Step 2: Implement the HTTP Logic

Implement the required get_proxy method to query the vendor endpoint using httpx from tools/httpx_util.py, convert responses to IpInfoModel, and populate the cache:

    async def get_proxy(self, num: int) -> List[IpInfoModel]:
        # ① Try cache first (pattern from wandou_http_proxy.py lines 58-64)

        cached = self.ip_cache.load_all_ip(self.proxy_brand_name)
        if len(cached) >= num:
            return cached[:num]

        # ② Not enough → call remote API

        need = num - len(cached)
        self.params["count"] = need
        
        from tools.httpx_util import make_async_client
        from urllib.parse import urlencode
        import httpx
        
        async with make_async_client() as client:
            url = f"{self.api_uri}?{urlencode(self.params)}"
            resp = await client.get(url)
            data = resp.json().get("list", [])
            
            new_ips = []
            for item in data:
                model = IpInfoModel(
                    ip=item["ip"],
                    port=item["port"],
                    user=item.get("user", ""),
                    password=item.get("pwd", ""),
                    expired_time_ts=self._parse_expiry(item["expire"]),
                )
                # Store with TTL (base_proxy.py lines 58-66)

                key = f"{self.proxy_brand_name}_{model.ip}_{model.port}"
                ttl = model.expired_time_ts - utils.get_unix_timestamp()
                self.ip_cache.set_ip(key, model.model_dump_json(), ex=ttl)
                new_ips.append(model)
            
            return cached + new_ips

Step 3: Add a Factory Function

Expose a factory helper that reads credentials from environment variables, following the dual-case handling pattern seen in wandou_http_proxy.py lines 20-24:

def new_fastproxy_http() -> FastProxyHttp:
    """Factory that reads FASTPROXY_API_KEY from environment."""
    api_key = os.getenv("FASTPROXY_API_KEY") or "demo_key"
    return FastProxyHttp(api_key=api_key)

Step 4: Consume the Provider in Crawlers

Instantiate your provider in any spider or fetcher and call get_proxy() to obtain rotating proxies. The returned IpInfoModel objects integrate directly with httpx.AsyncClient or the CDP browser launcher in tools/cdp_browser.py:

from proxy.providers.fastproxy_http import new_fastproxy_http
import httpx

class MySpider:
    def __init__(self):
        self.proxy_provider = new_fastproxy_http()
    
    async def fetch_page(self, url: str):
        proxies = await self.proxy_provider.get_proxy(1)
        proxy = proxies[0]
        proxy_url = f"http://{proxy.user}:{proxy.password}@{proxy.ip}:{proxy.port}" \
                    if proxy.user else f"http://{proxy.ip}:{proxy.port}"
        
        async with httpx.AsyncClient(proxies=proxy_url) as client:
            return await client.get(url)

Complete Implementation Example

Here is a full working implementation for a hypothetical "FastProxy" service that demonstrates all integration points:


# File: proxy/providers/fastproxy_http.py

import os
from typing import List
from urllib.parse import urlencode

import httpx
from proxy.base_proxy import IpCache, ProxyProvider, IpGetError
from proxy.types import IpInfoModel
from tools import utils
from tools.httpx_util import make_async_client


class FastProxyHttp(ProxyProvider):
    def __init__(self, api_key: str, num: int = 50):
        self.proxy_brand_name = "FASTPROXY"
        self.api_uri = "https://api.fastproxy.com/get"
        self.params = {"key": api_key, "count": num}
        self.ip_cache = IpCache()

    async def get_proxy(self, num: int) -> List[IpInfoModel]:
        # Check cache first

        cached = self.ip_cache.load_all_ip(self.proxy_brand_name)
        if len(cached) >= num:
            return cached[:num]

        # Fetch remaining from API

        need = num - len(cached)
        self.params["count"] = need
        
        async with make_async_client() as client:
            url = f"{self.api_uri}?{urlencode(self.params)}"
            utils.logger.info(f"[FastProxyHttp] Requesting {url}")
            resp = await client.get(url)
            data = resp.json().get("list", [])
            
            now = utils.get_unix_timestamp()
            new_ips = []
            for item in data:
                model = IpInfoModel(
                    ip=item["ip"],
                    port=item["port"],
                    user=item.get("user", ""),
                    password=item.get("pwd", ""),
                    expired_time_ts=utils.get_unix_time_from_time_str(item["expire"]),
                )
                key = f"FASTPROXY_{model.ip}_{model.port}"
                ttl = model.expired_time_ts - now
                self.ip_cache.set_ip(key, model.model_dump_json(), ex=ttl)
                new_ips.append(model)
            
            return cached + new_ips


def new_fastproxy_http() -> FastProxyHttp:
    """Factory reading FASTPROXY_API_KEY from environment."""
    api_key = os.getenv("FASTPROXY_API_KEY") or "demo_key"
    return FastProxyHttp(api_key=api_key)

Usage remains consistent across all MediaCrawler spiders:

import asyncio
from proxy.providers.fastproxy_http import new_fastproxy_http

async def demo():
    provider = new_fastproxy_http()
    proxies = await provider.get_proxy(5)
    for p in proxies:
        print(f"Proxy: http://{p.ip}:{p.port} (expires {p.expired_time_ts})")

asyncio.run(demo())

Why This Design Works

Separation of concerns isolates vendor-specific HTTP formats in dedicated provider files while the crawler core remains agnostic to proxy implementation details.

Cache-driven efficiency via IpCache minimizes external API calls and respects provider rate limits through Redis TTL management, as implemented in base_proxy.py lines 58-66.

Async-first architecture ensures all network calls use httpx.AsyncClient (from tools/httpx_util.py), keeping the crawler non-blocking during proxy acquisition.

Pluggable extensibility means adding a new provider requires only creating a new file under proxy/providers/ and exposing a factory function—no modifications to existing crawling logic are necessary.

Summary

  • Abstract contract: Implement ProxyProvider.get_proxy() in proxy/base_proxy.py to define how your crawler receives proxies.
  • Caching layer: Leverage IpCache to store proxies in Redis with automatic TTL expiration, reducing API costs and preventing expired proxy usage.
  • Vendor isolation: Place concrete implementations in proxy/providers/ and use IpInfoModel from proxy/types.py as the universal data transfer object.
  • Factory pattern: Provide environment-based initialization functions (like new_wandou_http_proxy) to keep credentials out of source code.
  • Zero core changes: The modular design allows integration of any HTTP proxy service without modifying MediaCrawler's spider or browser automation code.

Frequently Asked Questions

What is the minimum code required to add a new proxy provider?

You must create a class inheriting from ProxyProvider and implement the async def get_proxy(self, num: int) -> List[IpInfoModel] method. This method handles HTTP communication with the vendor API, converts responses to IpInfoModel instances, and returns them. Optionally add a factory function to read API keys from environment variables, as shown in wandou_http_proxy.py lines 108-124.

How does MediaCrawler handle proxy expiration and caching?

The IpCache class in proxy/base_proxy.py manages a Redis-backed cache with TTL support. When you call set_ip(key, value, ex=ttl), the cache automatically expires entries based on the proxy provider's reported expire_time. The get_proxy method checks this cache first (returning valid entries immediately) before making expensive HTTP calls to refill the pool.

Can I use the same provider with both httpx and CDP browser automation?

Yes. The IpInfoModel objects returned by get_proxy() contain standard fields (ip, port, user, password) compatible with both httpx.AsyncClient proxies and the Chrome DevTools Protocol (CDP) browser launcher in tools/cdp_browser.py. Simply format the proxy URL as http://user:pass@ip:port or http://ip:port depending on authentication requirements.

Where should I store API credentials for third-party proxy services?

Store credentials in environment variables and read them via factory functions like new_fastproxy_http(). This pattern, demonstrated in wandou_http_proxy.py lines 20-24, keeps sensitive keys out of version control while allowing runtime configuration through .env files or container orchestration secrets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →