# Integrating Third-Party Proxy Provider APIs with MediaCrawler: A Complete Implementation Guide

> Easily integrate third-party proxy provider APIs with MediaCrawler. Our guide shows how to implement a single method for seamless proxy handling in your crawling projects.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**MediaCrawler abstracts proxy handling behind a modular, async-first architecture that lets you plug in any HTTP-based proxy service by implementing a single abstract method, without touching core crawling logic.**

MediaCrawler is an open-source social media crawling framework that delegates proxy management to a pluggable provider system. When you need to integrate a new third-party proxy provider API, you create a concrete implementation of the `ProxyProvider` abstract base class and leverage the built-in `IpCache` for Redis-backed storage. This design keeps vendor-specific HTTP logic isolated while the crawler automatically handles rotation, expiration, and fallback to remote APIs when local caches deplete.

## Understanding the Proxy Architecture

MediaCrawler organizes proxy functionality into three distinct layers that separate concerns between abstract contracts, caching infrastructure, and vendor-specific implementations.

### The ProxyProvider Abstract Base Class

The foundation resides in [`proxy/base_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/base_proxy.py), where the `ProxyProvider` ABC defines the contract every provider must fulfill. The critical method signature appears at lines 42-50:

```python
class ProxyProvider(ABC):
    @abstractmethod
    async def get_proxy(self, num: int) -> List[IpInfoModel]:
        """Fetch proxies from the provider."""
        pass

```

Any concrete provider must implement this async method to return a list of `IpInfoModel` objects, which are Pydantic models defined in [`proxy/types.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/types.py) containing fields for `ip`, `port`, `user`, `password`, and `expired_time_ts`.

### The IpCache Layer

Directly beneath the provider interface sits `IpCache`, implemented in [`proxy/base_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/base_proxy.py) lines 54-84. This helper stores fetched proxies in Redis with TTL management derived from the provider's `expire_time`. The cache logic ensures that:

- Expired entries are purged automatically via Redis TTL
- Valid proxies serve from local memory before hitting external APIs
- Rate limits imposed by proxy vendors are respected

### Reference Implementation: WanDou Provider

The [`proxy/providers/wandou_http_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/providers/wandou_http_proxy.py) file demonstrates a production-ready integration. It shows how to translate vendor-specific JSON responses into standardized `IpInfoModel` instances while handling environment-based credential injection. The factory helper at lines 108-124 (`new_wandou_http_proxy`) reads the `WANDOU_APP_KEY` environment variable and returns a configured provider instance ready for immediate use.

## Step-by-Step Integration Guide

Adding a new third-party proxy provider requires four discrete steps that keep your changes isolated to a single new file under `proxy/providers/`.

### Step 1: Create the Provider Module

Create a new Python file in `proxy/providers/` (e.g., [`fastproxy_http.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/fastproxy_http.py)) and inherit from `ProxyProvider`:

```python
import os
from typing import List
from proxy.base_proxy import ProxyProvider, IpCache
from proxy.types import IpInfoModel

class FastProxyHttp(ProxyProvider):
    def __init__(self, api_key: str, num: int = 50):
        self.proxy_brand_name = "FASTPROXY"
        self.api_uri = "https://api.fastproxy.com/get"
        self.params = {"key": api_key, "count": num}
        self.ip_cache = IpCache()

```

### Step 2: Implement the HTTP Logic

Implement the required `get_proxy` method to query the vendor endpoint using `httpx` from [`tools/httpx_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/httpx_util.py), convert responses to `IpInfoModel`, and populate the cache:

```python
    async def get_proxy(self, num: int) -> List[IpInfoModel]:
        # ① Try cache first (pattern from wandou_http_proxy.py lines 58-64)

        cached = self.ip_cache.load_all_ip(self.proxy_brand_name)
        if len(cached) >= num:
            return cached[:num]

        # ② Not enough → call remote API

        need = num - len(cached)
        self.params["count"] = need
        
        from tools.httpx_util import make_async_client
        from urllib.parse import urlencode
        import httpx
        
        async with make_async_client() as client:
            url = f"{self.api_uri}?{urlencode(self.params)}"
            resp = await client.get(url)
            data = resp.json().get("list", [])
            
            new_ips = []
            for item in data:
                model = IpInfoModel(
                    ip=item["ip"],
                    port=item["port"],
                    user=item.get("user", ""),
                    password=item.get("pwd", ""),
                    expired_time_ts=self._parse_expiry(item["expire"]),
                )
                # Store with TTL (base_proxy.py lines 58-66)

                key = f"{self.proxy_brand_name}_{model.ip}_{model.port}"
                ttl = model.expired_time_ts - utils.get_unix_timestamp()
                self.ip_cache.set_ip(key, model.model_dump_json(), ex=ttl)
                new_ips.append(model)
            
            return cached + new_ips

```

### Step 3: Add a Factory Function

Expose a factory helper that reads credentials from environment variables, following the dual-case handling pattern seen in [`wandou_http_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/wandou_http_proxy.py) lines 20-24:

```python
def new_fastproxy_http() -> FastProxyHttp:
    """Factory that reads FASTPROXY_API_KEY from environment."""
    api_key = os.getenv("FASTPROXY_API_KEY") or "demo_key"
    return FastProxyHttp(api_key=api_key)

```

### Step 4: Consume the Provider in Crawlers

Instantiate your provider in any spider or fetcher and call `get_proxy()` to obtain rotating proxies. The returned `IpInfoModel` objects integrate directly with `httpx.AsyncClient` or the CDP browser launcher in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py):

```python
from proxy.providers.fastproxy_http import new_fastproxy_http
import httpx

class MySpider:
    def __init__(self):
        self.proxy_provider = new_fastproxy_http()
    
    async def fetch_page(self, url: str):
        proxies = await self.proxy_provider.get_proxy(1)
        proxy = proxies[0]
        proxy_url = f"http://{proxy.user}:{proxy.password}@{proxy.ip}:{proxy.port}" \
                    if proxy.user else f"http://{proxy.ip}:{proxy.port}"
        
        async with httpx.AsyncClient(proxies=proxy_url) as client:
            return await client.get(url)

```

## Complete Implementation Example

Here is a full working implementation for a hypothetical "FastProxy" service that demonstrates all integration points:

```python

# File: proxy/providers/fastproxy_http.py

import os
from typing import List
from urllib.parse import urlencode

import httpx
from proxy.base_proxy import IpCache, ProxyProvider, IpGetError
from proxy.types import IpInfoModel
from tools import utils
from tools.httpx_util import make_async_client


class FastProxyHttp(ProxyProvider):
    def __init__(self, api_key: str, num: int = 50):
        self.proxy_brand_name = "FASTPROXY"
        self.api_uri = "https://api.fastproxy.com/get"
        self.params = {"key": api_key, "count": num}
        self.ip_cache = IpCache()

    async def get_proxy(self, num: int) -> List[IpInfoModel]:
        # Check cache first

        cached = self.ip_cache.load_all_ip(self.proxy_brand_name)
        if len(cached) >= num:
            return cached[:num]

        # Fetch remaining from API

        need = num - len(cached)
        self.params["count"] = need
        
        async with make_async_client() as client:
            url = f"{self.api_uri}?{urlencode(self.params)}"
            utils.logger.info(f"[FastProxyHttp] Requesting {url}")
            resp = await client.get(url)
            data = resp.json().get("list", [])
            
            now = utils.get_unix_timestamp()
            new_ips = []
            for item in data:
                model = IpInfoModel(
                    ip=item["ip"],
                    port=item["port"],
                    user=item.get("user", ""),
                    password=item.get("pwd", ""),
                    expired_time_ts=utils.get_unix_time_from_time_str(item["expire"]),
                )
                key = f"FASTPROXY_{model.ip}_{model.port}"
                ttl = model.expired_time_ts - now
                self.ip_cache.set_ip(key, model.model_dump_json(), ex=ttl)
                new_ips.append(model)
            
            return cached + new_ips


def new_fastproxy_http() -> FastProxyHttp:
    """Factory reading FASTPROXY_API_KEY from environment."""
    api_key = os.getenv("FASTPROXY_API_KEY") or "demo_key"
    return FastProxyHttp(api_key=api_key)

```

Usage remains consistent across all MediaCrawler spiders:

```python
import asyncio
from proxy.providers.fastproxy_http import new_fastproxy_http

async def demo():
    provider = new_fastproxy_http()
    proxies = await provider.get_proxy(5)
    for p in proxies:
        print(f"Proxy: http://{p.ip}:{p.port} (expires {p.expired_time_ts})")

asyncio.run(demo())

```

## Why This Design Works

**Separation of concerns** isolates vendor-specific HTTP formats in dedicated provider files while the crawler core remains agnostic to proxy implementation details.

**Cache-driven efficiency** via `IpCache` minimizes external API calls and respects provider rate limits through Redis TTL management, as implemented in [`base_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/base_proxy.py) lines 58-66.

**Async-first architecture** ensures all network calls use `httpx.AsyncClient` (from [`tools/httpx_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/httpx_util.py)), keeping the crawler non-blocking during proxy acquisition.

**Pluggable extensibility** means adding a new provider requires only creating a new file under `proxy/providers/` and exposing a factory function—no modifications to existing crawling logic are necessary.

## Summary

- **Abstract contract**: Implement `ProxyProvider.get_proxy()` in [`proxy/base_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/base_proxy.py) to define how your crawler receives proxies.
- **Caching layer**: Leverage `IpCache` to store proxies in Redis with automatic TTL expiration, reducing API costs and preventing expired proxy usage.
- **Vendor isolation**: Place concrete implementations in `proxy/providers/` and use `IpInfoModel` from [`proxy/types.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/types.py) as the universal data transfer object.
- **Factory pattern**: Provide environment-based initialization functions (like `new_wandou_http_proxy`) to keep credentials out of source code.
- **Zero core changes**: The modular design allows integration of any HTTP proxy service without modifying MediaCrawler's spider or browser automation code.

## Frequently Asked Questions

### What is the minimum code required to add a new proxy provider?

You must create a class inheriting from `ProxyProvider` and implement the `async def get_proxy(self, num: int) -> List[IpInfoModel]` method. This method handles HTTP communication with the vendor API, converts responses to `IpInfoModel` instances, and returns them. Optionally add a factory function to read API keys from environment variables, as shown in [`wandou_http_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/wandou_http_proxy.py) lines 108-124.

### How does MediaCrawler handle proxy expiration and caching?

The `IpCache` class in [`proxy/base_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/base_proxy.py) manages a Redis-backed cache with TTL support. When you call `set_ip(key, value, ex=ttl)`, the cache automatically expires entries based on the proxy provider's reported `expire_time`. The `get_proxy` method checks this cache first (returning valid entries immediately) before making expensive HTTP calls to refill the pool.

### Can I use the same provider with both httpx and CDP browser automation?

Yes. The `IpInfoModel` objects returned by `get_proxy()` contain standard fields (`ip`, `port`, `user`, `password`) compatible with both `httpx.AsyncClient` proxies and the Chrome DevTools Protocol (CDP) browser launcher in [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py). Simply format the proxy URL as `http://user:pass@ip:port` or `http://ip:port` depending on authentication requirements.

### Where should I store API credentials for third-party proxy services?

Store credentials in environment variables and read them via factory functions like `new_fastproxy_http()`. This pattern, demonstrated in [`wandou_http_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/wandou_http_proxy.py) lines 20-24, keeps sensitive keys out of version control while allowing runtime configuration through `.env` files or container orchestration secrets.