# How to Configure an IP Proxy Pool with the KuaiDaili Provider in MediaCrawler

> Learn how to configure an IP proxy pool with the KuaiDaili provider in MediaCrawler. Easily set up credentials and environment variables for efficient proxy usage.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-06-30

---

**To configure the IP proxy pool with the KuaiDaili provider, set `ENABLE_IP_PROXY = True` and `IP_PROXY_PROVIDER_NAME = "kuaidaili"` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), then export the four KuaiDaili credential environment variables (`KDL_SECERT_ID`, `KDL_SIGNATURE`, `KDL_USER_NAME`, `KDL_USER_PWD`) before launching the crawler.**

MediaCrawler is an open-source asynchronous crawling framework that ships with a built-in proxy rotation system. When you configure the IP proxy pool with the KuaiDaili provider, the framework automatically fetches fresh residential or datacenter IPs from the KuaiDaili API, caches them with expiration timestamps, and injects them into HTTP requests. This setup relies on environment variables for secure credential storage and specific global configuration flags to activate the provider.

## Configure the KuaiDaili Provider

### Enable the Proxy Subsystem

Start by activating the proxy feature in the global configuration. Open [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and set the required flags:

```python

# config/base_config.py

ENABLE_IP_PROXY = True
IP_PROXY_PROVIDER_NAME = "kuaidaili"  # Must match ProviderNameEnum.KUAI_DAI_LI

IP_PROXY_POOL_COUNT = 5               # Number of proxies to maintain (default: 2)

```

These values are read when `create_ip_pool()` (defined in [`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py)) initializes the crawler. The `IP_PROXY_PROVIDER_NAME` string maps to the `new_kuai_daili_proxy()` factory function via the `IpProxyProvider` registry.

### Set KuaiDaili API Credentials

The provider implementation in [`proxy/providers/kuaidl_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/providers/kuaidl_proxy.py) reads credentials exclusively from environment variables. It accepts both uppercase and lowercase variants:

- `KDL_SECERT_ID` (or `kdl_secret_id`): Your KuaiDaili secret ID
- `KDL_SIGNATURE` (or `kdl_signature`): The request signature string
- `KDL_USER_NAME` (or `kdl_user_name`): Account username
- `KDL_USER_PWD` (or `kdl_user_pwd`): Account password

If any variable is missing, the code substitutes a placeholder string (e.g., `"your_kuaidaili_secret_id"`), causing an authentication error when the pool first attempts to fetch proxies. Export these in your shell before running the crawler:

```bash
export KDL_SECERT_ID="your_actual_secret_id"
export KDL_SIGNATURE="your_actual_signature"
export KDL_USER_NAME="your_username"
export KDL_USER_PWD="your_password"

```

### Adjust Pool Validation (Optional)

When calling `create_ip_pool()`, you can enable real-time proxy validation by passing `enable_validate_ip=True`. This triggers the `_is_valid_proxy` check, which verifies that each fetched proxy is reachable before adding it to the pool. For high-speed crawling where latency matters, you can disable this validation by setting the parameter to `False`.

## Initialize and Use the Proxy Pool

MediaCrawler initializes the pool automatically for standard crawlers, but you can manually instantiate it in custom scripts. The `ProxyIpPool` class orchestrates proxy retrieval, caching, and rotation.

```python
import asyncio
from proxy.proxy_ip_pool import create_ip_pool
from tools.httpx_util import make_async_client

async def fetch_with_rotation():
    # Initialize pool with 5 proxies; validation enabled

    ip_pool = await create_ip_pool(ip_pool_count=5, enable_validate_ip=True)
    
    # Get a valid proxy (auto-refreshes if expired)

    proxy_info = await ip_pool.get_or_refresh_proxy()
    
    # Construct proxy URL from IpInfoModel attributes

    proxy_url = f"http://{proxy_info.user}:{proxy_info.password}@{proxy_info.ip}:{proxy_info.port}"
    
    async with make_async_client(proxy=proxy_url) as client:
        resp = await client.get("https://httpbin.org/ip")
        print(f"External IP: {resp.json()}")

asyncio.run(fetch_with_rotation())

```

The `IpInfoModel` object returned by `get_or_refresh_proxy()` contains `ip`, `port`, `user`, `password`, and an absolute expiration timestamp (`expired_time_ts`). The pool stores these objects in an `IpCache` and automatically replenishes the pool when proxies expire or when the count drops below `IP_PROXY_POOL_COUNT`.

## Architecture and Key Implementation Files

Understanding the component hierarchy ensures you can debug configuration issues effectively:

- **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)**: Contains global toggles (`ENABLE_IP_PROXY`, `IP_PROXY_PROVIDER_NAME`, `IP_PROXY_POOL_COUNT`) that dictate whether `create_ip_pool()` builds a proxy-enabled crawler.

- **[`proxy/providers/kuaidl_proxy.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/providers/kuaidl_proxy.py)**: Implements the `KuaiDaiLiProxy` class and the `new_kuaidaili_proxy()` factory. This file handles credential loading from environment variables, API request construction, JSON response parsing, and wrapping results in `IpInfoModel` instances. It also manages the `IpCache` integration to avoid redundant API calls.

- **[`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py)**: Defines the `ProxyIpPool` class and the `create_ip_pool()` entry point. This module selects the correct provider from the `IpProxyProvider` dictionary, validates proxies (if enabled), and serves fresh IPs via `get_or_refresh_proxy()`.

- **[`proxy/types.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/types.py)**: Houses the Pydantic `IpInfoModel` used throughout the system to standardize proxy metadata and the `ProviderNameEnum` for provider validation.

The execution flow proceeds as follows: `create_ip_pool()` reads the global config → selects `kuaidaili` → calls `new_kuaidaili_proxy()` → instantiates `KuaiDaiLiProxy` → fetches proxies via `get_proxy(pool_size)` → wraps responses in `IpInfoModel` → caches them for the duration of their validity period.

## Summary

- Activate the feature by setting **`ENABLE_IP_PROXY = True`** and **`IP_PROXY_PROVIDER_NAME = "kuaidaili"`** in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py).
- Provide credentials via the **`KDL_SECERT_ID`**, **`KDL_SIGNATURE`**, **`KDL_USER_NAME`**, and **`KDL_USER_PWD`** environment variables (uppercase preferred, lowercase accepted).
- Control pool size with **`IP_PROXY_POOL_COUNT`** (default 2) and toggle validation via the **`enable_validate_ip`** parameter in `create_ip_pool()`.
- Retrieve active proxies using **`await ip_pool.get_or_refresh_proxy()`**, which returns an **`IpInfoModel`** containing host, port, credentials, and expiration timestamp.
- Pass the proxy URL to **`make_async_client(proxy=...)`** to route HTTP traffic through the rotated IP.

## Frequently Asked Questions

### What happens if I forget to set the environment variables?

The `KuaiDaiLiProxy` class substitutes placeholder strings like `"your_kuaidaili_secret_id"` for missing variables. When the pool first attempts to populate itself, the KuaiDaili API rejects these invalid credentials, raising a clear authentication error that indicates which environment variable is required.

### Can I use a static proxy instead of KuaiDaili for testing?

Yes. Set `IP_PROXY_PROVIDER_NAME = "static"` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and export `STATIC_PROXY_URL="http://user:pass@host:port"`. The `StaticProxyProvider` class parses this URL into an `IpInfoModel` with no expiration, allowing you to test crawler logic without consuming paid API credits.

### How does the pool handle expired proxies?

Each `IpInfoModel` carries an `expired_time_ts` timestamp set when the proxy is fetched. The `get_or_refresh_proxy()` method checks this value; if the current time exceeds the timestamp, it triggers a fresh API call to KuaiDaili via the provider’s `get_proxy()` method and updates the internal cache with new IPs.

### What is the difference between `get_proxy()` and `get_or_refresh_proxy()`?

`get_proxy()` returns the next available proxy from the internal list without checking expiration, suitable for high-throughput scenarios where you want minimal overhead. `get_or_refresh_proxy()` validates the timestamp and automatically replenishes the pool if the proxy is expired or the pool is empty, ensuring you never use a dead IP.