How to Integrate Custom Proxy Providers with MediaCrawler: Static, Kuaidaili, and Wandouhttp Configuration
MediaCrawler supports three proxy provider types—kuaidaili, wandouhttp, and static—and integrates them through the ProxyIpPool class in proxy/proxy_ip_pool.py, allowing you to route all HTTP requests through custom proxy endpoints by setting ENABLE_IP_PROXY, IP_PROXY_PROVIDER_NAME, and provider-specific URLs in config/base_config.py or via CLI flags.
The MediaCrawler repository provides robust proxy integration for web scraping tasks across multiple platforms. Whether you need residential rotating proxies or a dedicated static IP, understanding how to integrate custom proxy providers ensures reliable request routing and helps bypass rate limits imposed by target platforms.
Supported Proxy Provider Types
MediaCrawler implements a pluggable proxy architecture supporting three distinct providers. The system uses the IP_PROXY_PROVIDER_NAME configuration to determine which provider class to instantiate.
- kuaidaili: Commercial rotating proxy service integration
- wandouhttp: Alternative commercial proxy provider
- static: Single persistent proxy URL for consistent outbound IP addressing
Each provider must implement the interface expected by ProxyIpPool, though the static provider offers the simplest integration for self-managed or corporate proxy infrastructure.
Configuring Static Proxy Providers
The static provider (StaticProxyProvider) requires minimal setup and is ideal for environments where you control the proxy endpoint. This provider parses a single URL and returns a persistent IpInfoModel instance with an effectively infinite expiration time.
Configuration File Method
Edit config/base_config.py to enable and configure the static proxy:
# config/base_config.py
ENABLE_IP_PROXY = True # Activate proxy support globally
IP_PROXY_PROVIDER_NAME = "static" # Select static provider implementation
STATIC_PROXY_URL = "http://user:passwd@proxy.example.com:8080"
According to the MediaCrawler source code, these variables control proxy initialization in ProxyIpPool.__init__(), where the provider name is case-insensitive.
Command Line Interface Method
Override configuration file settings using CLI arguments defined in cmd_arg/arg.py:
python -m MediaCrawler.main \
--enable_ip_proxy true \
--ip_proxy_provider_name static \
--static_proxy_url "http://user:passwd@proxy.example.com:8080"
This approach is useful for containerized deployments or temporary proxy switching without modifying source files.
URL Parsing and Authentication
The StaticProxyProvider class in proxy/proxy_ip_pool.py parses STATIC_PROXY_URL using Python's urlparse module. Supported formats include:
http://your_home_domain:porthttp://user:password@your_home_domain:port
Both http and https schemes are supported. The parser extracts hostname, port, and optional credentials, defaulting to port 443 for HTTPS and 80 for HTTP when not specified.
How Proxy Integration Works Under the Hood
Understanding the internal flow helps debug connectivity issues and implement custom providers.
StaticProxyProvider Implementation
Located in proxy/proxy_ip_pool.py, the StaticProxyProvider.get_proxy() method constructs an IpInfoModel instance:
# proxy/proxy_ip_pool.py – StaticProxyProvider.get_proxy()
proxy_url = getattr(config, "STATIC_PROXY_URL", "")
parsed = urlparse(proxy_url)
scheme = parsed.scheme or "http"
ip = parsed.hostname or ""
port = parsed.port or (443 if scheme == "https" else 80)
return [
IpInfoModel(
ip=ip,
port=port,
user=unquote(parsed.username or ""),
password=unquote(parsed.password or ""),
protocol=f"{scheme}://",
expired_time_ts=int(time.time()) + 99999999, # Never expires
)
]
The far-future expired_time_ts ensures the proxy remains in the pool indefinitely, unlike commercial providers that rotate IPs based on expiration timestamps.
ProxyIpPool Integration
The ProxyIpPool class manages provider instantiation and proxy retrieval. When initialized, it reads IP_PROXY_PROVIDER_NAME and creates the appropriate provider instance. The get_or_refresh_proxy() method returns the current valid proxy, automatically handling any refresh logic required by dynamic providers.
Request-Level Proxy Application
Platform-specific clients (such as media_platform/zhihu/client.py) inherit from ProxyRefreshMixin, which automatically calls self.proxy_ip_pool.get_or_refresh_proxy() before each HTTP request:
# Inside API client implementations
proxy = await self.proxy_ip_pool.get_or_refresh_proxy()
proxy_url = f"{proxy.protocol}{proxy.ip}:{proxy.port}"
async with make_async_client(proxy=proxy_url) as client:
resp = await client.get("https://api.targetplatform.com/endpoint")
This ensures every request routes through the configured proxy without manual intervention in crawler logic.
Third-Party Provider Architecture
While static proxies suit fixed infrastructure, MediaCrawler's architecture supports commercial providers like kuaidaili and wandouhttp through the same ProxyIpPool interface. These providers typically implement API-based IP rotation, fetching fresh endpoints from provider APIs rather than parsing static URLs. The configuration pattern remains identical—set IP_PROXY_PROVIDER_NAME to the respective provider key and supply API credentials in the corresponding configuration variables.
Summary
- MediaCrawler supports three proxy provider types: kuaidaili, wandouhttp, and static, configured via
config/base_config.pyor CLI arguments incmd_arg/arg.py. - Static proxies are implemented by the
StaticProxyProviderclass inproxy/proxy_ip_pool.py, parsing URLs fromSTATIC_PROXY_URLinto persistentIpInfoModelinstances. - Authentication is supported through standard URL encoding (
user:password@host) and extracted viaurlparsein the provider implementation. - The
ProxyIpPoolclass manages provider lifecycle, whileProxyRefreshMixinautomatically applies proxies to HTTP clients in platform-specific crawlers likemedia_platform/zhihu/client.py. - Static proxies receive far-future expiration timestamps (99999999 seconds), ensuring they never rotate out of the pool during execution.
Frequently Asked Questions
What proxy formats does MediaCrawler support?
MediaCrawler accepts standard URL formats including http://host:port and http://user:password@host:port for static providers. The StaticProxyProvider uses Python's urlparse to extract scheme, hostname, port, and credentials, defaulting to port 80 for HTTP and 443 for HTTPS when ports are omitted.
How do I enable authentication for static proxies?
Embed credentials directly in the STATIC_PROXY_URL using URL-encoded syntax: http://username:password@proxy.example.com:8080. The StaticProxyProvider.get_proxy() method extracts these via parsed.username and parsed.password, applying unquote() to handle special characters, then stores them in the IpInfoModel for request-level authentication.
Can I switch between proxy providers without restarting?
Yes, when using the command line interface. The cmd_arg/arg.py module exposes --ip_proxy_provider_name and related flags, allowing runtime provider selection. However, configuration file changes in config/base_config.py require a process restart to take effect since values are loaded at initialization into the global config object.
Where does MediaCrawler apply the proxy in the request lifecycle?
Proxies are applied immediately before HTTP execution through the ProxyRefreshMixin class. Inheriting clients (such as those in media_platform/*/client.py) call self.proxy_ip_pool.get_or_refresh_proxy() to retrieve the current IpInfoModel, construct the proxy URL, and pass it to make_async_client(), ensuring every request routes through the configured endpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →