Setting Up Distributed Crawling with Redis Cache in MediaCrawler
MediaCrawler enables distributed crawling by using Redis as a central coordination layer, allowing multiple workers to share login sessions, throttle requests globally, and deduplicate tasks across different machines.
MediaCrawler is a multi‑platform web crawling framework that scales horizontally through Redis‑based coordination. By leveraging the RedisCache wrapper and configuration flags in config/base_config.py, you can deploy worker pools across separate hosts while maintaining shared state for authentication cookies and rate‑limiting locks. This architecture eliminates redundant QR‑code logins and prevents race conditions when multiple instances target the same content sources.
How Redis Fits Into the Architecture
The framework implements a thin async‑compatible Redis client that acts as a universal backing store for distributed coordination.
The RedisCache Wrapper
Located in database/redis_cache.py, the RedisCache class provides a generic key‑value interface (set, get, expire, delete) that the rest of the codebase consumes. When ENABLE_REDIS_CACHE is set to True in config/base_config.py, the framework initializes this client via get_redis_client() from database/db.py, which reads the connection string from config/db_config.py.
Distributed Coordination Patterns
MediaCrawler uses Redis for three critical coordination tasks:
- Login State Sharing – After a worker authenticates with a platform (e.g., Zhihu, Xiaohongshu), it serializes the session cookies to Redis with a TTL, allowing other workers to reuse the session without re‑authenticating.
- Global Request Throttling – Workers implement distributed rate limiting using Redis
SETNXlocks to enforce global request windows across all instances. - Task Deduplication – URLs and pagination cursors are stored with expiration times, ensuring that if one worker fails, another can resume from the same checkpoint.
Configuration Steps for Distributed Mode
Enable distributed crawling by configuring the Redis connection and cache flags.
1. Configure the Redis Connection
Edit config/db_config.py to point to your Redis instance:
# config/db_config.py
REDIS_URL = "redis://localhost:6379/0" # Update with your Redis endpoint
2. Enable the Distributed Cache
Set the feature flag in config/base_config.py:
# config/base_config.py
ENABLE_REDIS_CACHE = True # Activates Redis backing for sessions and throttling
3. Launch Multiple Workers
With the configuration in place, start multiple worker processes. Each worker runs main.py independently and automatically connects to the shared Redis instance to coordinate work:
# Terminal 1
uv run main.py --platform zhihu --login_type qrcode
# Terminal 2 (or separate host)
uv run main.py --platform zhihu --login_type qrcode
Implementation Patterns
The following patterns demonstrate how MediaCrawler uses Redis to coordinate distributed workers.
Sharing Login Sessions Across Workers
After authentication, workers store cookies in Redis so that peers can retrieve them:
import json
from database.db import get_redis_client
async def store_session(platform: str, cookies: dict):
"""Store login cookies for other workers to reuse."""
redis = await get_redis_client()
key = f"{platform}:session"
await redis.set(key, json.dumps(cookies), ex=3600) # 1 hour TTL
Other workers retrieve the session before creating browser contexts, eliminating redundant logins.
Distributed Request Throttling
To enforce global rate limits, workers acquire Redis locks before sending requests:
import asyncio
from database.db import get_redis_client
async def throttled_request(url: str):
"""Ensure only one worker hits this URL at a time across the cluster."""
redis = await get_redis_client()
lock_key = f"lock:{url}"
while not await redis.setnx(lock_key, "1"):
await asyncio.sleep(0.1) # Wait for lock release
await redis.expire(lock_key, 5) # Auto‑release after 5 seconds
try:
# Actual request logic here (Playwright, httpx, etc.)
return await fetch_url(url)
finally:
await redis.delete(lock_key) # Explicit cleanup
Consuming Tasks From a Redis Queue
For URL queue coordination, workers pop tasks from a Redis list to ensure each URL is processed exactly once:
async def process_queue():
redis = await get_redis_client()
while True:
url = await redis.lpop("crawl:queue")
if url is None:
break # Queue exhausted
decoded_url = url.decode()
await throttled_request(decoded_url)
Key Source Files
These files implement the distributed caching infrastructure:
config/db_config.py– ContainsREDIS_URLconfiguration for the connection string.config/base_config.py– DefinesENABLE_REDIS_CACHEtoggle to activate distributed mode.database/db.py– Implementsget_redis_client()factory that creates the shared async Redis connection.database/redis_cache.py– Provides theRedisCachewrapper withset,get,expire, anddeletemethods.proxy/proxy_ip_pool.py– Demonstrates Redis list operations for managing shared resources (IP proxy pools).test/test_redis_cache.py– Unit tests verifying the Redis wrapper functionality.main.py– Entry point that respectsENABLE_REDIS_CACHEand initializes workers.
Summary
- MediaCrawler uses Redis as a central store for sessions, locks, and temporary crawl results when
ENABLE_REDIS_CACHEis enabled. - Configure
REDIS_URLinconfig/db_config.pyto point all workers to the same Redis instance. - Share login state by storing cookies in Redis with TTL, preventing duplicate authentications across workers.
- Throttle globally using Redis
SETNXlocks to enforce rate limits across the entire cluster. - Scale horizontally by launching multiple
main.pyinstances on different machines connected to the same Redis endpoint.
Frequently Asked Questions
How do I verify that Redis caching is actually working?
Check that ENABLE_REDIS_CACHE is set to True in config/base_config.py and that your REDIS_URL in config/db_config.py is accessible from all worker nodes. You can monitor Redis using redis-cli MONITOR to confirm that keys are being written when workers log in or throttle requests.
Can I use Redis Sentinel or Cluster mode with MediaCrawler?
The get_redis_client() function in database/db.py uses aioredis (or compatible async Redis clients). You can modify the connection initialization to support Sentinel by passing a sentinel configuration instead of a single URL, though you may need to extend the RedisCache wrapper to handle Sentinel discovery logic.
What happens if the Redis connection drops during crawling?
The RedisCache wrapper does not implement automatic reconnection logic by default. If the connection fails, workers will raise connection errors and stop. For production deployments, implement a connection pool with retry logic in database/db.py or ensure your Redis instance is highly available behind a failover proxy.
Is Redis required for single‑machine crawling?
No. When ENABLE_REDIS_CACHE is False (the default), MediaCrawler operates in standalone mode using in‑memory storage via ExpiringLocalCache in database/expiring_local_cache.py. Redis is only required when you need to coordinate multiple workers across different processes or machines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →