Setting Up Distributed Crawling with Redis Cache in MediaCrawler

MediaCrawler enables distributed crawling by using Redis as a central coordination layer, allowing multiple workers to share login sessions, throttle requests globally, and deduplicate tasks across different machines.

MediaCrawler is a multi‑platform web crawling framework that scales horizontally through Redis‑based coordination. By leveraging the RedisCache wrapper and configuration flags in config/base_config.py, you can deploy worker pools across separate hosts while maintaining shared state for authentication cookies and rate‑limiting locks. This architecture eliminates redundant QR‑code logins and prevents race conditions when multiple instances target the same content sources.

How Redis Fits Into the Architecture

The framework implements a thin async‑compatible Redis client that acts as a universal backing store for distributed coordination.

The RedisCache Wrapper

Located in database/redis_cache.py, the RedisCache class provides a generic key‑value interface (set, get, expire, delete) that the rest of the codebase consumes. When ENABLE_REDIS_CACHE is set to True in config/base_config.py, the framework initializes this client via get_redis_client() from database/db.py, which reads the connection string from config/db_config.py.

Distributed Coordination Patterns

MediaCrawler uses Redis for three critical coordination tasks:

  • Login State Sharing – After a worker authenticates with a platform (e.g., Zhihu, Xiaohongshu), it serializes the session cookies to Redis with a TTL, allowing other workers to reuse the session without re‑authenticating.
  • Global Request Throttling – Workers implement distributed rate limiting using Redis SETNX locks to enforce global request windows across all instances.
  • Task Deduplication – URLs and pagination cursors are stored with expiration times, ensuring that if one worker fails, another can resume from the same checkpoint.

Configuration Steps for Distributed Mode

Enable distributed crawling by configuring the Redis connection and cache flags.

1. Configure the Redis Connection

Edit config/db_config.py to point to your Redis instance:


# config/db_config.py

REDIS_URL = "redis://localhost:6379/0"  # Update with your Redis endpoint

2. Enable the Distributed Cache

Set the feature flag in config/base_config.py:


# config/base_config.py

ENABLE_REDIS_CACHE = True  # Activates Redis backing for sessions and throttling

3. Launch Multiple Workers

With the configuration in place, start multiple worker processes. Each worker runs main.py independently and automatically connects to the shared Redis instance to coordinate work:


# Terminal 1

uv run main.py --platform zhihu --login_type qrcode

# Terminal 2 (or separate host)

uv run main.py --platform zhihu --login_type qrcode

Implementation Patterns

The following patterns demonstrate how MediaCrawler uses Redis to coordinate distributed workers.

Sharing Login Sessions Across Workers

After authentication, workers store cookies in Redis so that peers can retrieve them:

import json
from database.db import get_redis_client

async def store_session(platform: str, cookies: dict):
    """Store login cookies for other workers to reuse."""
    redis = await get_redis_client()
    key = f"{platform}:session"
    await redis.set(key, json.dumps(cookies), ex=3600)  # 1 hour TTL

Other workers retrieve the session before creating browser contexts, eliminating redundant logins.

Distributed Request Throttling

To enforce global rate limits, workers acquire Redis locks before sending requests:

import asyncio
from database.db import get_redis_client

async def throttled_request(url: str):
    """Ensure only one worker hits this URL at a time across the cluster."""
    redis = await get_redis_client()
    lock_key = f"lock:{url}"
    
    while not await redis.setnx(lock_key, "1"):
        await asyncio.sleep(0.1)  # Wait for lock release

    
    await redis.expire(lock_key, 5)  # Auto‑release after 5 seconds

    
    try:
        # Actual request logic here (Playwright, httpx, etc.)

        return await fetch_url(url)
    finally:
        await redis.delete(lock_key)  # Explicit cleanup

Consuming Tasks From a Redis Queue

For URL queue coordination, workers pop tasks from a Redis list to ensure each URL is processed exactly once:

async def process_queue():
    redis = await get_redis_client()
    
    while True:
        url = await redis.lpop("crawl:queue")
        if url is None:
            break  # Queue exhausted

        
        decoded_url = url.decode()
        await throttled_request(decoded_url)

Key Source Files

These files implement the distributed caching infrastructure:

  • config/db_config.py – Contains REDIS_URL configuration for the connection string.
  • config/base_config.py – Defines ENABLE_REDIS_CACHE toggle to activate distributed mode.
  • database/db.py – Implements get_redis_client() factory that creates the shared async Redis connection.
  • database/redis_cache.py – Provides the RedisCache wrapper with set, get, expire, and delete methods.
  • proxy/proxy_ip_pool.py – Demonstrates Redis list operations for managing shared resources (IP proxy pools).
  • test/test_redis_cache.py – Unit tests verifying the Redis wrapper functionality.
  • main.py – Entry point that respects ENABLE_REDIS_CACHE and initializes workers.

Summary

  • MediaCrawler uses Redis as a central store for sessions, locks, and temporary crawl results when ENABLE_REDIS_CACHE is enabled.
  • Configure REDIS_URL in config/db_config.py to point all workers to the same Redis instance.
  • Share login state by storing cookies in Redis with TTL, preventing duplicate authentications across workers.
  • Throttle globally using Redis SETNX locks to enforce rate limits across the entire cluster.
  • Scale horizontally by launching multiple main.py instances on different machines connected to the same Redis endpoint.

Frequently Asked Questions

How do I verify that Redis caching is actually working?

Check that ENABLE_REDIS_CACHE is set to True in config/base_config.py and that your REDIS_URL in config/db_config.py is accessible from all worker nodes. You can monitor Redis using redis-cli MONITOR to confirm that keys are being written when workers log in or throttle requests.

Can I use Redis Sentinel or Cluster mode with MediaCrawler?

The get_redis_client() function in database/db.py uses aioredis (or compatible async Redis clients). You can modify the connection initialization to support Sentinel by passing a sentinel configuration instead of a single URL, though you may need to extend the RedisCache wrapper to handle Sentinel discovery logic.

What happens if the Redis connection drops during crawling?

The RedisCache wrapper does not implement automatic reconnection logic by default. If the connection fails, workers will raise connection errors and stop. For production deployments, implement a connection pool with retry logic in database/db.py or ensure your Redis instance is highly available behind a failover proxy.

Is Redis required for single‑machine crawling?

No. When ENABLE_REDIS_CACHE is False (the default), MediaCrawler operates in standalone mode using in‑memory storage via ExpiringLocalCache in database/expiring_local_cache.py. Redis is only required when you need to coordinate multiple workers across different processes or machines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →