# Setting Up Distributed Crawling with Redis Cache in MediaCrawler

> Set up distributed crawling with Redis cache in MediaCrawler. Coordinate multiple workers, share sessions, throttle requests, and deduplicate tasks globally for efficient web scraping.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: how-to-guide
- Published: 2026-07-31

---

**MediaCrawler enables distributed crawling by using Redis as a central coordination layer, allowing multiple workers to share login sessions, throttle requests globally, and deduplicate tasks across different machines.**

MediaCrawler is a multi‑platform web crawling framework that scales horizontally through Redis‑based coordination. By leveraging the `RedisCache` wrapper and configuration flags in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), you can deploy worker pools across separate hosts while maintaining shared state for authentication cookies and rate‑limiting locks. This architecture eliminates redundant QR‑code logins and prevents race conditions when multiple instances target the same content sources.

## How Redis Fits Into the Architecture

The framework implements a thin async‑compatible Redis client that acts as a universal backing store for distributed coordination.

### The RedisCache Wrapper

Located in [`database/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/redis_cache.py), the `RedisCache` class provides a generic key‑value interface (`set`, `get`, `expire`, `delete`) that the rest of the codebase consumes. When `ENABLE_REDIS_CACHE` is set to `True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py), the framework initializes this client via `get_redis_client()` from [`database/db.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db.py), which reads the connection string from [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py).

### Distributed Coordination Patterns

MediaCrawler uses Redis for three critical coordination tasks:

- **Login State Sharing** – After a worker authenticates with a platform (e.g., Zhihu, Xiaohongshu), it serializes the session cookies to Redis with a TTL, allowing other workers to reuse the session without re‑authenticating.
- **Global Request Throttling** – Workers implement distributed rate limiting using Redis `SETNX` locks to enforce global request windows across all instances.
- **Task Deduplication** – URLs and pagination cursors are stored with expiration times, ensuring that if one worker fails, another can resume from the same checkpoint.

## Configuration Steps for Distributed Mode

Enable distributed crawling by configuring the Redis connection and cache flags.

### 1. Configure the Redis Connection

Edit [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) to point to your Redis instance:

```python

# config/db_config.py

REDIS_URL = "redis://localhost:6379/0"  # Update with your Redis endpoint

```

### 2. Enable the Distributed Cache

Set the feature flag in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py):

```python

# config/base_config.py

ENABLE_REDIS_CACHE = True  # Activates Redis backing for sessions and throttling

```

### 3. Launch Multiple Workers

With the configuration in place, start multiple worker processes. Each worker runs [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) independently and automatically connects to the shared Redis instance to coordinate work:

```bash

# Terminal 1

uv run main.py --platform zhihu --login_type qrcode

# Terminal 2 (or separate host)

uv run main.py --platform zhihu --login_type qrcode

```

## Implementation Patterns

The following patterns demonstrate how MediaCrawler uses Redis to coordinate distributed workers.

### Sharing Login Sessions Across Workers

After authentication, workers store cookies in Redis so that peers can retrieve them:

```python
import json
from database.db import get_redis_client

async def store_session(platform: str, cookies: dict):
    """Store login cookies for other workers to reuse."""
    redis = await get_redis_client()
    key = f"{platform}:session"
    await redis.set(key, json.dumps(cookies), ex=3600)  # 1 hour TTL

```

Other workers retrieve the session before creating browser contexts, eliminating redundant logins.

### Distributed Request Throttling

To enforce global rate limits, workers acquire Redis locks before sending requests:

```python
import asyncio
from database.db import get_redis_client

async def throttled_request(url: str):
    """Ensure only one worker hits this URL at a time across the cluster."""
    redis = await get_redis_client()
    lock_key = f"lock:{url}"
    
    while not await redis.setnx(lock_key, "1"):
        await asyncio.sleep(0.1)  # Wait for lock release

    
    await redis.expire(lock_key, 5)  # Auto‑release after 5 seconds

    
    try:
        # Actual request logic here (Playwright, httpx, etc.)

        return await fetch_url(url)
    finally:
        await redis.delete(lock_key)  # Explicit cleanup

```

### Consuming Tasks From a Redis Queue

For URL queue coordination, workers pop tasks from a Redis list to ensure each URL is processed exactly once:

```python
async def process_queue():
    redis = await get_redis_client()
    
    while True:
        url = await redis.lpop("crawl:queue")
        if url is None:
            break  # Queue exhausted

        
        decoded_url = url.decode()
        await throttled_request(decoded_url)

```

## Key Source Files

These files implement the distributed caching infrastructure:

- **[`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py)** – Contains `REDIS_URL` configuration for the connection string.
- **[`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py)** – Defines `ENABLE_REDIS_CACHE` toggle to activate distributed mode.
- **[`database/db.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db.py)** – Implements `get_redis_client()` factory that creates the shared async Redis connection.
- **[`database/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/redis_cache.py)** – Provides the `RedisCache` wrapper with `set`, `get`, `expire`, and `delete` methods.
- **[`proxy/proxy_ip_pool.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/proxy/proxy_ip_pool.py)** – Demonstrates Redis list operations for managing shared resources (IP proxy pools).
- **[`test/test_redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/test/test_redis_cache.py)** – Unit tests verifying the Redis wrapper functionality.
- **[`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)** – Entry point that respects `ENABLE_REDIS_CACHE` and initializes workers.

## Summary

- **MediaCrawler uses Redis** as a central store for sessions, locks, and temporary crawl results when `ENABLE_REDIS_CACHE` is enabled.
- **Configure `REDIS_URL`** in [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) to point all workers to the same Redis instance.
- **Share login state** by storing cookies in Redis with TTL, preventing duplicate authentications across workers.
- **Throttle globally** using Redis `SETNX` locks to enforce rate limits across the entire cluster.
- **Scale horizontally** by launching multiple [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) instances on different machines connected to the same Redis endpoint.

## Frequently Asked Questions

### How do I verify that Redis caching is actually working?

Check that `ENABLE_REDIS_CACHE` is set to `True` in [`config/base_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/base_config.py) and that your `REDIS_URL` in [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) is accessible from all worker nodes. You can monitor Redis using `redis-cli MONITOR` to confirm that keys are being written when workers log in or throttle requests.

### Can I use Redis Sentinel or Cluster mode with MediaCrawler?

The `get_redis_client()` function in [`database/db.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db.py) uses `aioredis` (or compatible async Redis clients). You can modify the connection initialization to support Sentinel by passing a `sentinel` configuration instead of a single URL, though you may need to extend the `RedisCache` wrapper to handle Sentinel discovery logic.

### What happens if the Redis connection drops during crawling?

The `RedisCache` wrapper does not implement automatic reconnection logic by default. If the connection fails, workers will raise connection errors and stop. For production deployments, implement a connection pool with retry logic in [`database/db.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/db.py) or ensure your Redis instance is highly available behind a failover proxy.

### Is Redis required for single‑machine crawling?

No. When `ENABLE_REDIS_CACHE` is `False` (the default), MediaCrawler operates in standalone mode using in‑memory storage via `ExpiringLocalCache` in [`database/expiring_local_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/database/expiring_local_cache.py). Redis is only required when you need to coordinate multiple workers across different processes or machines.