# How MediaCrawler Uses Redis Cache to Manage Session State and Deduplication

> Discover how MediaCrawler leverages Redis cache for efficient session state management and request deduplication. Learn about Redis sets and TTL expiration for optimized performance.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: performance
- Published: 2026-07-02

---

**MediaCrawler implements a Redis-backed caching layer to persist transient session data and prevent duplicate requests by storing session identifiers and content hashes in Redis sets with automatic TTL expiration.**

MediaCrawler is an open-source multi-platform content crawler that leverages Redis to maintain runtime state across distributed crawling jobs. The implementation abstracts cache operations behind a generic `AbstractCache` interface while using Redis for fast, atomic storage of authentication tokens, pagination cursors, and deduplication sets.

## Architecture Overview

The caching architecture follows a layered abstraction that separates the storage backend from the crawler logic:

| Layer | Component | Role |
|-------|-----------|------|
| **Cache Interface** | [`cache/abs_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/abs_cache.py) (`AbstractCache`) | Defines the core API – `get`, `set`, `exists`, `delete`, and `expire`. |
| **Redis Implementation** | [`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py) (`RedisCache`) | Concrete class that forwards interface calls to a Redis server using `redis-py`. |
| **Factory** | [`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py) (`CacheFactory`) | Instantiates the appropriate cache implementation based on configuration settings. |
| **Session & Deduplication Logic** | [`tools/crawler_util.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/crawler_util.py), [`tools/cdp_browser.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/tools/cdp_browser.py) | Stores per-session data and checks deduplication sets during crawling operations. |

The `RedisCache` class wraps standard Redis commands to provide the generic interface, while crawler modules interact only with the abstract methods, allowing easy switching between Redis, in-memory, or other backends.

## Session State Management

When a crawling job initiates, MediaCrawler generates a unique session identifier (typically a UUID) and stores transient runtime data in Redis with a short TTL. This ensures that crashed or hung jobs do not leave stale data in the cache indefinitely.

**Key session data includes:**

- **Authentication tokens** – Short-lived API tokens required for platform authentication
- **Pagination cursors** – Next-page tokens returned by platform APIs
- **Rate-limit counters** – Request counts tracked within the current rate-limit window

Session data is stored as JSON strings under keys formatted as `session:{session_id}` with a default expiry of 300 seconds (5 minutes).

```python
import json
from uuid import uuid4
from cache.cache_factory import CacheFactory

# Initialize cache via factory

cache = CacheFactory.get_cache(cache_type="redis")

# Create session

session_id = str(uuid4())
session_key = f"session:{session_id}"
session_payload = {
    "auth_token": "abc123",
    "page_cursor": None,
    "request_count": 0,
}

# Store with 5-minute TTL (as implemented in cache/redis_cache.py)

cache.set(session_key, json.dumps(session_payload), ttl=300)

```

To retrieve the state later in the crawling lifecycle:

```python
raw_data = cache.get(session_key)
if raw_data:
    session_data = json.loads(raw_data)
    # Resume crawling with restored state

```

## Request Deduplication

To avoid fetching the same resource multiple times—which wastes bandwidth and risks triggering anti-scraping defenses—MediaCrawler maintains a **deduplication set** in Redis for each crawler instance.

The implementation uses Redis sets (O(1) complexity) to store hashed content identifiers or URLs. Before processing any content, the crawler checks the set; if the ID is absent, it processes the content and adds the ID to the set atomically.

```python
dedup_key = "dedup:weibo"
content_id = "1234567890"

# Check if already processed (O(1) operation)

if not cache.sismember(dedup_key, content_id):
    # Process new content

    process_weibo_post(content_id)
    
    # Record in deduplication set

    cache.sadd(dedup_key, content_id)
    
    # Set 24-hour expiry to prevent unbounded growth

    cache.expire(dedup_key, 86400)
else:
    # Skip duplicate

    logger.debug("Skipping duplicate post %s", content_id)

```

Because Redis Set operations are atomic, this logic remains reliable even when multiple crawler workers run in parallel against the same Redis instance.

## Implementation Examples

### Initializing the Cache via Factory

The `CacheFactory` centralizes cache instantiation, reading configuration from [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) to determine which backend to use:

```python
from cache.cache_factory import CacheFactory

# Returns RedisCache instance when configured for Redis

cache = CacheFactory.get_cache(cache_type="redis")

```

*Source:* [[`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py)

### Storing and Retrieving Session Data

The `set` method in `RedisCache` supports optional TTL parameters to ensure automatic cleanup:

```python

# Store session data with explicit expiry

cache.set(f"session:{session_id}", json.dumps(payload), ttl=300)

# Retrieve and parse

data = cache.get(f"session:{session_id}")
if data:
    session = json.loads(data)

```

*Source:* [[`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py)](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py)

### Checking for Duplicate Content

The deduplication workflow leverages set-specific methods implemented in the `RedisCache` class:

```python

# Check membership (returns boolean)

is_duplicate = cache.sismember("dedup:tiktok", video_id)

# Add to set (returns number of elements added)

cache.sadd("dedup:tiktok", video_id)

# Ensure the set expires after 24 hours

cache.expire("dedup:tiktok", 86400)

```

## Key Files and Components

| File | Purpose |
|------|---------|
| [`cache/abs_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/abs_cache.py) | Defines the `AbstractCache` interface with methods `get`, `set`, `exists`, `delete`, and `expire`. |
| [`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py) | Implements `RedisCache` class wrapping `redis.Redis` client for session storage and set operations. |
| [`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py) | Factory pattern implementation that returns `RedisCache` instances based on configuration. |
| [`test/test_redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/test/test_redis_cache.py) | Unit tests validating storage, retrieval, TTL enforcement, and set operations. |

## Summary

- The **`RedisCache`** class in [`cache/redis_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py) implements the **`AbstractCache`** interface to provide Redis-backed storage for session and deduplication data.
- **Session state** persists as JSON strings with configurable TTLs (default 300 seconds), automatically expiring if jobs crash or complete without cleanup.
- **Deduplication** uses Redis sets via `sismember` and `sadd` for O(1) duplicate detection, supporting parallel crawler workers through atomic operations.
- The **`CacheFactory`** in [`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py) decouples cache instantiation from crawler logic, allowing backend switching without code changes.
- Automatic TTL expiration eliminates the need for manual cache cleanup, keeping memory usage bounded during long-running crawls.

## Frequently Asked Questions

### What is the default TTL for session data in MediaCrawler?

Session data defaults to a **5-minute (300 seconds)** TTL when stored via `cache.set()`. This value is configurable per call, ensuring that transient authentication tokens and pagination cursors expire automatically if the crawler stalls or crashes.

### How does MediaCrawler prevent processing the same content twice?

MediaCrawler uses **Redis sets** to track processed content IDs. Before fetching a resource, it calls `sismember` to check the deduplication set in O(1) time. If the ID is absent, the crawler processes the content and adds the ID via `sadd`, with an explicit `expire` call to limit the set's lifetime to 24 hours.

### Which file defines the contract for cache implementations?

The abstract base class **`AbstractCache`** is defined in [`cache/abs_cache.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/abs_cache.py). It specifies the required interface methods—`get`, `set`, `exists`, `delete`, and `expire`—that all cache backends, including `RedisCache`, must implement.

### How do I instantiate the Redis cache in MediaCrawler?

Use the **`CacheFactory.get_cache()`** method in [`cache/cache_factory.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py), passing `cache_type="redis"`. The factory reads the configuration from [`config/db_config.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/config/db_config.py) and returns a properly configured `RedisCache` instance ready for session and deduplication operations.