How MediaCrawler Uses Redis Cache to Manage Session State and Deduplication

MediaCrawler implements a Redis-backed caching layer to persist transient session data and prevent duplicate requests by storing session identifiers and content hashes in Redis sets with automatic TTL expiration.

MediaCrawler is an open-source multi-platform content crawler that leverages Redis to maintain runtime state across distributed crawling jobs. The implementation abstracts cache operations behind a generic AbstractCache interface while using Redis for fast, atomic storage of authentication tokens, pagination cursors, and deduplication sets.

Architecture Overview

The caching architecture follows a layered abstraction that separates the storage backend from the crawler logic:

Layer Component Role
Cache Interface cache/abs_cache.py (AbstractCache) Defines the core API – get, set, exists, delete, and expire.
Redis Implementation cache/redis_cache.py (RedisCache) Concrete class that forwards interface calls to a Redis server using redis-py.
Factory cache/cache_factory.py (CacheFactory) Instantiates the appropriate cache implementation based on configuration settings.
Session & Deduplication Logic tools/crawler_util.py, tools/cdp_browser.py Stores per-session data and checks deduplication sets during crawling operations.

The RedisCache class wraps standard Redis commands to provide the generic interface, while crawler modules interact only with the abstract methods, allowing easy switching between Redis, in-memory, or other backends.

Session State Management

When a crawling job initiates, MediaCrawler generates a unique session identifier (typically a UUID) and stores transient runtime data in Redis with a short TTL. This ensures that crashed or hung jobs do not leave stale data in the cache indefinitely.

Key session data includes:

  • Authentication tokens – Short-lived API tokens required for platform authentication
  • Pagination cursors – Next-page tokens returned by platform APIs
  • Rate-limit counters – Request counts tracked within the current rate-limit window

Session data is stored as JSON strings under keys formatted as session:{session_id} with a default expiry of 300 seconds (5 minutes).

import json
from uuid import uuid4
from cache.cache_factory import CacheFactory

# Initialize cache via factory

cache = CacheFactory.get_cache(cache_type="redis")

# Create session

session_id = str(uuid4())
session_key = f"session:{session_id}"
session_payload = {
    "auth_token": "abc123",
    "page_cursor": None,
    "request_count": 0,
}

# Store with 5-minute TTL (as implemented in cache/redis_cache.py)

cache.set(session_key, json.dumps(session_payload), ttl=300)

To retrieve the state later in the crawling lifecycle:

raw_data = cache.get(session_key)
if raw_data:
    session_data = json.loads(raw_data)
    # Resume crawling with restored state

Request Deduplication

To avoid fetching the same resource multiple times—which wastes bandwidth and risks triggering anti-scraping defenses—MediaCrawler maintains a deduplication set in Redis for each crawler instance.

The implementation uses Redis sets (O(1) complexity) to store hashed content identifiers or URLs. Before processing any content, the crawler checks the set; if the ID is absent, it processes the content and adds the ID to the set atomically.

dedup_key = "dedup:weibo"
content_id = "1234567890"

# Check if already processed (O(1) operation)

if not cache.sismember(dedup_key, content_id):
    # Process new content

    process_weibo_post(content_id)
    
    # Record in deduplication set

    cache.sadd(dedup_key, content_id)
    
    # Set 24-hour expiry to prevent unbounded growth

    cache.expire(dedup_key, 86400)
else:
    # Skip duplicate

    logger.debug("Skipping duplicate post %s", content_id)

Because Redis Set operations are atomic, this logic remains reliable even when multiple crawler workers run in parallel against the same Redis instance.

Implementation Examples

Initializing the Cache via Factory

The CacheFactory centralizes cache instantiation, reading configuration from config/db_config.py to determine which backend to use:

from cache.cache_factory import CacheFactory

# Returns RedisCache instance when configured for Redis

cache = CacheFactory.get_cache(cache_type="redis")

Source: [cache/cache_factory.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py)

Storing and Retrieving Session Data

The set method in RedisCache supports optional TTL parameters to ensure automatic cleanup:


# Store session data with explicit expiry

cache.set(f"session:{session_id}", json.dumps(payload), ttl=300)

# Retrieve and parse

data = cache.get(f"session:{session_id}")
if data:
    session = json.loads(data)

Source: [cache/redis_cache.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py)

Checking for Duplicate Content

The deduplication workflow leverages set-specific methods implemented in the RedisCache class:


# Check membership (returns boolean)

is_duplicate = cache.sismember("dedup:tiktok", video_id)

# Add to set (returns number of elements added)

cache.sadd("dedup:tiktok", video_id)

# Ensure the set expires after 24 hours

cache.expire("dedup:tiktok", 86400)

Key Files and Components

File Purpose
cache/abs_cache.py Defines the AbstractCache interface with methods get, set, exists, delete, and expire.
cache/redis_cache.py Implements RedisCache class wrapping redis.Redis client for session storage and set operations.
cache/cache_factory.py Factory pattern implementation that returns RedisCache instances based on configuration.
test/test_redis_cache.py Unit tests validating storage, retrieval, TTL enforcement, and set operations.

Summary

  • The RedisCache class in cache/redis_cache.py implements the AbstractCache interface to provide Redis-backed storage for session and deduplication data.
  • Session state persists as JSON strings with configurable TTLs (default 300 seconds), automatically expiring if jobs crash or complete without cleanup.
  • Deduplication uses Redis sets via sismember and sadd for O(1) duplicate detection, supporting parallel crawler workers through atomic operations.
  • The CacheFactory in cache/cache_factory.py decouples cache instantiation from crawler logic, allowing backend switching without code changes.
  • Automatic TTL expiration eliminates the need for manual cache cleanup, keeping memory usage bounded during long-running crawls.

Frequently Asked Questions

What is the default TTL for session data in MediaCrawler?

Session data defaults to a 5-minute (300 seconds) TTL when stored via cache.set(). This value is configurable per call, ensuring that transient authentication tokens and pagination cursors expire automatically if the crawler stalls or crashes.

How does MediaCrawler prevent processing the same content twice?

MediaCrawler uses Redis sets to track processed content IDs. Before fetching a resource, it calls sismember to check the deduplication set in O(1) time. If the ID is absent, the crawler processes the content and adds the ID via sadd, with an explicit expire call to limit the set's lifetime to 24 hours.

Which file defines the contract for cache implementations?

The abstract base class AbstractCache is defined in cache/abs_cache.py. It specifies the required interface methods—get, set, exists, delete, and expire—that all cache backends, including RedisCache, must implement.

How do I instantiate the Redis cache in MediaCrawler?

Use the CacheFactory.get_cache() method in cache/cache_factory.py, passing cache_type="redis". The factory reads the configuration from config/db_config.py and returns a properly configured RedisCache instance ready for session and deduplication operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →