How MediaCrawler Uses Redis Cache to Manage Session State and Deduplication
MediaCrawler implements a Redis-backed caching layer to persist transient session data and prevent duplicate requests by storing session identifiers and content hashes in Redis sets with automatic TTL expiration.
MediaCrawler is an open-source multi-platform content crawler that leverages Redis to maintain runtime state across distributed crawling jobs. The implementation abstracts cache operations behind a generic AbstractCache interface while using Redis for fast, atomic storage of authentication tokens, pagination cursors, and deduplication sets.
Architecture Overview
The caching architecture follows a layered abstraction that separates the storage backend from the crawler logic:
| Layer | Component | Role |
|---|---|---|
| Cache Interface | cache/abs_cache.py (AbstractCache) |
Defines the core API – get, set, exists, delete, and expire. |
| Redis Implementation | cache/redis_cache.py (RedisCache) |
Concrete class that forwards interface calls to a Redis server using redis-py. |
| Factory | cache/cache_factory.py (CacheFactory) |
Instantiates the appropriate cache implementation based on configuration settings. |
| Session & Deduplication Logic | tools/crawler_util.py, tools/cdp_browser.py |
Stores per-session data and checks deduplication sets during crawling operations. |
The RedisCache class wraps standard Redis commands to provide the generic interface, while crawler modules interact only with the abstract methods, allowing easy switching between Redis, in-memory, or other backends.
Session State Management
When a crawling job initiates, MediaCrawler generates a unique session identifier (typically a UUID) and stores transient runtime data in Redis with a short TTL. This ensures that crashed or hung jobs do not leave stale data in the cache indefinitely.
Key session data includes:
- Authentication tokens – Short-lived API tokens required for platform authentication
- Pagination cursors – Next-page tokens returned by platform APIs
- Rate-limit counters – Request counts tracked within the current rate-limit window
Session data is stored as JSON strings under keys formatted as session:{session_id} with a default expiry of 300 seconds (5 minutes).
import json
from uuid import uuid4
from cache.cache_factory import CacheFactory
# Initialize cache via factory
cache = CacheFactory.get_cache(cache_type="redis")
# Create session
session_id = str(uuid4())
session_key = f"session:{session_id}"
session_payload = {
"auth_token": "abc123",
"page_cursor": None,
"request_count": 0,
}
# Store with 5-minute TTL (as implemented in cache/redis_cache.py)
cache.set(session_key, json.dumps(session_payload), ttl=300)
To retrieve the state later in the crawling lifecycle:
raw_data = cache.get(session_key)
if raw_data:
session_data = json.loads(raw_data)
# Resume crawling with restored state
Request Deduplication
To avoid fetching the same resource multiple times—which wastes bandwidth and risks triggering anti-scraping defenses—MediaCrawler maintains a deduplication set in Redis for each crawler instance.
The implementation uses Redis sets (O(1) complexity) to store hashed content identifiers or URLs. Before processing any content, the crawler checks the set; if the ID is absent, it processes the content and adds the ID to the set atomically.
dedup_key = "dedup:weibo"
content_id = "1234567890"
# Check if already processed (O(1) operation)
if not cache.sismember(dedup_key, content_id):
# Process new content
process_weibo_post(content_id)
# Record in deduplication set
cache.sadd(dedup_key, content_id)
# Set 24-hour expiry to prevent unbounded growth
cache.expire(dedup_key, 86400)
else:
# Skip duplicate
logger.debug("Skipping duplicate post %s", content_id)
Because Redis Set operations are atomic, this logic remains reliable even when multiple crawler workers run in parallel against the same Redis instance.
Implementation Examples
Initializing the Cache via Factory
The CacheFactory centralizes cache instantiation, reading configuration from config/db_config.py to determine which backend to use:
from cache.cache_factory import CacheFactory
# Returns RedisCache instance when configured for Redis
cache = CacheFactory.get_cache(cache_type="redis")
Source: [cache/cache_factory.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/cache_factory.py)
Storing and Retrieving Session Data
The set method in RedisCache supports optional TTL parameters to ensure automatic cleanup:
# Store session data with explicit expiry
cache.set(f"session:{session_id}", json.dumps(payload), ttl=300)
# Retrieve and parse
data = cache.get(f"session:{session_id}")
if data:
session = json.loads(data)
Source: [cache/redis_cache.py](https://github.com/NanmiCoder/MediaCrawler/blob/main/cache/redis_cache.py)
Checking for Duplicate Content
The deduplication workflow leverages set-specific methods implemented in the RedisCache class:
# Check membership (returns boolean)
is_duplicate = cache.sismember("dedup:tiktok", video_id)
# Add to set (returns number of elements added)
cache.sadd("dedup:tiktok", video_id)
# Ensure the set expires after 24 hours
cache.expire("dedup:tiktok", 86400)
Key Files and Components
| File | Purpose |
|---|---|
cache/abs_cache.py |
Defines the AbstractCache interface with methods get, set, exists, delete, and expire. |
cache/redis_cache.py |
Implements RedisCache class wrapping redis.Redis client for session storage and set operations. |
cache/cache_factory.py |
Factory pattern implementation that returns RedisCache instances based on configuration. |
test/test_redis_cache.py |
Unit tests validating storage, retrieval, TTL enforcement, and set operations. |
Summary
- The
RedisCacheclass incache/redis_cache.pyimplements theAbstractCacheinterface to provide Redis-backed storage for session and deduplication data. - Session state persists as JSON strings with configurable TTLs (default 300 seconds), automatically expiring if jobs crash or complete without cleanup.
- Deduplication uses Redis sets via
sismemberandsaddfor O(1) duplicate detection, supporting parallel crawler workers through atomic operations. - The
CacheFactoryincache/cache_factory.pydecouples cache instantiation from crawler logic, allowing backend switching without code changes. - Automatic TTL expiration eliminates the need for manual cache cleanup, keeping memory usage bounded during long-running crawls.
Frequently Asked Questions
What is the default TTL for session data in MediaCrawler?
Session data defaults to a 5-minute (300 seconds) TTL when stored via cache.set(). This value is configurable per call, ensuring that transient authentication tokens and pagination cursors expire automatically if the crawler stalls or crashes.
How does MediaCrawler prevent processing the same content twice?
MediaCrawler uses Redis sets to track processed content IDs. Before fetching a resource, it calls sismember to check the deduplication set in O(1) time. If the ID is absent, the crawler processes the content and adds the ID via sadd, with an explicit expire call to limit the set's lifetime to 24 hours.
Which file defines the contract for cache implementations?
The abstract base class AbstractCache is defined in cache/abs_cache.py. It specifies the required interface methods—get, set, exists, delete, and expire—that all cache backends, including RedisCache, must implement.
How do I instantiate the Redis cache in MediaCrawler?
Use the CacheFactory.get_cache() method in cache/cache_factory.py, passing cache_type="redis". The factory reads the configuration from config/db_config.py and returns a properly configured RedisCache instance ready for session and deduplication operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →