LiteLLM Caching Layer Architecture: Redis vs In-Memory Implementation Guide
LiteLLM implements a dual-backend caching architecture that caches both HTTP client objects and model responses using either an in-process dictionary or a Redis server, both conforming to the abstract BaseCache interface for seamless interchangeability.
LiteLLM provides a unified Python SDK for interacting with over 100 large language model providers. To optimize performance and reduce redundant API latency, the library incorporates a sophisticated caching layer that persists both client connections and completion responses. This technical deep dive examines the BerriAI/litellm repository's caching implementation, contrasting the zero-configuration in-memory backend with the production-grade Redis alternative.
How LiteLLM Caching Works
The caching layer in LiteLLM is built around an abstract BaseCache class defined in litellm/caching/base_cache.py. This interface standardizes three core operations: get_cache(), set_cache(), and flush_cache(). Both the in-memory and Redis implementations inherit from this base, enabling runtime swapping without code changes. When a completion request is initiated, LiteLLM generates a deterministic cache key from the model name, provider endpoint, and request payload, then queries the configured backend—either a local Python dictionary or a remote Redis store—to retrieve cached clients or previous responses. An auxiliary pattern using InMemoryFile for batch operations is also demonstrated in litellm/router_utils/batch_utils.py, illustrating the library's consistent approach to process-local storage.
In-Memory Cache Implementation
Architecture and Scope
The in-memory backend, implemented in litellm/caching/in_memory_cache.py, provides a process-local LRU (Least Recently Used) cache using a standard Python dict. This implementation is exposed globally through the LLMClientCache singleton instantiated in litellm/__init__.py as in_memory_llm_clients_cache. The cache lives only as long as the Python interpreter runs, making it suitable for single-process scripts, unit tests, and development environments where persistence across restarts is not required.
Usage Example
No external services or environment variables are required to use the in-memory cache. The singleton is created automatically upon importing the relevant modules:
import litellm
# Access the singleton cache instance
cache = litellm.in_memory_llm_clients_cache
# Check for existing OpenAI client using the cache key format
client = cache.get_cache("openai:gpt-4o")
if client is None:
client = litellm.OpenAI()
cache.set_cache("openai:gpt-4o", client)
# Subsequent calls reuse the cached HTTP connection
response = litellm.completion(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello"}]
)
Entries are evicted based on LRU policy or cleared manually using cache.flush_cache(), which is useful for testing scenarios requiring clean state.
Redis Cache Implementation
Architecture and Serialization
For production deployments requiring cache sharing across multiple workers or containers, litellm/caching/redis_cache.py implements the RedisCache class. This backend serializes values using Python's pickle module before storing them in Redis, enabling complex objects like HTTP clients to be persisted across process boundaries. Unlike the in-memory variant, this implementation supports configurable Time-To-Live (TTL) eviction, with a default expiration of 1 hour per entry.
Configuration and Environment Variables
Switching to Redis requires setting specific environment variables before initializing the cache:
import os
import litellm
# Configure Redis backend
os.environ["LLM_CACHE"] = "redis"
os.environ["REDIS_HOST"] = "redis.example.com"
os.environ["REDIS_PORT"] = "6379"
os.environ["REDIS_PASSWORD"] = "secure-password"
# Lazy instantiation on first access
cache = litellm.redis_cache
# Store with custom TTL (30 minutes)
cache.set_cache(
key="openai:gpt-4o",
value=litellm.OpenAI(),
ttl_secs=1800
)
The RedisCache also provides fallback mechanisms to the in-memory cache for certain client objects, ensuring resilience if Redis connectivity is temporarily unavailable.
Cache Key Generation and Lookup Flow
The caching mechanism operates through a deterministic four-step process visible in the integration points such as litellm/llms/openai/common_utils.py:
- Key Generation – LiteLLM constructs a unique cache key combining the provider endpoint, model identifier, and hashed request parameters.
- Lookup – The system queries either
InMemoryCache.get_cache()orRedisCache.get_cache()depending on theLLM_CACHEenvironment variable. - Hit Handling – If a valid entry exists (within TTL for Redis), the deserialized client or cached response returns immediately, bypassing external API calls.
- Miss Handling – On cache miss, LiteLLM instantiates a new client or fetches from the provider, then stores the result via
set_cache()for future requests.
This architecture ensures that both client connection overhead and expensive LLM inference costs are minimized across repeated identical calls.
Performance Characteristics and Trade-offs
In-Memory Cache:
- Zero network latency: Direct dictionary lookups provide microsecond-level access times.
- No operational overhead: Requires no external services or configuration management.
- Process isolation: Each worker maintains independent cache state, potentially duplicating memory usage and API warmup costs across multiple instances.
- Volatility: All cached data is lost on process restart or deployment, as entries are stored only in the Python process heap.
Redis Cache:
- Distributed consistency: All workers, containers, and serverless functions share identical cache state, eliminating redundant warmups and ensuring cache hits across the entire infrastructure.
- Persistence: Survives process restarts and supports serverless runtimes that frequently recycle execution environments.
- Operational complexity: Requires running and maintaining a Redis server (or managed service like AWS ElastiCache), adding infrastructure overhead.
- Serialization overhead:
pickleserialization/deserialization and network round-trips add marginal latency compared to local dictionary access, though this is typically negligible relative to LLM API latency.
Summary
- LiteLLM implements a pluggable caching architecture via the
BaseCacheabstraction inlitellm/caching/base_cache.py, supporting seamless backend swaps through environment variables. - The in-memory cache (
litellm/caching/in_memory_cache.py) provides zero-configuration, process-local storage ideal for development and single-instance deployments via theLLMClientCachesingleton. - The Redis cache (
litellm/caching/redis_cache.py) enables distributed, persistent caching across microservices usingpickleserialization and configurable TTL via thecache_ttl_secsparameter. - Cache keys are deterministically generated from request parameters, ensuring identical calls receive cached responses regardless of backend selection.
- Configuration is environment-driven via
LLM_CACHE,REDIS_HOST,REDIS_PORT, andREDIS_PASSWORDvariables, with lazy instantiation on first use.
Frequently Asked Questions
How do I switch from in-memory caching to Redis in LiteLLM?
Set the LLM_CACHE environment variable to "redis" and provide connection details via REDIS_HOST, REDIS_PORT, and REDIS_PASSWORD. The RedisCache instance is lazily created on first access, allowing runtime backend switching without modifying application code or cache lookup logic.
What is the default TTL for cached entries in LiteLLM's Redis implementation?
The default TTL is 1 hour (3600 seconds). You can override this per-entry using the ttl_secs parameter when calling set_cache(), or modify the default behavior in the RedisCache class implementation.
Does LiteLLM cache both LLM responses and HTTP client connections?
Yes. The caching layer handles both client objects (such as OpenAI HTTP clients cached to avoid connection establishment overhead) and model responses (to avoid redundant API calls for identical prompts). The specific caching strategy varies by provider integration, with client caching logic visible in files like litellm/llms/openai/common_utils.py.
How can I manually clear the cache during testing?
Both backends implement a flush_cache() method. For the in-memory cache, call litellm.in_memory_llm_clients_cache.flush_cache(). For Redis, use litellm.redis_cache.flush_cache(). This is particularly useful in CI/CD pipelines or unit tests requiring isolated state between test cases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →