How to Configure LiteLLM Caching: Redis, In-Memory, and Disk Backends Explained
LiteLLM offers a pluggable caching architecture built on the BaseCache abstraction that supports in-memory, Redis, disk, and dual-cache backends to reduce LLM token costs and latency through configurable TTL and cache control policies.
LiteLLM enables intelligent request/response caching to minimize API costs and improve response times across multiple LLM providers. Whether you are running the LiteLLM Proxy or integrating caching directly into Python applications, understanding how to configure LiteLLM caching is essential for production deployments. The BerriAI/litellm repository implements this through a modular system centered on the BaseCache interface in litellm/caching/base_cache.py.
Core Caching Architecture
The LiteLLM caching system relies on a clean separation between the abstraction layer and concrete storage implementations.
The BaseCache Abstraction
At the foundation lies BaseCache, defined in [litellm/caching/base_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py). This abstract base class mandates four core methods for all cache backends:
get(key: str) -> Any: Retrieves cached responsesset(key: str, value: Any, ttl: Optional[float] = None) -> None: Stores responses with optional TTLdelete(key: str) -> None: Removes specific entriesclear() -> None: Flushes the entire cache
Available Cache Backends
LiteLLM ships with several concrete implementations, each inheriting from BaseCache:
- InMemoryCache: LRU-based storage in [
litellm/caching/in_memory_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py), ideal for single-process applications - RedisCache: Persistent distributed storage in [
litellm/caching/redis_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py), sharing cache across multiple proxy instances - DiskCache: SQLite-based persistence in [
litellm/caching/disk_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/disk_cache.py), suitable for low-traffic environments without Redis infrastructure - DualCache: Tiered storage in [
litellm/caching/dual_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py), combining fast in-memory lookups with persistent backup
Request Flow and Coordination
When you configure LiteLLM caching, requests flow through several coordinated components:
- Cache Control Hook: The
cache_control_checkhook in [litellm/proxy/hooks/cache_control_check.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py) inspects payloads forcache_controlflags (ephemeralvspermanent) - Cache Coordinator: The
CacheCoordinatorin [litellm/proxy/common_utils/cache_coordinator.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py) builds deterministic cache keys from model name, prompt, temperature, and other parameters - Router Integration: The router checks the cache via [
litellm/router_utils/prompt_caching_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py) before calling providers, storing results after successful completions
Configuring In-Memory Caching
For single-process applications or development environments, in-memory caching provides the lowest latency.
import litellm
# Configure global in-memory cache with 1-hour TTL
litellm.set_cache(
cache=litellm.InMemoryCache(maxsize=10_000, ttl=3600)
)
response = litellm.completion(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Explain caching in LiteLLM"}],
)
The InMemoryCache class in [litellm/caching/in_memory_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py) implements an LRU (Least Recently Used) eviction policy. Data persists only for the process lifetime, making this unsuitable for distributed deployments but perfect for reducing repeated identical prompts within a single session.
Configuring Redis Caching
For production deployments requiring persistence across multiple LiteLLM Proxy instances, configure RedisCache from [litellm/caching/redis_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py):
import litellm
from litellm.caching.redis_cache import RedisCache
redis_cache = RedisCache(
redis_url="redis://:password@localhost:6379/0",
ttl=86400, # 24 hours
namespace="litellm"
)
litellm.set_cache(cache=redis_cache)
response = litellm.completion(
model="gpt-4",
messages=[{"role": "user", "content": "What is a dual cache?"}]
)
RedisCache supports connection pooling, namespace isolation, and configurable TTL per entry. This backend ensures that cache hits serve identical requests across server restarts and multiple proxy replicas.
Configuring Disk-Based Caching
When Redis infrastructure is unavailable, DiskCache provides SQLite-backed persistence through [litellm/caching/disk_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/disk_cache.py):
from litellm.caching.disk_cache import DiskCache
disk_cache = DiskCache(
cache_dir="/tmp/litellm_cache",
ttl=7200 # 2 hours
)
litellm.set_cache(cache=disk_cache)
This backend stores responses on local disk, surviving process restarts while requiring no external dependencies beyond SQLite.
Dual Cache Configuration
For optimal performance, DualCache in [litellm/caching/dual_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py) combines a fast primary cache (in-memory) with a persistent secondary cache (Redis or disk):
from litellm.caching.dual_cache import DualCache
from litellm.caching.in_memory_cache import InMemoryCache
from litellm.caching.redis_cache import RedisCache
dual = DualCache(
primary_cache=InMemoryCache(ttl=300), # 5 minutes
secondary_cache=RedisCache(redis_url="redis://localhost:6379/0", ttl=86400)
)
litellm.set_cache(cache=dual)
Reads check the in-memory cache first, then fall back to Redis. Writes populate both layers, ensuring hot data remains in RAM while cold data persists to external storage.
Proxy Configuration via YAML
When running the LiteLLM Proxy server, configure caching declaratively in your configuration file:
model_list:
- model_name: "gpt-3.5-turbo"
litellm_params:
model: "gpt-3.5-turbo"
cache_responses: true # Enable caching for this model
cache_control: "ephemeral" # Use in-memory only, skip persistent storage
cache_retention: 3600 # TTL in seconds (1 hour)
Start the proxy with:
litellm --config cache_config.yaml --port 4000
The proxy applies these settings through the CacheCoordinator in [litellm/proxy/common_utils/cache_coordinator.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py), respecting per-model TTL and cache control policies.
Per-Request Cache Overrides
Override global settings for individual completion calls using the cache_* parameter family:
litellm.completion(
model="gpt-4",
messages=[{"role": "user", "content": "Explain per-request TTL"}],
cache_responses=True, # Enable caching for this call only
cache_retention=1800, # 30-minute TTL for this entry
cache_control={"type": "ephemeral"} # Force in-memory storage
)
These parameters flow through the router to [litellm/router_utils/prompt_caching_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py), allowing dynamic cache policies based on request context.
Cache Control and TTL Management
LiteLLM respects two cache control modes via the hook in [litellm/proxy/hooks/cache_control_check.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py):
- Ephemeral: Stores data only in memory, bypassing Redis/disk persistence
- Permanent: Writes to all configured persistent backends
TTL (time-to-live) defaults to model-specific cache_retention values but can be overridden programmatically or via YAML. When TTL expires, the CacheCoordinator automatically invalidates entries on the next lookup.
Monitoring Cache Performance
The LiteLLM Proxy exposes cache metrics through the /cache/status endpoint defined in [litellm/proxy/management_endpoints/cache_settings_endpoints.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/management_endpoints/cache_settings_endpoints.py):
curl http://localhost:4000/cache/status
This returns JSON containing hit/miss ratios, total cached entries, and per-model cost breakdowns using the pricing schema from [litellm/types/management_endpoints/cache_settings_endpoints.py](https://github.com/BerriAI/litellm/blob/main/litellm/types/management_endpoints/cache_settings_endpoints.py).
Summary
- BaseCache abstraction in
litellm/caching/base_cache.pyprovides the unified interface for all storage backends - Four primary backends support diverse deployment needs: InMemoryCache for speed, RedisCache for distributed persistence, DiskCache for local file storage, and DualCache for tiered optimization
- Two configuration methods exist: programmatic
litellm.set_cache()for Python applications and YAML-based proxy configuration for server deployments - Cache control modes (
ephemeralvspermanent) and per-request TTL overrides enable fine-grained cache policies - Management endpoints at
/cache/statusprovide real-time visibility into hit rates and cost savings
Frequently Asked Questions
What is the difference between ephemeral and permanent cache control in LiteLLM?
Ephemeral caching stores responses only in memory and bypasses persistent backends like Redis or disk, making it ideal for sensitive data that should not survive process restarts. Permanent caching writes to all configured persistent storage layers, ensuring cache hits across proxy restarts and distributed instances. The cache_control_check hook in litellm/proxy/hooks/cache_control_check.py automatically routes requests based on these flags.
How does LiteLLM generate cache keys?
The CacheCoordinator in litellm/proxy/common_utils/cache_coordinator.py generates deterministic cache keys by hashing the model name, message content, temperature, max_tokens, and other request parameters. This ensures that identical prompts with identical parameters receive cached responses while slight variations trigger fresh API calls.
Can I use multiple cache backends simultaneously?
Yes, through the DualCache implementation in litellm/caching/dual_cache.py. Configure a fast primary cache (typically InMemoryCache) for hot data and a secondary persistent cache (RedisCache or DiskCache) for long-term storage. The system checks the primary cache first, falls back to the secondary on miss, and writes results to both layers.
How does LiteLLM handle pricing for cached responses?
According to the schema in litellm/types/management_endpoints/cache_settings_endpoints.py, LiteLLM tracks cache_read_input_token_cost and cache_creation_input_token_cost separately from standard input costs. The proxy calculates savings based on these rates when serving cached responses, allowing accurate cost attribution in the /cache/status metrics endpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →