How to Configure LiteLLM Caching: Redis, In-Memory, and Disk Backends Explained

LiteLLM offers a pluggable caching architecture built on the BaseCache abstraction that supports in-memory, Redis, disk, and dual-cache backends to reduce LLM token costs and latency through configurable TTL and cache control policies.

LiteLLM enables intelligent request/response caching to minimize API costs and improve response times across multiple LLM providers. Whether you are running the LiteLLM Proxy or integrating caching directly into Python applications, understanding how to configure LiteLLM caching is essential for production deployments. The BerriAI/litellm repository implements this through a modular system centered on the BaseCache interface in litellm/caching/base_cache.py.

Core Caching Architecture

The LiteLLM caching system relies on a clean separation between the abstraction layer and concrete storage implementations.

The BaseCache Abstraction

At the foundation lies BaseCache, defined in [litellm/caching/base_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py). This abstract base class mandates four core methods for all cache backends:

  • get(key: str) -> Any: Retrieves cached responses
  • set(key: str, value: Any, ttl: Optional[float] = None) -> None: Stores responses with optional TTL
  • delete(key: str) -> None: Removes specific entries
  • clear() -> None: Flushes the entire cache

Available Cache Backends

LiteLLM ships with several concrete implementations, each inheriting from BaseCache:

Request Flow and Coordination

When you configure LiteLLM caching, requests flow through several coordinated components:

  1. Cache Control Hook: The cache_control_check hook in [litellm/proxy/hooks/cache_control_check.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py) inspects payloads for cache_control flags (ephemeral vs permanent)
  2. Cache Coordinator: The CacheCoordinator in [litellm/proxy/common_utils/cache_coordinator.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py) builds deterministic cache keys from model name, prompt, temperature, and other parameters
  3. Router Integration: The router checks the cache via [litellm/router_utils/prompt_caching_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py) before calling providers, storing results after successful completions

Configuring In-Memory Caching

For single-process applications or development environments, in-memory caching provides the lowest latency.

import litellm

# Configure global in-memory cache with 1-hour TTL

litellm.set_cache(
    cache=litellm.InMemoryCache(maxsize=10_000, ttl=3600)
)

response = litellm.completion(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "Explain caching in LiteLLM"}],
)

The InMemoryCache class in [litellm/caching/in_memory_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py) implements an LRU (Least Recently Used) eviction policy. Data persists only for the process lifetime, making this unsuitable for distributed deployments but perfect for reducing repeated identical prompts within a single session.

Configuring Redis Caching

For production deployments requiring persistence across multiple LiteLLM Proxy instances, configure RedisCache from [litellm/caching/redis_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py):

import litellm
from litellm.caching.redis_cache import RedisCache

redis_cache = RedisCache(
    redis_url="redis://:password@localhost:6379/0",
    ttl=86400,  # 24 hours

    namespace="litellm"
)

litellm.set_cache(cache=redis_cache)

response = litellm.completion(
    model="gpt-4",
    messages=[{"role": "user", "content": "What is a dual cache?"}]
)

RedisCache supports connection pooling, namespace isolation, and configurable TTL per entry. This backend ensures that cache hits serve identical requests across server restarts and multiple proxy replicas.

Configuring Disk-Based Caching

When Redis infrastructure is unavailable, DiskCache provides SQLite-backed persistence through [litellm/caching/disk_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/disk_cache.py):

from litellm.caching.disk_cache import DiskCache

disk_cache = DiskCache(
    cache_dir="/tmp/litellm_cache",
    ttl=7200  # 2 hours

)

litellm.set_cache(cache=disk_cache)

This backend stores responses on local disk, surviving process restarts while requiring no external dependencies beyond SQLite.

Dual Cache Configuration

For optimal performance, DualCache in [litellm/caching/dual_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py) combines a fast primary cache (in-memory) with a persistent secondary cache (Redis or disk):

from litellm.caching.dual_cache import DualCache
from litellm.caching.in_memory_cache import InMemoryCache
from litellm.caching.redis_cache import RedisCache

dual = DualCache(
    primary_cache=InMemoryCache(ttl=300),  # 5 minutes

    secondary_cache=RedisCache(redis_url="redis://localhost:6379/0", ttl=86400)
)

litellm.set_cache(cache=dual)

Reads check the in-memory cache first, then fall back to Redis. Writes populate both layers, ensuring hot data remains in RAM while cold data persists to external storage.

Proxy Configuration via YAML

When running the LiteLLM Proxy server, configure caching declaratively in your configuration file:

model_list:
  - model_name: "gpt-3.5-turbo"
    litellm_params:
      model: "gpt-3.5-turbo"
      cache_responses: true          # Enable caching for this model

      cache_control: "ephemeral"    # Use in-memory only, skip persistent storage

      cache_retention: 3600         # TTL in seconds (1 hour)

Start the proxy with:

litellm --config cache_config.yaml --port 4000

The proxy applies these settings through the CacheCoordinator in [litellm/proxy/common_utils/cache_coordinator.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py), respecting per-model TTL and cache control policies.

Per-Request Cache Overrides

Override global settings for individual completion calls using the cache_* parameter family:

litellm.completion(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain per-request TTL"}],
    cache_responses=True,          # Enable caching for this call only

    cache_retention=1800,          # 30-minute TTL for this entry

    cache_control={"type": "ephemeral"}  # Force in-memory storage

)

These parameters flow through the router to [litellm/router_utils/prompt_caching_cache.py](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py), allowing dynamic cache policies based on request context.

Cache Control and TTL Management

LiteLLM respects two cache control modes via the hook in [litellm/proxy/hooks/cache_control_check.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py):

  • Ephemeral: Stores data only in memory, bypassing Redis/disk persistence
  • Permanent: Writes to all configured persistent backends

TTL (time-to-live) defaults to model-specific cache_retention values but can be overridden programmatically or via YAML. When TTL expires, the CacheCoordinator automatically invalidates entries on the next lookup.

Monitoring Cache Performance

The LiteLLM Proxy exposes cache metrics through the /cache/status endpoint defined in [litellm/proxy/management_endpoints/cache_settings_endpoints.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/management_endpoints/cache_settings_endpoints.py):

curl http://localhost:4000/cache/status

This returns JSON containing hit/miss ratios, total cached entries, and per-model cost breakdowns using the pricing schema from [litellm/types/management_endpoints/cache_settings_endpoints.py](https://github.com/BerriAI/litellm/blob/main/litellm/types/management_endpoints/cache_settings_endpoints.py).

Summary

  • BaseCache abstraction in litellm/caching/base_cache.py provides the unified interface for all storage backends
  • Four primary backends support diverse deployment needs: InMemoryCache for speed, RedisCache for distributed persistence, DiskCache for local file storage, and DualCache for tiered optimization
  • Two configuration methods exist: programmatic litellm.set_cache() for Python applications and YAML-based proxy configuration for server deployments
  • Cache control modes (ephemeral vs permanent) and per-request TTL overrides enable fine-grained cache policies
  • Management endpoints at /cache/status provide real-time visibility into hit rates and cost savings

Frequently Asked Questions

What is the difference between ephemeral and permanent cache control in LiteLLM?

Ephemeral caching stores responses only in memory and bypasses persistent backends like Redis or disk, making it ideal for sensitive data that should not survive process restarts. Permanent caching writes to all configured persistent storage layers, ensuring cache hits across proxy restarts and distributed instances. The cache_control_check hook in litellm/proxy/hooks/cache_control_check.py automatically routes requests based on these flags.

How does LiteLLM generate cache keys?

The CacheCoordinator in litellm/proxy/common_utils/cache_coordinator.py generates deterministic cache keys by hashing the model name, message content, temperature, max_tokens, and other request parameters. This ensures that identical prompts with identical parameters receive cached responses while slight variations trigger fresh API calls.

Can I use multiple cache backends simultaneously?

Yes, through the DualCache implementation in litellm/caching/dual_cache.py. Configure a fast primary cache (typically InMemoryCache) for hot data and a secondary persistent cache (RedisCache or DiskCache) for long-term storage. The system checks the primary cache first, falls back to the secondary on miss, and writes results to both layers.

How does LiteLLM handle pricing for cached responses?

According to the schema in litellm/types/management_endpoints/cache_settings_endpoints.py, LiteLLM tracks cache_read_input_token_cost and cache_creation_input_token_cost separately from standard input costs. The proxy calculates savings based on these rates when serving cached responses, allowing accurate cost attribution in the /cache/status metrics endpoint.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →