# How to Configure LiteLLM Caching: Redis, In-Memory, and Disk Backends Explained

> Learn to configure LiteLLM caching with Redis, in-memory, and disk backends. Reduce LLM costs and latency using TTL and cache policies.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: how-to-guide
- Published: 2026-03-26

---

**LiteLLM offers a pluggable caching architecture built on the `BaseCache` abstraction that supports in-memory, Redis, disk, and dual-cache backends to reduce LLM token costs and latency through configurable TTL and cache control policies.**

LiteLLM enables intelligent request/response caching to minimize API costs and improve response times across multiple LLM providers. Whether you are running the LiteLLM Proxy or integrating caching directly into Python applications, understanding how to configure LiteLLM caching is essential for production deployments. The BerriAI/litellm repository implements this through a modular system centered on the `BaseCache` interface in [`litellm/caching/base_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py).

## Core Caching Architecture

The LiteLLM caching system relies on a clean separation between the abstraction layer and concrete storage implementations.

### The BaseCache Abstraction

At the foundation lies `BaseCache`, defined in [[`litellm/caching/base_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py). This abstract base class mandates four core methods for all cache backends:

- `get(key: str) -> Any`: Retrieves cached responses
- `set(key: str, value: Any, ttl: Optional[float] = None) -> None`: Stores responses with optional TTL
- `delete(key: str) -> None`: Removes specific entries
- `clear() -> None`: Flushes the entire cache

### Available Cache Backends

LiteLLM ships with several concrete implementations, each inheriting from `BaseCache`:

- **InMemoryCache**: LRU-based storage in [[`litellm/caching/in_memory_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py), ideal for single-process applications
- **RedisCache**: Persistent distributed storage in [[`litellm/caching/redis_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py), sharing cache across multiple proxy instances
- **DiskCache**: SQLite-based persistence in [[`litellm/caching/disk_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/disk_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/disk_cache.py), suitable for low-traffic environments without Redis infrastructure
- **DualCache**: Tiered storage in [[`litellm/caching/dual_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py), combining fast in-memory lookups with persistent backup

### Request Flow and Coordination

When you configure LiteLLM caching, requests flow through several coordinated components:

1. **Cache Control Hook**: The `cache_control_check` hook in [[`litellm/proxy/hooks/cache_control_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py) inspects payloads for `cache_control` flags (`ephemeral` vs `permanent`)
2. **Cache Coordinator**: The `CacheCoordinator` in [[`litellm/proxy/common_utils/cache_coordinator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py) builds deterministic cache keys from model name, prompt, temperature, and other parameters
3. **Router Integration**: The router checks the cache via [[`litellm/router_utils/prompt_caching_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py) before calling providers, storing results after successful completions

## Configuring In-Memory Caching

For single-process applications or development environments, in-memory caching provides the lowest latency.

```python
import litellm

# Configure global in-memory cache with 1-hour TTL

litellm.set_cache(
    cache=litellm.InMemoryCache(maxsize=10_000, ttl=3600)
)

response = litellm.completion(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "Explain caching in LiteLLM"}],
)

```

The `InMemoryCache` class in [[`litellm/caching/in_memory_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py) implements an LRU (Least Recently Used) eviction policy. Data persists only for the process lifetime, making this unsuitable for distributed deployments but perfect for reducing repeated identical prompts within a single session.

## Configuring Redis Caching

For production deployments requiring persistence across multiple LiteLLM Proxy instances, configure **RedisCache** from [[`litellm/caching/redis_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py):

```python
import litellm
from litellm.caching.redis_cache import RedisCache

redis_cache = RedisCache(
    redis_url="redis://:password@localhost:6379/0",
    ttl=86400,  # 24 hours

    namespace="litellm"
)

litellm.set_cache(cache=redis_cache)

response = litellm.completion(
    model="gpt-4",
    messages=[{"role": "user", "content": "What is a dual cache?"}]
)

```

**RedisCache** supports connection pooling, namespace isolation, and configurable TTL per entry. This backend ensures that cache hits serve identical requests across server restarts and multiple proxy replicas.

## Configuring Disk-Based Caching

When Redis infrastructure is unavailable, **DiskCache** provides SQLite-backed persistence through [[`litellm/caching/disk_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/disk_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/disk_cache.py):

```python
from litellm.caching.disk_cache import DiskCache

disk_cache = DiskCache(
    cache_dir="/tmp/litellm_cache",
    ttl=7200  # 2 hours

)

litellm.set_cache(cache=disk_cache)

```

This backend stores responses on local disk, surviving process restarts while requiring no external dependencies beyond SQLite.

## Dual Cache Configuration

For optimal performance, **DualCache** in [[`litellm/caching/dual_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py) combines a fast primary cache (in-memory) with a persistent secondary cache (Redis or disk):

```python
from litellm.caching.dual_cache import DualCache
from litellm.caching.in_memory_cache import InMemoryCache
from litellm.caching.redis_cache import RedisCache

dual = DualCache(
    primary_cache=InMemoryCache(ttl=300),  # 5 minutes

    secondary_cache=RedisCache(redis_url="redis://localhost:6379/0", ttl=86400)
)

litellm.set_cache(cache=dual)

```

Reads check the in-memory cache first, then fall back to Redis. Writes populate both layers, ensuring hot data remains in RAM while cold data persists to external storage.

## Proxy Configuration via YAML

When running the LiteLLM Proxy server, configure caching declaratively in your configuration file:

```yaml
model_list:
  - model_name: "gpt-3.5-turbo"
    litellm_params:
      model: "gpt-3.5-turbo"
      cache_responses: true          # Enable caching for this model

      cache_control: "ephemeral"    # Use in-memory only, skip persistent storage

      cache_retention: 3600         # TTL in seconds (1 hour)

```

Start the proxy with:

```bash
litellm --config cache_config.yaml --port 4000

```

The proxy applies these settings through the `CacheCoordinator` in [[`litellm/proxy/common_utils/cache_coordinator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py), respecting per-model TTL and cache control policies.

## Per-Request Cache Overrides

Override global settings for individual completion calls using the `cache_*` parameter family:

```python
litellm.completion(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain per-request TTL"}],
    cache_responses=True,          # Enable caching for this call only

    cache_retention=1800,          # 30-minute TTL for this entry

    cache_control={"type": "ephemeral"}  # Force in-memory storage

)

```

These parameters flow through the router to [[`litellm/router_utils/prompt_caching_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py)](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/prompt_caching_cache.py), allowing dynamic cache policies based on request context.

## Cache Control and TTL Management

LiteLLM respects two cache control modes via the hook in [[`litellm/proxy/hooks/cache_control_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py):

- **Ephemeral**: Stores data only in memory, bypassing Redis/disk persistence
- **Permanent**: Writes to all configured persistent backends

TTL (time-to-live) defaults to model-specific `cache_retention` values but can be overridden programmatically or via YAML. When TTL expires, the `CacheCoordinator` automatically invalidates entries on the next lookup.

## Monitoring Cache Performance

The LiteLLM Proxy exposes cache metrics through the `/cache/status` endpoint defined in [[`litellm/proxy/management_endpoints/cache_settings_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/management_endpoints/cache_settings_endpoints.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/management_endpoints/cache_settings_endpoints.py):

```bash
curl http://localhost:4000/cache/status

```

This returns JSON containing hit/miss ratios, total cached entries, and per-model cost breakdowns using the pricing schema from [[`litellm/types/management_endpoints/cache_settings_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/management_endpoints/cache_settings_endpoints.py)](https://github.com/BerriAI/litellm/blob/main/litellm/types/management_endpoints/cache_settings_endpoints.py).

## Summary

- **BaseCache abstraction** in [`litellm/caching/base_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py) provides the unified interface for all storage backends
- **Four primary backends** support diverse deployment needs: InMemoryCache for speed, RedisCache for distributed persistence, DiskCache for local file storage, and DualCache for tiered optimization
- **Two configuration methods** exist: programmatic `litellm.set_cache()` for Python applications and YAML-based proxy configuration for server deployments
- **Cache control modes** (`ephemeral` vs `permanent`) and per-request TTL overrides enable fine-grained cache policies
- **Management endpoints** at `/cache/status` provide real-time visibility into hit rates and cost savings

## Frequently Asked Questions

### What is the difference between ephemeral and permanent cache control in LiteLLM?

**Ephemeral caching** stores responses only in memory and bypasses persistent backends like Redis or disk, making it ideal for sensitive data that should not survive process restarts. **Permanent caching** writes to all configured persistent storage layers, ensuring cache hits across proxy restarts and distributed instances. The `cache_control_check` hook in [`litellm/proxy/hooks/cache_control_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/cache_control_check.py) automatically routes requests based on these flags.

### How does LiteLLM generate cache keys?

The `CacheCoordinator` in [`litellm/proxy/common_utils/cache_coordinator.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/common_utils/cache_coordinator.py) generates deterministic cache keys by hashing the model name, message content, temperature, max_tokens, and other request parameters. This ensures that identical prompts with identical parameters receive cached responses while slight variations trigger fresh API calls.

### Can I use multiple cache backends simultaneously?

Yes, through the **DualCache** implementation in [`litellm/caching/dual_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/dual_cache.py). Configure a fast primary cache (typically InMemoryCache) for hot data and a secondary persistent cache (RedisCache or DiskCache) for long-term storage. The system checks the primary cache first, falls back to the secondary on miss, and writes results to both layers.

### How does LiteLLM handle pricing for cached responses?

According to the schema in [`litellm/types/management_endpoints/cache_settings_endpoints.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/management_endpoints/cache_settings_endpoints.py), LiteLLM tracks `cache_read_input_token_cost` and `cache_creation_input_token_cost` separately from standard input costs. The proxy calculates savings based on these rates when serving cached responses, allowing accurate cost attribution in the `/cache/status` metrics endpoint.