# LiteLLM Caching Layer Architecture: Redis vs In-Memory Implementation Guide

> Explore LiteLLM's caching layer architecture. Compare Redis vs in-memory implementations and understand how to integrate them seamlessly using the BaseCache interface for optimal performance.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: architecture
- Published: 2026-03-26

---

**LiteLLM implements a dual-backend caching architecture that caches both HTTP client objects and model responses using either an in-process dictionary or a Redis server, both conforming to the abstract `BaseCache` interface for seamless interchangeability.**

LiteLLM provides a unified Python SDK for interacting with over 100 large language model providers. To optimize performance and reduce redundant API latency, the library incorporates a sophisticated caching layer that persists both client connections and completion responses. This technical deep dive examines the `BerriAI/litellm` repository's caching implementation, contrasting the zero-configuration in-memory backend with the production-grade Redis alternative.

## How LiteLLM Caching Works

The caching layer in LiteLLM is built around an abstract `BaseCache` class defined in [`litellm/caching/base_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py). This interface standardizes three core operations: `get_cache()`, `set_cache()`, and `flush_cache()`. Both the in-memory and Redis implementations inherit from this base, enabling runtime swapping without code changes. When a completion request is initiated, LiteLLM generates a deterministic cache key from the model name, provider endpoint, and request payload, then queries the configured backend—either a local Python dictionary or a remote Redis store—to retrieve cached clients or previous responses. An auxiliary pattern using `InMemoryFile` for batch operations is also demonstrated in [`litellm/router_utils/batch_utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/batch_utils.py), illustrating the library's consistent approach to process-local storage.

## In-Memory Cache Implementation

### Architecture and Scope

The in-memory backend, implemented in [`litellm/caching/in_memory_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py), provides a process-local **LRU (Least Recently Used)** cache using a standard Python `dict`. This implementation is exposed globally through the `LLMClientCache` singleton instantiated in [`litellm/__init__.py`](https://github.com/BerriAI/litellm/blob/main/litellm/__init__.py) as `in_memory_llm_clients_cache`. The cache lives only as long as the Python interpreter runs, making it suitable for single-process scripts, unit tests, and development environments where persistence across restarts is not required.

### Usage Example

No external services or environment variables are required to use the in-memory cache. The singleton is created automatically upon importing the relevant modules:

```python
import litellm

# Access the singleton cache instance

cache = litellm.in_memory_llm_clients_cache

# Check for existing OpenAI client using the cache key format

client = cache.get_cache("openai:gpt-4o")
if client is None:
    client = litellm.OpenAI()
    cache.set_cache("openai:gpt-4o", client)

# Subsequent calls reuse the cached HTTP connection

response = litellm.completion(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello"}]
)

```

Entries are evicted based on LRU policy or cleared manually using `cache.flush_cache()`, which is useful for testing scenarios requiring clean state.

## Redis Cache Implementation

### Architecture and Serialization

For production deployments requiring cache sharing across multiple workers or containers, [`litellm/caching/redis_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py) implements the `RedisCache` class. This backend serializes values using Python's `pickle` module before storing them in Redis, enabling complex objects like HTTP clients to be persisted across process boundaries. Unlike the in-memory variant, this implementation supports configurable **Time-To-Live (TTL)** eviction, with a default expiration of 1 hour per entry.

### Configuration and Environment Variables

Switching to Redis requires setting specific environment variables before initializing the cache:

```python
import os
import litellm

# Configure Redis backend

os.environ["LLM_CACHE"] = "redis"
os.environ["REDIS_HOST"] = "redis.example.com"
os.environ["REDIS_PORT"] = "6379"
os.environ["REDIS_PASSWORD"] = "secure-password"

# Lazy instantiation on first access

cache = litellm.redis_cache

# Store with custom TTL (30 minutes)

cache.set_cache(
    key="openai:gpt-4o",
    value=litellm.OpenAI(),
    ttl_secs=1800
)

```

The `RedisCache` also provides fallback mechanisms to the in-memory cache for certain client objects, ensuring resilience if Redis connectivity is temporarily unavailable.

## Cache Key Generation and Lookup Flow

The caching mechanism operates through a deterministic four-step process visible in the integration points such as [`litellm/llms/openai/common_utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/openai/common_utils.py):

1. **Key Generation** – LiteLLM constructs a unique cache key combining the provider endpoint, model identifier, and hashed request parameters.
2. **Lookup** – The system queries either `InMemoryCache.get_cache()` or `RedisCache.get_cache()` depending on the `LLM_CACHE` environment variable.
3. **Hit Handling** – If a valid entry exists (within TTL for Redis), the deserialized client or cached response returns immediately, bypassing external API calls.
4. **Miss Handling** – On cache miss, LiteLLM instantiates a new client or fetches from the provider, then stores the result via `set_cache()` for future requests.

This architecture ensures that both client connection overhead and expensive LLM inference costs are minimized across repeated identical calls.

## Performance Characteristics and Trade-offs

**In-Memory Cache:**
- **Zero network latency**: Direct dictionary lookups provide microsecond-level access times.
- **No operational overhead**: Requires no external services or configuration management.
- **Process isolation**: Each worker maintains independent cache state, potentially duplicating memory usage and API warmup costs across multiple instances.
- **Volatility**: All cached data is lost on process restart or deployment, as entries are stored only in the Python process heap.

**Redis Cache:**
- **Distributed consistency**: All workers, containers, and serverless functions share identical cache state, eliminating redundant warmups and ensuring cache hits across the entire infrastructure.
- **Persistence**: Survives process restarts and supports serverless runtimes that frequently recycle execution environments.
- **Operational complexity**: Requires running and maintaining a Redis server (or managed service like AWS ElastiCache), adding infrastructure overhead.
- **Serialization overhead**: `pickle` serialization/deserialization and network round-trips add marginal latency compared to local dictionary access, though this is typically negligible relative to LLM API latency.

## Summary

- LiteLLM implements a **pluggable caching architecture** via the `BaseCache` abstraction in [`litellm/caching/base_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/base_cache.py), supporting seamless backend swaps through environment variables.
- The **in-memory cache** ([`litellm/caching/in_memory_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/in_memory_cache.py)) provides zero-configuration, process-local storage ideal for development and single-instance deployments via the `LLMClientCache` singleton.
- The **Redis cache** ([`litellm/caching/redis_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/caching/redis_cache.py)) enables distributed, persistent caching across microservices using `pickle` serialization and configurable TTL via the `cache_ttl_secs` parameter.
- Cache keys are deterministically generated from request parameters, ensuring identical calls receive cached responses regardless of backend selection.
- Configuration is environment-driven via `LLM_CACHE`, `REDIS_HOST`, `REDIS_PORT`, and `REDIS_PASSWORD` variables, with lazy instantiation on first use.

## Frequently Asked Questions

### How do I switch from in-memory caching to Redis in LiteLLM?

Set the `LLM_CACHE` environment variable to `"redis"` and provide connection details via `REDIS_HOST`, `REDIS_PORT`, and `REDIS_PASSWORD`. The `RedisCache` instance is lazily created on first access, allowing runtime backend switching without modifying application code or cache lookup logic.

### What is the default TTL for cached entries in LiteLLM's Redis implementation?

The default TTL is **1 hour (3600 seconds)**. You can override this per-entry using the `ttl_secs` parameter when calling `set_cache()`, or modify the default behavior in the `RedisCache` class implementation.

### Does LiteLLM cache both LLM responses and HTTP client connections?

Yes. The caching layer handles both **client objects** (such as OpenAI HTTP clients cached to avoid connection establishment overhead) and **model responses** (to avoid redundant API calls for identical prompts). The specific caching strategy varies by provider integration, with client caching logic visible in files like [`litellm/llms/openai/common_utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/openai/common_utils.py).

### How can I manually clear the cache during testing?

Both backends implement a `flush_cache()` method. For the in-memory cache, call `litellm.in_memory_llm_clients_cache.flush_cache()`. For Redis, use `litellm.redis_cache.flush_cache()`. This is particularly useful in CI/CD pipelines or unit tests requiring isolated state between test cases.