How Bella OpenAPI Implements Multi-Level Caching with Caffeine, Redis, and JetCache

Bella OpenAPI uses JetCache as a unified façade to combine Caffeine (local in-memory) and Redis via Redisson (remote distributed) into a two-tier cache that serves hot data from the JVM while maintaining consistency across cluster nodes.

The multi-level caching architecture in Bella OpenAPI addresses the latency and consistency challenges of high-throughput API services. By leveraging JetCache's abstraction layer, the system automatically coordinates between ultra-fast local caching provided by Caffeine and distributed synchronization powered by Redis, ensuring optimal performance without sacrificing data coherence across multiple server instances.

Understanding the Multi-Level Caching Architecture

Bella OpenAPI implements a two-level cache (CacheType.BOTH) for each logical cache region. This architecture separates concerns between speed and distribution:

  • Local Layer (Caffeine): Provides microsecond-level access within a single JVM instance, configured with size limits and short TTL to prevent memory bloat.
  • Remote Layer (Redis/Redisson): Offers millisecond-level access shared across all cluster nodes, ensuring consistency with longer TTL and key namespacing.

The CacheManager bean, auto-configured at startup, reads the YAML configuration and instantiates these layered caches, while individual services can further customize behavior through QuickConfig builders.

Configuration and Setup

Local Cache Configuration (Caffeine)

The local caching tier is defined in api/server/src/main/resources/application.yml under the jetcache.local section. Bella OpenAPI configures Caffeine with conservative limits to balance speed and memory usage:

jetcache:
  local:
    default:
      type: caffeine
      keyConvertor: fastjson2
      limit: 100
      expireAfterWriteInMillis: 120000  # 2 minutes

This configuration limits each local cache region to 100 entries with a 2-minute expiration, ensuring that stale data does not linger in individual JVMs while frequently accessed keys remain in memory.

Remote Cache Configuration (Redis/Redisson)

The remote tier uses Redisson as the Redis client, configured in the same YAML file under jetcache.remote:

jetcache:
  remote:
    default:
      type: redisson
      keyConvertor: fastjson2
      keyPrefix: bella-openapi-
      expireAfterWriteInMillis: 600000  # 10 minutes

      broadcast: false

The 10-minute TTL provides a longer-lived shared cache across the cluster, while the bella-openapi- prefix prevents key collisions with other services. The broadcast: false setting indicates that invalidation messages are not broadcast to other JVMs, relying instead on TTL expiration for consistency.

How the Two-Level Cache Works

Lookup Flow and Population

When a service method annotated with @Cached or a programmatic cache access is invoked, JetCache executes the following sequence:

  1. Local Lookup: Query the Caffeine cache within the current JVM. If present, return immediately (microsecond latency).
  2. Remote Fallback: On local miss, query Redis via Redisson. If present, return the value and backfill the local Caffeine cache so subsequent requests hit the local tier.
  3. Source Query: If both caches miss, execute the underlying method, store the result in both tiers, and return.

This read-through behavior ensures that hot keys naturally migrate to the local cache while the remote layer maintains a consistent view across the cluster.

Cache Penetration Protection

Bella OpenAPI enables cache penetration protection (stampede prevention) on all caches to prevent thundering herd problems when popular keys expire. In ModelService.java, the QuickConfig builder explicitly sets:

.penetrationProtect(true)
.penetrationProtectTimeout(Duration.ofSeconds(10))

This ensures that when a cache entry expires, only one thread is allowed to load the value from the database while others wait up to 10 seconds, preventing database overload during high-traffic events.

Implementation Patterns in Bella OpenAPI

Annotation-Driven Caching

The most common pattern in Bella OpenAPI uses JetCache's @Cached annotation to declaratively apply two-level caching. In EndpointService.java, endpoint details are cached with:

@Cached(
    name = "endpoint:details:",
    key = "#condition.endpoint + ':' + #identity",
    cacheType = CacheType.BOTH,
    condition = "(#condition.modelName == null || #condition.modelName == '')"
)
public EndpointDetails getEndpointDetails(Condition condition, String identity) {
    // business logic
}

The cacheType = CacheType.BOTH parameter instructs JetCache to use both Caffeine and Redis layers, while the condition ensures caching only occurs when specific criteria are met.

Programmatic Cache Creation with QuickConfig

For caches requiring custom parameters beyond the defaults, Bella OpenAPI uses QuickConfig in ModelService.java:

@PostConstruct
public void init() {
    // Model terminal cache - local only, high capacity
    QuickConfig terminalConfig = QuickConfig.newBuilder("model:terminal:")
            .cacheType(CacheType.LOCAL)
            .limit(500)
            .expire(Duration.ofMinutes(10))
            .penetrationProtect(true)
            .penetrationProtectTimeout(Duration.ofSeconds(10))
            .build();
    cacheManager.getOrCreateCache(terminalConfig);
    
    // Model metadata cache - two-level with extreme remote TTL
    QuickConfig modelCacheConfig = QuickConfig.newBuilder(modelMapCacheKey)
            .expire(Duration.ofDays(3650))      // 10 years in Redis
            .localExpire(Duration.ofMinutes(5)) // 5 min local freshness
            .localLimit(1)
            .cacheType(CacheType.BOTH)
            .syncLocal(true)
            .penetrationProtect(true)
            .penetrationProtectTimeout(Duration.ofSeconds(10))
            .build();
    cacheManager.getOrCreateCache(modelCacheConfig);
}

This pattern allows fine-grained control over cache regions, such as setting a 10-year TTL for model metadata in Redis while keeping only the most recent entry locally for 5 minutes.

Standalone Caffeine for Specialized Use Cases

In the SDK module, Bella OpenAPI uses a standalone Caffeine instance for caching compiled Groovy scripts, bypassing JetCache entirely. In GroovyExecutor.java:

private static final Cache<String, Script> scriptCache = Caffeine.newBuilder()
        .maximumSize(100)
        .expireAfterAccess(30, TimeUnit.MINUTES)
        .build();

private static Script getCompiledScript(String scriptText) {
    return scriptCache.get(scriptText, key -> {
        GroovyShell shell = new GroovyShell(classLoader);
        return shell.parse(key);
    });
}

This specialized cache avoids the overhead of serialization required for Redis, keeping compiled script objects directly in heap memory for rapid execution.

Summary

Bella OpenAPI's multi-level caching architecture delivers high-performance data access through these key mechanisms:

  • JetCache unifies Caffeine and Redis into a transparent two-level cache (CacheType.BOTH), automatically handling local-to-remote fallback and backfill.
  • Configuration-driven defaults in application.yml establish baseline limits (100 entries/2min local, 10min remote) while QuickConfig enables per-cache customization.
  • Penetration protection prevents cache stampedes by serializing load operations for expired keys across both annotation-driven and programmatic caches.
  • Hybrid usage patterns combine annotation-driven caching for service methods with standalone Caffeine instances for specialized objects like compiled Groovy scripts.

Frequently Asked Questions

What is the lookup order in Bella OpenAPI's multi-level cache?

The lookup follows a strict top-down sequence: first the Caffeine local cache within the JVM, then the Redis remote cache via Redisson if the local cache misses. When a value is found in Redis, it is automatically backfilled into the local Caffeine cache so subsequent requests are served locally without network overhead.

How does Bella OpenAPI prevent cache stampedes?

Bella OpenAPI enables cache penetration protection on all JetCache instances by setting penetrationProtect=true and a 10-second timeout (penetrationProtectTimeout). When a popular cache entry expires, only one thread is permitted to load the value from the database while other concurrent requests wait for the result, preventing thundering herd scenarios that could overwhelm backend services.

Why use both Caffeine and Redis instead of just one?

The combination leverages the strengths of each technology: Caffeine provides microsecond-level access for hot data within a single JVM without network serialization overhead, while Redis ensures consistency across distributed instances and survives JVM restarts. Using only Caffeine would cause cache fragmentation across the cluster; using only Redis would introduce unnecessary network latency for data that could be served locally.

Where is the cache configuration defined in the codebase?

The primary configuration resides in api/server/src/main/resources/application.yml under the jetcache section, defining default behaviors for both local (Caffeine) and remote (Redisson) tiers. Additionally, programmatic configurations using QuickConfig are implemented in service classes such as api/server/src/main/java/com/ke/bella/openapi/service/ModelService.java for cache regions requiring custom TTLs, size limits, or penetration protection settings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →