Moka Cache Performance Tuning: Optimizing Builder Configuration for High-Concurrency Rust Workloads

Moka cache performance tuning relies on configuring the CacheBuilder with segmented hash tables, fast hashers, size-based weighers, and aggressive house-keeper intervals to minimize thread contention and memory pressure in the moka-rs/moka library.

Moka is a high-performance, thread-safe in-memory cache for Rust optimized for concurrent workloads. Effective Moka cache performance tuning requires understanding the internal architecture—including the CacheBuilder in src/sync/builder.rs, the segmented lock-free hash table, and the timer-wheel house-keeper—to optimize for your specific latency and throughput requirements.

Core Architecture and Tuning Levers

The CacheBuilder and BaseCache Internal Structure

The CacheBuilder struct in src/sync/builder.rs accumulates all configuration knobs before instantiating the cache via the internal with_everything method (lines 30–50). This factory method creates a BaseCache (defined in src/sync/base_cache.rs) which holds the lock-free hash table, timer wheel for expirations, and house-keeper state. Every tuning decision flows through this builder pipeline, from capacity limits to hash algorithms.

Segmented Hash Tables for Concurrency

When segments(N) is invoked, the builder creates a SegmentedCache that distributes keys across N independent hash tables, each with its own lock-free bucket array. According to the source in src/sync/builder.rs (lines 100–126), this sharding dramatically reduces cross-thread contention under high parallelism by isolating writes to separate segments.

Timer Wheel and House-Keeper

Moka stores expiration timestamps in a hierarchical timer wheel implemented in src/common/timer_wheel.rs. The house-keeper (configured via HousekeeperConfig in src/common/concurrent/housekeeper.rs) runs periodically (defaulting to every 5 seconds) to advance the wheel and evict expired entries. The configuration struct exposes two critical fields:

  • run_interval: Duration between maintenance ticks
  • shards_per_run: Number of timer-wheel shards processed per tick

Lowering run_interval improves expiration latency but increases CPU overhead, while adjusting shards_per_run controls the amortized work per maintenance cycle.

Critical Configuration Levers

Capacity and Segmentation (max_capacity and segments)

Set max_capacity to a realistic bound to prevent unbounded memory growth. For workloads with heavy concurrent writes, increase segments via CacheBuilder::segments(16) or higher. The BaseCache stores each segment as an independent lock-free table, reducing cache-line bouncing across CPU cores.

Hash Algorithm Selection (build_with_hasher)

The default SipHash provides DoS resistance but adds latency for simple keys. For trusted internal workloads, replace it with a faster algorithm like ahash using build_with_hasher (lines 226–247 in src/sync/builder.rs):

use moka::sync::Cache;
use ahash::RandomState;

let cache = Cache::builder()
    .max_capacity(100_000)
    .build_with_hasher(RandomState::default());

This bypasses the cryptographic overhead of SipHash while maintaining excellent distribution for integer or string keys.

Size-Based Eviction with Weighers

When values vary significantly in size (e.g., large byte buffers), use CacheBuilder::weigher to define a closure that returns a u32 weight per entry. The cache tracks total weighted size via BaseCache::evict and triggers eviction when the sum exceeds max_capacity:

use moka::sync::Cache;
use std::sync::Arc;

let cache = Cache::builder()
    .max_capacity(32 * 1024 * 1024) // 32 MiB total weight
    .weigher(|_k: &u64, v: &Arc<Vec<u8>>| v.len() as u32)
    .build();

The weigher closure runs on every insert, allowing eviction based on actual memory footprint rather than entry count.

Expiration Policies and House-Keeper Tuning

For aggressive TTL enforcement, reduce the run_interval in HousekeeperConfig. The policy::ExpirationPolicy in src/policy/expiration_policy.rs handles TTL, TTI, and per-entry Expiry traits. Short durations increase churn and CPU usage, while long durations risk stale data.

Invalidation Support Overhead

The support_invalidation_closures feature enables conditional bulk removal via invalidate_entries_if, but incurs extra memory per entry and occasional lock costs in BaseCache. Omit this method if you only perform single-key invalidation to save hundreds of bytes per cache and reduce latency.

Implementation Examples

Example 1: High-Concurrency Write Optimization

Use 16 segments and a fast hasher to minimize contention during bulk inserts:

use moka::sync::Cache;
use ahash::RandomState;

fn main() {
    let cache = Cache::builder()
        .max_capacity(1_000_000)
        .segments(16)  // Spread across 16 shards
        .build_with_hasher(RandomState::default());

    std::thread::scope(|s| {
        for i in 0..8 {
            s.spawn(move || {
                for key in (i * 125_000)..((i + 1) * 125_000) {
                    cache.insert(key, format!("value-{key}"));
                }
            });
        }
    });

    println!("Entries inserted: {}", cache.entry_count());
}

Example 2: Memory-Aware Caching with Weighers

Evict based on byte size rather than entry count for variable-size payloads:

use moka::sync::Cache;
use std::sync::Arc;

fn main() {
    let cache = Cache::builder()
        .max_capacity(32 * 1024 * 1024) // 32 MiB limit
        .weigher(|_k: &u32, v: &Arc<Vec<u8>>| v.len() as u32)
        .build();

    for i in 0..1000 {
        let data = Arc::new(vec![0u8; 64 * 1024]); // 64 KiB each
        cache.insert(i, data);
    }

    println!("Weighted size: {} bytes", cache.weighted_size());
}

Example 3: Aggressive Expiration and House-Keeper Tuning

Force stale entry eviction every second for time-sensitive data:

use moka::sync::{Cache, HousekeeperConfig};
use std::time::Duration;

fn main() {
    let cache = Cache::builder()
        .max_capacity(10_000)
        .time_to_live(Duration::from_secs(10))
        .housekeeper_config(HousekeeperConfig {
            run_interval: Duration::from_secs(1),
            shards_per_run: 8,
        })
        .build();

    cache.insert("session", "data");
    std::thread::sleep(Duration::from_secs(12));
    
    assert!(cache.get(&"session").is_none());
}

Example 4: Disabling Invalidation Closures for Low Latency

Remove overhead from conditional invalidation if unused:

use moka::sync::Cache;

fn main() {
    let cache = Cache::builder()
        .max_capacity(50_000)
        // Deliberately omit .support_invalidation_closures()
        .build();
    
    // Standard insert/get operations only
    cache.insert(1, "value");
}

Summary

  • Use segments(N) to reduce lock contention when multiple threads write concurrently; each segment operates as an independent lock-free hash table.
  • Configure a fast hasher via build_with_hasher (e.g., ahash::RandomState) for simple keys to reduce hashing latency compared to the default SipHash.
  • Implement a weigher closure for variable-size values to ensure max_capacity reflects actual memory usage rather than entry count.
  • Tune HousekeeperConfig by adjusting run_interval and shards_per_run to balance expiration freshness against CPU overhead.
  • Disable support_invalidation_closures unless you require conditional bulk invalidation, as it adds per-entry memory overhead and occasional synchronization costs.
  • Prefer TinyLFU (the default eviction policy in src/policy/eviction_policy.rs) for general workloads with recurring keys; consider FIFO only if hit rates are uniformly low and you need simpler bookkeeping.

Frequently Asked Questions

How do I minimize thread contention in a high-write Moka cache?

Configure the cache with segments(16) or higher via CacheBuilder::segments in src/sync/builder.rs (lines 100–126). This spreads the lock-free hash table across independent shards, ensuring that concurrent writes rarely collide on the same cache lines. Combine this with a fast hasher like ahash to remove hashing bottlenecks.

What is the performance impact of the house-keeper configuration?

The HousekeeperConfig struct controls how often the timer wheel advances to evict expired entries. A shorter run_interval (e.g., 1 second) reduces the window for stale data but increases CPU cycles spent on maintenance. Conversely, longer intervals reduce CPU usage but may retain expired entries longer. The shards_per_run field further controls the amortized cost per tick by limiting how many wheel shards are processed per interval.

Should I use the synchronous or asynchronous Moka cache for maximum performance?

The synchronous cache in src/sync/cache.rs provides the lowest latency because it avoids the extra Arc indirection and runtime coordination required by the async implementation in src/future/*.rs. Use the async cache only when your application already requires an async runtime; for pure throughput optimization, prefer the sync variant.

When should I use a custom weigher instead of entry-count limits?

Use a custom weigher (configured via CacheBuilder::weigher in src/sync/builder.rs) when your values have highly variable sizes, such as large byte vectors or nested structures. This ensures that max_capacity represents actual weighted size (e.g., bytes) rather than entry count, preventing memory pressure from a few oversized entries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →