How OmniRoute's Semantic Cache Works: Implementation Guide for LLM Response Caching
OmniRoute's semantic cache is a two-tier storage system that caches deterministic LLM responses (temperature = 0) to reduce latency and cost, using SHA-256 request signatures for cache keys with LRU memory and SQLite persistence layers.
The semantic cache in OmniRoute is designed specifically for LLM workloads where identical prompts should yield identical outputs. This deep dive examines the complete architecture as implemented in diegosouzapw/OmniRoute, covering signature generation, storage mechanics, cache policies, and practical integration patterns.
Cache Architecture and Design Principles
OmniRoute's semantic cache enforces determinism as a prerequisite for caching. Only requests with temperature: 0 qualify—this prevents caching stochastic completions that would vary between identical prompts. The cache operates as a read-through, write-through layer integrated directly into the chat request pipeline.
Two-Tier Storage System
The cache uses complementary storage layers defined in src/lib/semanticCache.ts:
- In-memory LRU: Fast, bounded by
SEMANTIC_CACHE_MAX_SIZE(entry count) andSEMANTIC_CACHE_MAX_BYTES(memory). Initialized at lines 10-17. - SQLite table
semantic_cache: Persistent storage across restarts, with CRUD operations at lines 62-88.
This design optimizes for the Pareto distribution of cache access patterns: hot requests hit memory, while the SQLite backing ensures durability and larger capacity.
Signature Generation: The Cache Key
Every cacheable request is identified by a deterministic SHA-256 signature generated by generateSignature() at lines 49-71. The signature incorporates:
- Model name (e.g.,
"gpt-4o-mini") - Normalized message array (order and content)
temperature(must equal 0)top_pvalue- Optional API-key prefix for multi-tenant isolation
// Generate a cache signature for deterministic requests
import { generateSignature } from "@/lib/semanticCache";
const signature = generateSignature(
"gpt-4o-mini",
[{ role: "user", content: "Explain quantum tunneling." }],
0, // temperature: REQUIRED to be exactly 0
1, // top_p
"apiKey-12345" // optional: isolates cache by API key
);
The SHA-256 hash ensures collision-resistant, fixed-length keys suitable for both LRU and SQLite indexing.
Cache Lookup and Write Flows
Read Path: LRU-First with SQLite Fallback
getCachedResponse() (lines 6-14) implements a promotion-aware lookup:
- Check in-memory LRU first
- On miss, query SQLite
semantic_cachetable - Promote SQLite hits back into LRU
- Update hit metrics
// Cache lookup pattern in request handlers
import { getCachedResponse } from "@/lib/semanticCache";
const cached = await getCachedResponse(signature);
if (cached) {
// Return cached JSON directly—bypass provider entirely
return Response.json(cached.response);
}
Write Path: Dual-Layer Persistence
After receiving a provider response, setCachedResponse() (lines 63-71) writes to both layers:
// Store fresh response after provider completes
import { setCachedResponse } from "@/lib/semanticCache";
await setCachedResponse(
signature,
"gpt-4o-mini",
providerResponse, // Full LLM JSON payload
estimatedTokensSaved, // For cost analytics
1800000 // TTL: 30 minutes (default from SEMANTIC_CACHE_TTL_MS)
);
The default 30-minute TTL balances freshness with hit rate—adjustable per-write or via environment variable.
Cacheability Rules and Safety
OmniRoute strictly validates cache eligibility through isCacheableForRead() and isCacheableForWrite() (lines 87-100). A request is cacheable only when:
X-OmniRoute-No-Cache: trueheader is absent- Request body explicitly sets
temperature: 0
This dual-gate prevents accidental caching of non-deterministic outputs and provides escape hatches for cache-bypass scenarios.
Cache Invalidation Strategies
OmniRoute provides granular invalidation helpers at lines 92-99 and related functions:
// Invalidate all entries for a specific model (e.g., after upgrade)
import { invalidateByModel } from "@/lib/semanticCache";
const removed = await invalidateByModel("gpt-4o-mini");
console.info(`Invalidated ${removed} cached rows`);
// Additional helpers (by signature, age, or complete purge)
// invalidateBySignature(signature)
// invalidateOlderThan(timestampMs)
// invalidateAll()
Invalidation clears LRU entries before SQLite deletion to maintain consistency.
Metrics and Observability
The semantic cache records operational data in two dimensions:
cache_metricstable: Persistent storage of hits, misses, and tokens saved- LRU runtime stats: Exposed via
/api/cache/statsendpoint
Use getCacheStats() (lines 66-78) and incrementMetric() for custom instrumentation.
Integration in Request Pipeline
The semantic cache hooks into the chat core at two integration points in open-sse/handlers/chatCore/:
| File | Responsibility |
|---|---|
[semanticCache.ts](https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/open-sse/handlers/chatCore/semanticCache.ts) |
Pre-request lookup and post-response cache decision |
[semanticCacheStore.ts](https://github.com/diegosouzapw/OmniRoute/blob/release/v3.8.51/open-sse/handlers/chatCore/semanticCacheStore.ts) |
Buffers and stores complete streaming responses |
This architecture keeps cache logic decoupled from provider-specific implementations while ensuring zero overhead for non-cacheable requests.
Summary
- Determinism requirement: Only
temperature: 0requests are cached, enforced byisCacheableForRead/Write - SHA-256 signatures:
generateSignature()creates collision-resistant, normalized keys including optional API-key isolation - Two-tier storage: LRU memory for speed, SQLite
semantic_cachetable for persistence - 30-minute default TTL: Configurable via
SEMANTIC_CACHE_TTL_MSor per-write override - Promotion on miss: SQLite hits are re-inserted into LRU to optimize subsequent accesses
- Multi-model invalidation:
invalidateByModel()and helpers support operational cache management - Integrated metrics: Track hits, misses, and cost savings via
cache_metricsand runtime stats
Frequently Asked Questions
How does OmniRoute ensure cached responses are deterministic?
OmniRoute requires explicit temperature: 0 in the request body and validates absence of the X-OmniRoute-No-Cache: true header. The isCacheableForRead() and isCacheableForWrite() functions in src/lib/semanticCache.ts enforce these rules before any cache operation.
What happens when the in-memory LRU is full?
The LRU evicts least-recently-used entries based on configured SEMANTIC_CACHE_MAX_SIZE (entry count) and SEMANTIC_CACHE_MAX_BYTES (memory). Evicted entries remain in SQLite and can be promoted back to LRU on subsequent hits.
How does multi-tenant API key isolation work?
The optional apiKeyPrefix parameter in generateSignature() prepends a tenant identifier to the hash input. This creates namespace-isolated cache keys, preventing cross-tenant data leakage while allowing shared infrastructure.
Can I disable or bypass the semantic cache for specific requests?
Yes. Add the header X-OmniRoute-No-Cache: true to any request, or ensure temperature is not exactly 0. The cache integration checks these conditions before lookup and will always forward non-cacheable requests to the provider.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →