How WorkWeave Router's Semantic Cache Determines Response Equivalence
WorkWeave Router determines response equivalence by comparing the cosine similarity of L2-normalized prompt embeddings against a configurable threshold, returning cached responses when similarity exceeds the per-cluster or global limit.
The workweave/router repository implements an intelligent semantic cache that deduplicates near-duplicate non-streaming responses based on prompt semantics rather than exact string matching. This system enables efficient reuse of upstream LLM responses when incoming requests are sufficiently similar to previously cached queries, reducing latency and provider costs.
L2 Normalization and Cosine Similarity Calculation
The semantic cache requires that all embedding vectors supplied to it be L2-normalized. This normalization ensures that the cosine similarity between two embeddings can be computed efficiently as a simple dot product of the vectors. According to the source code in internal/router/cache/cache.go (lines 52-64), the cache implementation relies on this mathematical property to avoid expensive square-root calculations during the hot path of cache lookups.
When the Lookup method receives a query embedding, it compares this vector against stored embeddings using dot product operations, treating the result as the cosine similarity score.
Configurable Similarity Thresholds
WorkWeave Router supports granular control over equivalence detection through both global and cluster-specific thresholds. As implemented in internal/router/cache/cache.go (lines 33-40), the cache configuration includes:
DefaultThreshold: A global fallback value used when no cluster-specific setting existsPerClusterThreshold: A map allowing individual routing clusters to define their own sensitivity levels
During lookup evaluation, the system first checks for a cluster-specific threshold in the PerClusterThreshold map. If absent, it falls back to the DefaultThreshold, enabling fine-grained control over semantic matching strictness across different use cases.
The Lookup and Matching Process
The core equivalence logic resides in the Lookup method, which spans lines 44-71 in internal/router/cache/cache.go. This method executes the following sequence:
- Cluster Iteration: Iterates over the list of candidate cluster IDs supplied by the router
- Expiration Filtering: Retrieves stored entries for each cluster, discarding any where
now.Sub(e.storedAt) > cfg.TTL - Similarity Computation: Calculates
cosine(embedding, e.embedding)via dot product - Threshold Comparison: If the similarity is ≥ threshold, returns the cached response immediately, treating the new request as semantically equivalent
This process ensures that only fresh, sufficiently similar responses are considered for reuse.
Cache Storage and Key Generation
When storing responses via the Store method, the cache performs deep copying of embedding vectors to prevent external mutation of cached data. As detailed in internal/router/cache/cache.go (lines 37-51), the cache generates storage keys using the entryKeyFor helper function.
This function creates a truncated SHA-256 hash of the embedding vector, ensuring that identical embeddings map to the same storage key. Consequently, storing a new response with an identical embedding automatically replaces the previous entry rather than creating duplicates, maintaining cache integrity and preventing unbounded growth.
Implementation Example
The following Go code demonstrates creating a cache instance, storing a response with L2-normalized embeddings, and performing a lookup to test for semantic equivalence:
// Create a cache with default settings.
c := cache.New(cache.Config{})
// Assume we have an L2‑normalized embedding for the prompt.
emb := []float32{0.2, 0.5, 0.3, 0.9}
// Store a response for installation "inst-42", Anthropic format, cluster 7.
resp := cache.CachedResponse{
StatusCode: http.StatusOK,
Headers: http.Header{"Content-Type": []string{"application/json"}},
Body: []byte(`{"choice":"A"}`),
}
c.Store("inst-42", cache.FormatAnthropic, emb, 7, resp, "v1", 0)
// Later, a request with the same installation and a similar embedding arrives.
similarEmb := []float32{0.21, 0.49, 0.31, 0.88} // still L2‑normalized
cached, hit := c.Lookup("inst-42", cache.FormatAnthropic,
similarEmb, []int{7, 8}, "v1", 0)
if hit {
fmt.Println("Cache hit! Reusing response:", string(cached.Body))
} else {
fmt.Println("Cache miss – need to call upstream provider.")
}
In this example, the Lookup call returns the stored response when the cosine similarity between similarEmb and the cached embedding meets the per-cluster threshold (default 0.95).
Key Files and Architecture
The semantic cache implementation is concentrated in the following files within the workweave/router repository:
| File | Role |
|---|---|
internal/router/cache/cache.go |
Core cache implementation containing Lookup, Store, threshold logic, and cosine similarity calculations |
internal/router/cache/cache_internal_test.go |
Unit tests verifying cache behavior, bucket allocation, and TTL handling |
internal/router/cache/CLAUDE.md |
High-level design documentation for the cache package |
internal/router/cache/AGENTS.md |
Agent-focused operational documentation |
Summary
- Cosine similarity between L2-normalized embeddings serves as the mathematical foundation for determining semantic equivalence
- Configurable thresholds at both global and per-cluster levels provide flexibility for different routing strategies
- The
Lookupmethod incache.gofilters expired entries and compares embeddings against cluster-specific thresholds during the hot path - SHA-256 hashed keys ensure identical embeddings replace rather than duplicate existing cache entries
- The system requires input embeddings to be L2-normalized to enable efficient dot-product cosine calculations
Frequently Asked Questions
What embedding format does WorkWeave Router's semantic cache require?
The cache requires L2-normalized embedding vectors as input. This normalization ensures that cosine similarity can be computed as a simple dot product between vectors, eliminating the need for expensive magnitude calculations during cache lookups. Attempting to use non-normalized embeddings will result in incorrect similarity scores.
How does the cache handle different similarity thresholds for different clusters?
WorkWeave Router supports per-cluster threshold configuration through the PerClusterThreshold map in the cache configuration. When processing a lookup for a specific cluster, the system checks this map first; if no cluster-specific value exists, it falls back to the DefaultThreshold. This allows high-sensitivity clusters (e.g., creative writing) to use lower thresholds while strict clusters (e.g., code generation) require higher similarity.
What happens when multiple cached entries meet the similarity threshold?
During the Lookup process, the method iterates through candidate clusters in the order provided. The first entry that exceeds the similarity threshold and has not expired is returned immediately. This short-circuit evaluation ensures minimal latency, though the specific entry returned depends on the order of cluster IDs passed to the lookup function.
How does the cache prevent duplicate storage of identical embeddings?
When storing responses, the entryKeyFor helper function generates a truncated SHA-256 hash of the embedding vector. This hash serves as the storage key, ensuring that subsequent stores with identical embeddings overwrite the previous entry rather than creating duplicates. Combined with deep-copying of embedding data, this mechanism maintains cache consistency and prevents memory leaks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →