# How WorkWeave Router's Semantic Cache Determines Response Equivalence

> Discover how WorkWeave Router's semantic cache determines response equivalence using cosine similarity and configurable thresholds for efficient caching.

- Repository: [Weave/router](https://github.com/workweave/router)
- Tags: deep-dive
- Published: 2026-08-30

---

**WorkWeave Router determines response equivalence by comparing the cosine similarity of L2-normalized prompt embeddings against a configurable threshold, returning cached responses when similarity exceeds the per-cluster or global limit.**

The `workweave/router` repository implements an intelligent semantic cache that deduplicates near-duplicate non-streaming responses based on prompt semantics rather than exact string matching. This system enables efficient reuse of upstream LLM responses when incoming requests are sufficiently similar to previously cached queries, reducing latency and provider costs.

## L2 Normalization and Cosine Similarity Calculation

The semantic cache requires that all embedding vectors supplied to it be **L2-normalized**. This normalization ensures that the cosine similarity between two embeddings can be computed efficiently as a simple **dot product** of the vectors. According to the source code in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go) (lines 52-64), the cache implementation relies on this mathematical property to avoid expensive square-root calculations during the hot path of cache lookups.

When the `Lookup` method receives a query embedding, it compares this vector against stored embeddings using dot product operations, treating the result as the cosine similarity score.

## Configurable Similarity Thresholds

WorkWeave Router supports granular control over equivalence detection through both global and cluster-specific thresholds. As implemented in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go) (lines 33-40), the cache configuration includes:

- **`DefaultThreshold`**: A global fallback value used when no cluster-specific setting exists
- **`PerClusterThreshold`**: A map allowing individual routing clusters to define their own sensitivity levels

During lookup evaluation, the system first checks for a cluster-specific threshold in the `PerClusterThreshold` map. If absent, it falls back to the `DefaultThreshold`, enabling fine-grained control over semantic matching strictness across different use cases.

## The Lookup and Matching Process

The core equivalence logic resides in the `Lookup` method, which spans lines 44-71 in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go). This method executes the following sequence:

1. **Cluster Iteration**: Iterates over the list of candidate cluster IDs supplied by the router
2. **Expiration Filtering**: Retrieves stored entries for each cluster, discarding any where `now.Sub(e.storedAt) > cfg.TTL`
3. **Similarity Computation**: Calculates `cosine(embedding, e.embedding)` via dot product
4. **Threshold Comparison**: If the similarity is **≥ threshold**, returns the cached response immediately, treating the new request as semantically equivalent

This process ensures that only fresh, sufficiently similar responses are considered for reuse.

## Cache Storage and Key Generation

When storing responses via the `Store` method, the cache performs deep copying of embedding vectors to prevent external mutation of cached data. As detailed in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go) (lines 37-51), the cache generates storage keys using the `entryKeyFor` helper function.

This function creates a **truncated SHA-256 hash** of the embedding vector, ensuring that identical embeddings map to the same storage key. Consequently, storing a new response with an identical embedding automatically replaces the previous entry rather than creating duplicates, maintaining cache integrity and preventing unbounded growth.

## Implementation Example

The following Go code demonstrates creating a cache instance, storing a response with L2-normalized embeddings, and performing a lookup to test for semantic equivalence:

```go
// Create a cache with default settings.
c := cache.New(cache.Config{})

// Assume we have an L2‑normalized embedding for the prompt.
emb := []float32{0.2, 0.5, 0.3, 0.9}

// Store a response for installation "inst-42", Anthropic format, cluster 7.
resp := cache.CachedResponse{
    StatusCode: http.StatusOK,
    Headers:    http.Header{"Content-Type": []string{"application/json"}},
    Body:       []byte(`{"choice":"A"}`),
}
c.Store("inst-42", cache.FormatAnthropic, emb, 7, resp, "v1", 0)

// Later, a request with the same installation and a similar embedding arrives.
similarEmb := []float32{0.21, 0.49, 0.31, 0.88} // still L2‑normalized
cached, hit := c.Lookup("inst-42", cache.FormatAnthropic,
    similarEmb, []int{7, 8}, "v1", 0)

if hit {
    fmt.Println("Cache hit! Reusing response:", string(cached.Body))
} else {
    fmt.Println("Cache miss – need to call upstream provider.")
}

```

In this example, the `Lookup` call returns the stored response when the cosine similarity between `similarEmb` and the cached embedding meets the per-cluster threshold (default 0.95).

## Key Files and Architecture

The semantic cache implementation is concentrated in the following files within the `workweave/router` repository:

| File | Role |
|------|------|
| [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go) | Core cache implementation containing `Lookup`, `Store`, threshold logic, and cosine similarity calculations |
| [`internal/router/cache/cache_internal_test.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache_internal_test.go) | Unit tests verifying cache behavior, bucket allocation, and TTL handling |
| [`internal/router/cache/CLAUDE.md`](https://github.com/workweave/router/blob/main/internal/router/cache/CLAUDE.md) | High-level design documentation for the cache package |
| [`internal/router/cache/AGENTS.md`](https://github.com/workweave/router/blob/main/internal/router/cache/AGENTS.md) | Agent-focused operational documentation |

## Summary

- **Cosine similarity** between L2-normalized embeddings serves as the mathematical foundation for determining semantic equivalence
- **Configurable thresholds** at both global and per-cluster levels provide flexibility for different routing strategies
- The **`Lookup`** method in [`cache.go`](https://github.com/workweave/router/blob/main/cache.go) filters expired entries and compares embeddings against cluster-specific thresholds during the hot path
- **SHA-256 hashed keys** ensure identical embeddings replace rather than duplicate existing cache entries
- The system requires input embeddings to be **L2-normalized** to enable efficient dot-product cosine calculations

## Frequently Asked Questions

### What embedding format does WorkWeave Router's semantic cache require?

The cache requires **L2-normalized embedding vectors** as input. This normalization ensures that cosine similarity can be computed as a simple dot product between vectors, eliminating the need for expensive magnitude calculations during cache lookups. Attempting to use non-normalized embeddings will result in incorrect similarity scores.

### How does the cache handle different similarity thresholds for different clusters?

WorkWeave Router supports per-cluster threshold configuration through the `PerClusterThreshold` map in the cache configuration. When processing a lookup for a specific cluster, the system checks this map first; if no cluster-specific value exists, it falls back to the `DefaultThreshold`. This allows high-sensitivity clusters (e.g., creative writing) to use lower thresholds while strict clusters (e.g., code generation) require higher similarity.

### What happens when multiple cached entries meet the similarity threshold?

During the `Lookup` process, the method iterates through candidate clusters in the order provided. The first entry that exceeds the similarity threshold and has not expired is returned immediately. This short-circuit evaluation ensures minimal latency, though the specific entry returned depends on the order of cluster IDs passed to the lookup function.

### How does the cache prevent duplicate storage of identical embeddings?

When storing responses, the `entryKeyFor` helper function generates a **truncated SHA-256 hash** of the embedding vector. This hash serves as the storage key, ensuring that subsequent stores with identical embeddings overwrite the previous entry rather than creating duplicates. Combined with deep-copying of embedding data, this mechanism maintains cache consistency and prevents memory leaks.