How WorkWeave Router's Semantic Cache Determines Response Equivalence

WorkWeave Router determines response equivalence by comparing the cosine similarity of L2-normalized prompt embeddings against a configurable threshold, returning cached responses when similarity exceeds the per-cluster or global limit.

The workweave/router repository implements an intelligent semantic cache that deduplicates near-duplicate non-streaming responses based on prompt semantics rather than exact string matching. This system enables efficient reuse of upstream LLM responses when incoming requests are sufficiently similar to previously cached queries, reducing latency and provider costs.

L2 Normalization and Cosine Similarity Calculation

The semantic cache requires that all embedding vectors supplied to it be L2-normalized. This normalization ensures that the cosine similarity between two embeddings can be computed efficiently as a simple dot product of the vectors. According to the source code in internal/router/cache/cache.go (lines 52-64), the cache implementation relies on this mathematical property to avoid expensive square-root calculations during the hot path of cache lookups.

When the Lookup method receives a query embedding, it compares this vector against stored embeddings using dot product operations, treating the result as the cosine similarity score.

Configurable Similarity Thresholds

WorkWeave Router supports granular control over equivalence detection through both global and cluster-specific thresholds. As implemented in internal/router/cache/cache.go (lines 33-40), the cache configuration includes:

  • DefaultThreshold: A global fallback value used when no cluster-specific setting exists
  • PerClusterThreshold: A map allowing individual routing clusters to define their own sensitivity levels

During lookup evaluation, the system first checks for a cluster-specific threshold in the PerClusterThreshold map. If absent, it falls back to the DefaultThreshold, enabling fine-grained control over semantic matching strictness across different use cases.

The Lookup and Matching Process

The core equivalence logic resides in the Lookup method, which spans lines 44-71 in internal/router/cache/cache.go. This method executes the following sequence:

  1. Cluster Iteration: Iterates over the list of candidate cluster IDs supplied by the router
  2. Expiration Filtering: Retrieves stored entries for each cluster, discarding any where now.Sub(e.storedAt) > cfg.TTL
  3. Similarity Computation: Calculates cosine(embedding, e.embedding) via dot product
  4. Threshold Comparison: If the similarity is ≥ threshold, returns the cached response immediately, treating the new request as semantically equivalent

This process ensures that only fresh, sufficiently similar responses are considered for reuse.

Cache Storage and Key Generation

When storing responses via the Store method, the cache performs deep copying of embedding vectors to prevent external mutation of cached data. As detailed in internal/router/cache/cache.go (lines 37-51), the cache generates storage keys using the entryKeyFor helper function.

This function creates a truncated SHA-256 hash of the embedding vector, ensuring that identical embeddings map to the same storage key. Consequently, storing a new response with an identical embedding automatically replaces the previous entry rather than creating duplicates, maintaining cache integrity and preventing unbounded growth.

Implementation Example

The following Go code demonstrates creating a cache instance, storing a response with L2-normalized embeddings, and performing a lookup to test for semantic equivalence:

// Create a cache with default settings.
c := cache.New(cache.Config{})

// Assume we have an L2‑normalized embedding for the prompt.
emb := []float32{0.2, 0.5, 0.3, 0.9}

// Store a response for installation "inst-42", Anthropic format, cluster 7.
resp := cache.CachedResponse{
    StatusCode: http.StatusOK,
    Headers:    http.Header{"Content-Type": []string{"application/json"}},
    Body:       []byte(`{"choice":"A"}`),
}
c.Store("inst-42", cache.FormatAnthropic, emb, 7, resp, "v1", 0)

// Later, a request with the same installation and a similar embedding arrives.
similarEmb := []float32{0.21, 0.49, 0.31, 0.88} // still L2‑normalized
cached, hit := c.Lookup("inst-42", cache.FormatAnthropic,
    similarEmb, []int{7, 8}, "v1", 0)

if hit {
    fmt.Println("Cache hit! Reusing response:", string(cached.Body))
} else {
    fmt.Println("Cache miss – need to call upstream provider.")
}

In this example, the Lookup call returns the stored response when the cosine similarity between similarEmb and the cached embedding meets the per-cluster threshold (default 0.95).

Key Files and Architecture

The semantic cache implementation is concentrated in the following files within the workweave/router repository:

File Role
internal/router/cache/cache.go Core cache implementation containing Lookup, Store, threshold logic, and cosine similarity calculations
internal/router/cache/cache_internal_test.go Unit tests verifying cache behavior, bucket allocation, and TTL handling
internal/router/cache/CLAUDE.md High-level design documentation for the cache package
internal/router/cache/AGENTS.md Agent-focused operational documentation

Summary

  • Cosine similarity between L2-normalized embeddings serves as the mathematical foundation for determining semantic equivalence
  • Configurable thresholds at both global and per-cluster levels provide flexibility for different routing strategies
  • The Lookup method in cache.go filters expired entries and compares embeddings against cluster-specific thresholds during the hot path
  • SHA-256 hashed keys ensure identical embeddings replace rather than duplicate existing cache entries
  • The system requires input embeddings to be L2-normalized to enable efficient dot-product cosine calculations

Frequently Asked Questions

What embedding format does WorkWeave Router's semantic cache require?

The cache requires L2-normalized embedding vectors as input. This normalization ensures that cosine similarity can be computed as a simple dot product between vectors, eliminating the need for expensive magnitude calculations during cache lookups. Attempting to use non-normalized embeddings will result in incorrect similarity scores.

How does the cache handle different similarity thresholds for different clusters?

WorkWeave Router supports per-cluster threshold configuration through the PerClusterThreshold map in the cache configuration. When processing a lookup for a specific cluster, the system checks this map first; if no cluster-specific value exists, it falls back to the DefaultThreshold. This allows high-sensitivity clusters (e.g., creative writing) to use lower thresholds while strict clusters (e.g., code generation) require higher similarity.

What happens when multiple cached entries meet the similarity threshold?

During the Lookup process, the method iterates through candidate clusters in the order provided. The first entry that exceeds the similarity threshold and has not expired is returned immediately. This short-circuit evaluation ensures minimal latency, though the specific entry returned depends on the order of cluster IDs passed to the lookup function.

How does the cache prevent duplicate storage of identical embeddings?

When storing responses, the entryKeyFor helper function generates a truncated SHA-256 hash of the embedding vector. This hash serves as the storage key, ensuring that subsequent stores with identical embeddings overwrite the previous entry rather than creating duplicates. Combined with deep-copying of embedding data, this mechanism maintains cache consistency and prevents memory leaks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →