# How the Semantic Cache Improves WorkWeave Router Performance

> Discover how the WorkWeave Router's semantic cache boosts performance by eliminating redundant LLM calls. Reduce latency and API costs with intelligent prompt matching.

- Repository: [Weave/router](https://github.com/workweave/router)
- Tags: performance
- Published: 2026-08-30

---

**The WorkWeave Router uses an in-memory LRU semantic cache that eliminates redundant upstream LLM calls by matching near-identical prompts through cosine similarity of their embeddings, reducing latency and API costs.**

The `workweave/router` repository implements an intelligent caching layer that deduplicates non-streaming requests based on semantic meaning rather than exact text matching. By storing upstream responses and retrieving them when similar prompts arrive, the router avoids expensive network round-trips to LLM providers while maintaining per-installation isolation.

## What the Semantic Cache Stores

Each cached entry in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go) contains the complete upstream response in its original wire format, including HTTP status, headers, and body. According to lines 19-24 of the cache implementation, the `CachedResponse` struct preserves the full provider response to ensure downstream clients receive identical data whether served from cache or fresh upstream calls.

### Embedding Key Strategy

Rather than storing the full 3 KB embedding vector in the map, the cache generates a 16-byte SHA-256 digest of the L2-normalized embedding to use as the LRU key. As implemented in lines 23-26 of [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go), this hashing strategy minimizes memory overhead while maintaining unique identifier integrity for each prompt embedding.

## Configuration and Thresholds

The `DefaultConfig` struct defined in lines 45-55 of [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go) establishes baseline operational parameters. Default settings include a **similarity threshold of approximately 0.95**, a **TTL of one hour**, a **maximum body size of 1 MiB**, and limits on buckets per installation to prevent unbounded growth.

- **Default Threshold**: ~0.95 cosine similarity required for cache hits
- **TTL**: 1 hour before automatic eviction
- **Max Body Size**: 1 MiB (oversized responses are dropped)
- **Bucket Sharding**: Per-installation isolation prevents cross-tenant data leaks

## Cache Lookup Flow

The `Cache.Lookup` method in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go) (lines 42-45) orchestrates retrieval by accepting the installation ID, inbound format (Anthropic or OpenAI), the query embedding, and a list of candidate cluster IDs. This design allows the router to check multiple semantic clusters for a potential match before falling back to the upstream provider.

### Bucket Selection and TTL Eviction

For each candidate cluster ID, the lookup fetches the appropriate bucket from the LRU store. If a bucket is missing for a specific cluster, the lookup proceeds to the next cluster (lines 49-53). During this process, entries older than the TTL are evicted on-the-fly (lines 60-64) to ensure stale data never returns to clients.

### Similarity Matching

The router computes cosine similarity between the incoming prompt's embedding and stored embeddings within the targeted bucket. When the similarity score meets the per-cluster threshold defined in the configuration, the cached response returns immediately (lines 65-68), bypassing the upstream provider entirely.

## Cache Store Flow

After processing a request, the router invokes `Cache.Store` (lines 74-93 in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go)) to persist the response. The method validates that response bodies do not exceed the maximum size limit before storage. To prevent mutation issues, the implementation creates a deep copy of the embedding vector (lines 87-90) before insertion into the LRU cache.

## Performance Impact

By short-circuiting duplicate prompts, the semantic cache eliminates network round-trips to upstream LLM providers for requests that are semantically equivalent. This architecture delivers three primary benefits:

- **Latency Reduction**: Cache hits return in microseconds rather than seconds required for provider API calls
- **Cost Optimization**: Reduced quota consumption and lower billing from eliminated redundant requests
- **Tenant Isolation**: Bucket sharding ensures per-installation data separation, preventing cache pollution between different users or organizations

## Implementation Example

The following example demonstrates integrating the cache into a proxy handler:

```go
// Example: Using the semantic cache inside a proxy handler
func (s *Service) handleChat(ctx context.Context, req *ChatRequest) (*ChatResponse, error) {
    // 1️⃣ Compute an L2‑normalized embedding of the prompt (omitted here)
    embed := computeEmbedding(req.Prompt)

    // 2️⃣ Try to fetch a cached response
    if cached, ok := s.cache.Lookup(req.InstallationID, cache.FormatOpenAI,
        embed, []int{req.ClusterID}, req.ClusterVersion, req.KnobsHash); ok {
        // Cache hit – return the cached upstream response directly
        return &ChatResponse{Body: cached.Body, Headers: cached.Headers, Status: cached.StatusCode}, nil
    }

    // 3️⃣ No hit – forward the request to the upstream provider
    upstreamResp, err := s.provider.Call(ctx, req)
    if err != nil { return nil, err }

    // 4️⃣ Store the fresh response for future identical prompts
    s.cache.Store(req.InstallationID, cache.FormatOpenAI, embed,
        req.ClusterID, cache.CachedResponse{
            StatusCode: upstreamResp.StatusCode,
            Headers:    upstreamResp.Header,
            Body:       upstreamResp.Body,
        }, req.ClusterVersion, req.KnobsHash)

    return upstreamResp, nil
}

```

To instantiate the cache with custom limits in your application entry point:

```go
// Example: Creating the cache with custom limits (e.g., in `cmd/router/main.go`)
cfg := cache.Config{
    DefaultThreshold: 0.97,          // stricter similarity
    BucketSize:       2048,          // larger per‑bucket LRU
    MaxBucketsPerInstallation: 256, // tighter memory budget
    TTL:              30 * time.Minute,
}
semanticCache := cache.New(cfg)

```

## Summary

- The semantic cache stores full upstream HTTP responses keyed by SHA‑256 digests of L2‑normalized embeddings in [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go)
- Cosine similarity matching with a default threshold of 0.95 identifies near‑duplicate prompts without exact string matching
- TTL‑based eviction and size limits (1 MiB bodies) prevent unbounded memory growth while filtering oversized responses
- Per‑installation bucket sharding ensures multi‑tenant isolation through dedicated LRU buckets per installation ID
- Cache hits eliminate provider API calls, reducing latency from seconds to microseconds and lowering operational costs

## Frequently Asked Questions

### How does the semantic cache determine if two prompts are similar?

The cache computes **cosine similarity** between the L2‑normalized embedding of the incoming prompt and embeddings stored in the candidate bucket. If the similarity score meets or exceeds the configured threshold (default ~0.95 as defined in `DefaultConfig`), the prompts are considered semantically identical and the cached response is returned immediately without contacting the upstream provider.

### What happens to cached entries that exceed the TTL?

Entries older than the configured TTL (default one hour) are **evicted on‑the‑fly** during the lookup process. As implemented in lines 60-64 of [`internal/router/cache/cache.go`](https://github.com/workweave/router/blob/main/internal/router/cache/cache.go), the `Cache.Lookup` method checks timestamps during bucket iteration and removes expired items before evaluating similarity scores, ensuring stale responses never reach clients.

### How does the cache handle oversized responses?

The `Cache.Store` method enforces a **maximum body size of 1 MiB** by default (configurable via `MaxBodySize`). Responses exceeding this limit are dropped and not inserted into the cache, preventing memory exhaustion from large LLM outputs while ensuring the cache only retains viable, quick-to-retrieve entries as shown in lines 74-93 of the implementation.

### Is the semantic cache shared across all installations?

No. The cache implements **per‑installation isolation** through bucket sharding. Each installation ID maps to dedicated LRU buckets, ensuring that cached data from one tenant cannot be served to another, maintaining strict data separation boundaries within the `workweave/router` infrastructure.