How the Semantic Cache Improves WorkWeave Router Performance

The WorkWeave Router uses an in-memory LRU semantic cache that eliminates redundant upstream LLM calls by matching near-identical prompts through cosine similarity of their embeddings, reducing latency and API costs.

The workweave/router repository implements an intelligent caching layer that deduplicates non-streaming requests based on semantic meaning rather than exact text matching. By storing upstream responses and retrieving them when similar prompts arrive, the router avoids expensive network round-trips to LLM providers while maintaining per-installation isolation.

What the Semantic Cache Stores

Each cached entry in internal/router/cache/cache.go contains the complete upstream response in its original wire format, including HTTP status, headers, and body. According to lines 19-24 of the cache implementation, the CachedResponse struct preserves the full provider response to ensure downstream clients receive identical data whether served from cache or fresh upstream calls.

Embedding Key Strategy

Rather than storing the full 3 KB embedding vector in the map, the cache generates a 16-byte SHA-256 digest of the L2-normalized embedding to use as the LRU key. As implemented in lines 23-26 of internal/router/cache/cache.go, this hashing strategy minimizes memory overhead while maintaining unique identifier integrity for each prompt embedding.

Configuration and Thresholds

The DefaultConfig struct defined in lines 45-55 of internal/router/cache/cache.go establishes baseline operational parameters. Default settings include a similarity threshold of approximately 0.95, a TTL of one hour, a maximum body size of 1 MiB, and limits on buckets per installation to prevent unbounded growth.

  • Default Threshold: ~0.95 cosine similarity required for cache hits
  • TTL: 1 hour before automatic eviction
  • Max Body Size: 1 MiB (oversized responses are dropped)
  • Bucket Sharding: Per-installation isolation prevents cross-tenant data leaks

Cache Lookup Flow

The Cache.Lookup method in internal/router/cache/cache.go (lines 42-45) orchestrates retrieval by accepting the installation ID, inbound format (Anthropic or OpenAI), the query embedding, and a list of candidate cluster IDs. This design allows the router to check multiple semantic clusters for a potential match before falling back to the upstream provider.

Bucket Selection and TTL Eviction

For each candidate cluster ID, the lookup fetches the appropriate bucket from the LRU store. If a bucket is missing for a specific cluster, the lookup proceeds to the next cluster (lines 49-53). During this process, entries older than the TTL are evicted on-the-fly (lines 60-64) to ensure stale data never returns to clients.

Similarity Matching

The router computes cosine similarity between the incoming prompt's embedding and stored embeddings within the targeted bucket. When the similarity score meets the per-cluster threshold defined in the configuration, the cached response returns immediately (lines 65-68), bypassing the upstream provider entirely.

Cache Store Flow

After processing a request, the router invokes Cache.Store (lines 74-93 in internal/router/cache/cache.go) to persist the response. The method validates that response bodies do not exceed the maximum size limit before storage. To prevent mutation issues, the implementation creates a deep copy of the embedding vector (lines 87-90) before insertion into the LRU cache.

Performance Impact

By short-circuiting duplicate prompts, the semantic cache eliminates network round-trips to upstream LLM providers for requests that are semantically equivalent. This architecture delivers three primary benefits:

  • Latency Reduction: Cache hits return in microseconds rather than seconds required for provider API calls
  • Cost Optimization: Reduced quota consumption and lower billing from eliminated redundant requests
  • Tenant Isolation: Bucket sharding ensures per-installation data separation, preventing cache pollution between different users or organizations

Implementation Example

The following example demonstrates integrating the cache into a proxy handler:

// Example: Using the semantic cache inside a proxy handler
func (s *Service) handleChat(ctx context.Context, req *ChatRequest) (*ChatResponse, error) {
    // 1️⃣ Compute an L2‑normalized embedding of the prompt (omitted here)
    embed := computeEmbedding(req.Prompt)

    // 2️⃣ Try to fetch a cached response
    if cached, ok := s.cache.Lookup(req.InstallationID, cache.FormatOpenAI,
        embed, []int{req.ClusterID}, req.ClusterVersion, req.KnobsHash); ok {
        // Cache hit – return the cached upstream response directly
        return &ChatResponse{Body: cached.Body, Headers: cached.Headers, Status: cached.StatusCode}, nil
    }

    // 3️⃣ No hit – forward the request to the upstream provider
    upstreamResp, err := s.provider.Call(ctx, req)
    if err != nil { return nil, err }

    // 4️⃣ Store the fresh response for future identical prompts
    s.cache.Store(req.InstallationID, cache.FormatOpenAI, embed,
        req.ClusterID, cache.CachedResponse{
            StatusCode: upstreamResp.StatusCode,
            Headers:    upstreamResp.Header,
            Body:       upstreamResp.Body,
        }, req.ClusterVersion, req.KnobsHash)

    return upstreamResp, nil
}

To instantiate the cache with custom limits in your application entry point:

// Example: Creating the cache with custom limits (e.g., in `cmd/router/main.go`)
cfg := cache.Config{
    DefaultThreshold: 0.97,          // stricter similarity
    BucketSize:       2048,          // larger per‑bucket LRU
    MaxBucketsPerInstallation: 256, // tighter memory budget
    TTL:              30 * time.Minute,
}
semanticCache := cache.New(cfg)

Summary

  • The semantic cache stores full upstream HTTP responses keyed by SHA‑256 digests of L2‑normalized embeddings in internal/router/cache/cache.go
  • Cosine similarity matching with a default threshold of 0.95 identifies near‑duplicate prompts without exact string matching
  • TTL‑based eviction and size limits (1 MiB bodies) prevent unbounded memory growth while filtering oversized responses
  • Per‑installation bucket sharding ensures multi‑tenant isolation through dedicated LRU buckets per installation ID
  • Cache hits eliminate provider API calls, reducing latency from seconds to microseconds and lowering operational costs

Frequently Asked Questions

How does the semantic cache determine if two prompts are similar?

The cache computes cosine similarity between the L2‑normalized embedding of the incoming prompt and embeddings stored in the candidate bucket. If the similarity score meets or exceeds the configured threshold (default ~0.95 as defined in DefaultConfig), the prompts are considered semantically identical and the cached response is returned immediately without contacting the upstream provider.

What happens to cached entries that exceed the TTL?

Entries older than the configured TTL (default one hour) are evicted on‑the‑fly during the lookup process. As implemented in lines 60-64 of internal/router/cache/cache.go, the Cache.Lookup method checks timestamps during bucket iteration and removes expired items before evaluating similarity scores, ensuring stale responses never reach clients.

How does the cache handle oversized responses?

The Cache.Store method enforces a maximum body size of 1 MiB by default (configurable via MaxBodySize). Responses exceeding this limit are dropped and not inserted into the cache, preventing memory exhaustion from large LLM outputs while ensuring the cache only retains viable, quick-to-retrieve entries as shown in lines 74-93 of the implementation.

Is the semantic cache shared across all installations?

No. The cache implements per‑installation isolation through bucket sharding. Each installation ID maps to dedicated LRU buckets, ensuring that cached data from one tenant cannot be served to another, maintaining strict data separation boundaries within the workweave/router infrastructure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →