How CacheGuard and Prefix Replacement Optimizers Work at the Provider Level in Caveman
Both optimizers operate as lightweight, content-blind middleware in the Caveman provider proxy, using epoch identifiers and SHA-256 prefix hashes to enforce cache consistency and enable fast prefix substitution without parsing request payloads.
Caveman's provider proxy layer implements two specialized optimizers—CacheGuard and Prefix Replacement—that intercept requests before they reach downstream LLM providers. These components optimize performance and ensure correctness by operating solely on envelope metadata rather than inspecting actual request payloads. Understanding how these optimizers function at the provider level reveals how the Caveman platform maintains cache safety across provider upgrades while minimizing redundant transformations.
CacheGuard: Enforcing Epoch Consistency
CacheGuard guarantees that a provider's cache epoch does not drift while a prefix remains unchanged. Implemented primarily in shared/platform/cacheguard/cacheguard.go, this optimizer inspects every request's envelope metadata to validate consistency between the EpochID field and the PrefixSHA256 hash.
When the prefix changes between requests, CacheGuard flags a prefix-drift warning and forces a fresh transform. This mechanism prevents stale caches from being reused across incompatible provider versions, ensuring that semantic changes in the downstream provider trigger automatic cache invalidation.
The Inspect Method and Decision Logic
The core of CacheGuard resides in the Inspect() method, which accepts an input struct containing the EpochID, PrefixSHA256, and associated metadata. Based on this inspection, CacheGuard returns one of two decision constants:
- DecisionTransformLiveOnly – Freezes the prefix and rejects any request attempting to mutate it after the first inspection. This decision enables transformation while locking the cache state.
- DecisionPassThrough – Allows the request to proceed unchanged when the prefix is stable or when the provider reports no known boundary or adapter information.
Handling Prefix Drift
When CacheGuard detects that a request's prefix differs from the established epoch boundary, it generates a WarningPrefixDrift signal. This warning triggers an immediate cache invalidation, forcing the system to bypass the prefix cache and generate a fresh transformation. The detection logic ensures that any provider upgrade or configuration change that alters the prefix hash automatically invalidates cached entries without requiring manual intervention.
Prefix Replacement: Fast In-Memory Caching
Prefix Replacement functions as a high-speed, in-memory cache that substitutes incoming request prefixes with cached versions when the prefix is known to be stable. Unlike CacheGuard, which focuses on validation, this optimizer serves as a read-through cache implemented in proxy/internal/store/prefix_cache.go.
The store maintains a mapping of prefix → cached entry, allowing the proxy to serve cache hits without forwarding requests to the downstream provider. When a request arrives with a SHA-256 prefix hash that exists in the cache, the optimizer returns the cached transformed envelope immediately.
Lookup and Put Operations
The prefix_cache.go implementation provides two primary methods for cache management:
Lookup()– Accepts aPrefixSHA256value and returns the cached entry if present and valid. On a cache hit, the request bypasses the provider entirely.Put()– Writes a new entry to the cache after a successful transformation by the downstream provider. This method updates the internal mapping only when the complete transformed response is available.
Eviction Strategy and Fallback Behavior
The Prefix Replacement store implements automatic eviction to maintain cache integrity. Entries are removed under two conditions: when a write error occurs during the Put() operation, or when the underlying prefix changes (as detected by CacheGuard). If a lookup returns a miss, or if the cached entry is incomplete or corrupted, the request falls back to the normal processing path. The failed entry is immediately evicted to prevent serving stale data on subsequent requests.
Integration Flow in the Gateway
Both optimizers integrate into the request pipeline through proxy/internal/gateway/server.go, where they orchestrate the flow between the client and the provider. The execution follows a strict sequence to ensure safety before performance optimization.
- CacheGuard Inspection – The gateway invokes
guard.Inspect()with the request'sEpochIDandPrefixSHA256. If the decision isDecisionPassThrough, the request bypasses all caching logic. - Prefix Lookup – For requests receiving
DecisionTransformLiveOnly, the gateway queriesprefixStore.Lookup(). A hit returns the cached response immediately. - Provider Forwarding – On a cache miss, the gateway forwards the request to the provider (such as those defined in
proxy/providers/openai/openai.go), transforms the response, and caches the result viaprefixStore.Put().
// Example: Provider-level optimization in the gateway handler
guard := cacheguard.New()
result, err := guard.Inspect(cacheguard.Input{
EpochID: request.EpochID,
PrefixSHA256: request.PrefixSHA256,
})
if err != nil || result.Decision == cacheguard.DecisionPassThrough {
forward(request)
return
}
// Check prefix cache before hitting the provider
cached, err := prefixStore.Lookup(request.PrefixSHA256)
if err == nil && cached != nil {
serve(cached)
return
}
// Miss: forward to provider, then cache
resp := forward(request)
_ = prefixStore.Put(request.PrefixSHA256, resp) // Safe fallback on error
Provider-Agnostic Safety Guarantees
Both optimizers are content-blind by design, meaning they never parse the concrete payload (such as JSON or protobuf content). Instead, they rely exclusively on envelope metadata, making them safe to deploy across any downstream LLM or API provider. This architectural choice ensures that the optimization layer remains decoupled from provider-specific implementation details while guaranteeing that cache entries are never reused across incompatible provider versions.
Summary
- CacheGuard validates epoch consistency using
Inspect()inshared/platform/cacheguard/cacheguard.go, forcing fresh transforms whenPrefixSHA256changes and returningDecisionTransformLiveOnlyorDecisionPassThroughbased on metadata. - Prefix Replacement provides fast in-memory caching via
Lookup()andPut()inproxy/internal/store/prefix_cache.go, serving cached entries without provider round-trips while automatically evicting corrupted or stale data. - The gateway integration in
proxy/internal/gateway/server.gosequences these optimizers to validate safety before attempting cache retrieval, ensuring correct behavior across provider upgrades. - Both optimizers operate on envelope metadata only, making them provider-agnostic and safe for any downstream LLM integration.
Frequently Asked Questions
What is the relationship between CacheGuard and Prefix Replacement?
CacheGuard acts as a safety gate that validates epoch and prefix consistency before allowing cache utilization, while Prefix Replacement serves as the performance layer that stores and retrieves transformed responses. CacheGuard must approve a request via DecisionTransformLiveOnly before Prefix Replacement attempts a cache lookup, ensuring that stale prefixes never enter the cache.
How does CacheGuard detect provider version changes?
CacheGuard detects version changes through prefix drift. When a request arrives with a PrefixSHA256 hash that differs from the established epoch boundary, the Inspect() method generates a WarningPrefixDrift signal. This forces a cache miss and prevents the system from reusing transformations generated under previous provider semantics.
What happens when the prefix cache returns a miss?
When prefixStore.Lookup() returns a miss or encounters a corrupted entry, the request falls back to the standard provider forwarding path. After receiving and transforming the response from the downstream provider, the gateway executes prefixStore.Put() to populate the cache for subsequent requests, ensuring the next identical prefix hits the cache.
Are these optimizers specific to certain LLM providers?
No, both optimizers are provider-agnostic. They never inspect the actual request payload (JSON or protobuf) and instead operate solely on envelope metadata like EpochID and PrefixSHA256. This design allows them to function identically across different providers, as demonstrated by their integration points in proxy/providers/openai/openai.go and other provider implementations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →