Caveman Context Recovery (CCR): How It Preserves Exact Original Bytes in Lossy Compression

Caveman Context Recovery (CCR) is a content-addressable storage system that preserves the exact original bytes of any lossy-compressed payload by generating SHA-256-based handles and storing full payloads for deterministic retrieval.

Caveman Context Recovery (CCR) enables the Caveman engine to compress low-value context aggressively while maintaining the ability to recover original source material byte-for-byte. Implemented in the JuliusBrussee/caveman repository, CCR functions as a deterministic bridge between compression and data fidelity. This article examines the core architecture, dual-backend implementation, and operational security properties of Caveman Context Recovery.

Core Architecture of Caveman Context Recovery

CCR operates as a typed object store where identical payloads always map to the same identifier, ensuring idempotent storage. The system validates and normalizes all objects before persistence, guaranteeing consistent retrieval semantics across different runtime environments.

Handle Generation and Content Addressing

Each stored payload receives a handle derived from the first 16 bytes of its SHA-256 hash, hex-encoded and prefixed with ccr_ according to the technical documentation in docs/technical/context-recovery.md【L9-L14】. This content-addressable approach means that storing the same data twice returns the identical handle without duplicating storage. The handle serves as an immutable pointer to the original bytes, allowing agents to reference large payloads using compact identifiers.

The Object Model

CCR stores typed objects defined in engine/ccr/store.go【L74-L91】 through the Object struct, which encapsulates metadata including Type, Source, SessionID, timestamps, and the raw Data byte slice. Before storage, the prepareObject function validates and normalizes these fields【L101-L148】, ensuring that all objects meet structural requirements before backend insertion. This validation layer prevents corrupted or malformed data from entering the store.

Dual-Backend Storage Implementation

The store implementation splits by runtime environment to optimize for different deployment targets. As defined in engine/ccr/store.go【L10-L13】, CCR uses a SQLite database for native engines and a pure-Go in-memory map for WebAssembly environments, exposed through store_sqlite.go and store_wasm.go respectively. Both backends implement the identical API surface, rendering the engine agnostic to the underlying storage mechanism. This abstraction allows seamless operation across server-side and browser-based deployments without code changes.

Storage Constraints and Security Properties

CCR enforces operational boundaries through configurable capacity limits and explicit security assumptions. Understanding these constraints is critical for safe deployment in production environments.

Capacity Limits and Budget Enforcement

CCR maintains a configurable payload budget defaulting to 512 MiB. When PutObject receives a new object, it checks whether insertion would exceed this limit; if so, it returns ErrBudgetExceeded and preserves the original payload unchanged【L31-L34】【docs/technical/context-recovery.md L60-L66】. This fail-safe prevents unbounded storage growth while ensuring that compression operations never block on storage constraints.

Security Model and Trust Boundaries

CCR guarantees availability of original bytes but does not provide encryption, access control, or secret redaction【docs/technical/context-recovery.md L68-L80】. The system assumes a trusted local operator and should not be used to share handles across trust boundaries without additional authorization layers. Handles like ccr_0123456789abcdef0123456789abcdef reveal only the hash prefix, but the storage backend itself remains unencrypted, requiring filesystem or database-level security controls for sensitive data.

Working with CCR Objects

Retrieval mechanisms use typed identifiers and URI schemes to validate objects before returning data, ensuring type safety across the retrieval pipeline.

Typed Objects and URI Retrieval

Objects can be addressed using identifiers beginning with ccr_obj_. Clients retrieve them using a ccr:// URI scheme, which validates the pointer shape and object type before returning the underlying data【docs/technical/context-recovery.md L39-L47】. This URI-based addressing allows integration with standard HTTP-like retrieval patterns while maintaining CCR's validation requirements.

Operational Diagnostics

When retrieval fails, the engine advises checking runtime/store consistency, file existence, handle integrity, and capacity errors as documented in the operational checks section【L88-L97】. Common failure modes include attempting to retrieve handles from exhausted budgets or querying handles generated in different storage backends.

Practical Implementation Examples

The following examples demonstrate storing objects, retrieving original payloads, and using the CLI tool.

Storing an Object and Obtaining a Handle

obj := ccr.Object{
    Type:      ccr.ObjectFileObservation,
    SessionID: "session-123",
    Source:    "example.txt",
    Data:      []byte("the original file contents…"),
}
handle, err := store.PutObject(obj) // handle is a string like "ccr_…"
if err != nil {
    // handle ErrBudgetExceeded or other errors
}

Retrieving the Original Payload

recovery, err := store.Get(handle)
if err != nil {
    // handle ErrNotFound, etc.
}
original := recovery.Original // []byte with the exact bytes stored originally

Using the CLI Tool

caveman tools retrieve ccr_0123456789abcdef0123456789abcdef

Summary

  • Caveman Context Recovery (CCR) is a content-addressable store that preserves exact original bytes behind SHA-256-based handles, enabling lossless recovery of lossy-compressed data.
  • Handles are generated from the first 16 bytes of SHA-256 hashes (prefixed with ccr_), ensuring identical payloads always map to the same identifier.
  • Dual backends (SQLite for native, in-memory map for WebAssembly) provide runtime flexibility while exposing a unified API via engine/ccr/store.go.
  • Capacity enforcement defaults to 512 MiB with ErrBudgetExceeded returned when budgets are exhausted, preventing unbounded growth.
  • Security properties explicitly exclude encryption and access control; CCR assumes trusted local operators and requires external authorization for cross-boundary handle sharing.

Frequently Asked Questions

What exactly does Caveman Context Recovery store?

CCR stores typed objects containing metadata (type, source, session ID, timestamps) and the raw original byte slice. According to engine/ccr/store.go, these objects are validated through prepareObject before storage to ensure structural integrity. The system preserves the exact bytes of lossy-compressed payloads so they can be retrieved byte-for-byte later.

How does CCR handle different runtime environments?

CCR implements runtime-specific backends: a SQLite database for native engines and a pure-Go in-memory map for WebAssembly environments. Both implementations expose identical APIs, allowing the Caveman engine to operate without awareness of the underlying storage mechanism. This design enables consistent behavior across server-side and browser-based deployments.

What happens when the storage budget is exceeded?

When adding a new object would exceed the configured budget (default 512 MiB), the PutObject function returns ErrBudgetExceeded and leaves the original payload unchanged. This prevents storage exhaustion while ensuring that compression operations fail gracefully rather than blocking or corrupting data.

Is CCR suitable for storing sensitive or shared data?

No, CCR is explicitly designed for trusted local operators only. According to the technical documentation, CCR provides no encryption, access control, or secret redaction. Handles should not be shared across trust boundaries without additional authorization layers, as the storage backend remains unencrypted and accessible to anyone with filesystem or database access.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →