What Information Is Stored by Caveman Context Recovery?

Caveman Context Recovery (CCR) persists the byte-exact original payload, MIME type, compressor metadata, token accounting metrics, and optional auxiliary data, all indexed by a deterministic SHA‑256 recovery handle, alongside typed working‑memory objects containing provenance and lifecycle information.

Caveman Context Recovery (CCR) is the durable storage layer in the JuliusBrussee/caveman engine that guarantees lossless retrieval of payloads processed through the engine’s lossy “S4” compression. When content is compressed, CCR generates a content‑addressed recovery handle and stores comprehensive metadata alongside the original bytes in a local SQLite database, enabling exact reconstruction via the caveman_retrieve tool.

Recovery Handle and Payload Metadata

Content‑Addressed Identifier

In engine/ccr/store.go (lines 24‑27), the recovery handle is generated as a SHA‑256 hash of the original payload bytes. Because the handle is derived from the content itself, identical payloads automatically map to the same handle, ensuring natural deduplication across sessions. The handle appears externally with a ccr_ prefix (e.g., ccr_9f4a…) and serves as the primary key for all retrieval operations.

Stored Metadata Fields

According to engine/ccr/store.go (lines 82‑90), each recovery entry persists the following fields:

  • handle – The SHA‑256 identifier used as the database key.
  • content_type – MIME‑type or logical type of the payload (e.g., text, json, toon).
  • compressor – Name of the compression algorithm employed (caveman_compress, caveman_head, etc.).
  • tokens_before – Inferred token count of the original payload, used for budget accounting.
  • tokens_after – Inferred token count of the compressed representation.
  • original – The full, byte‑exact original payload stored as a []byte slice.
  • metadata – Optional auxiliary BLOB for provenance data, timestamps, or custom tags.

When a client invokes caveman_retrieve (defined in mcp/engine_tools.go), the engine queries these fields and returns the exact original bytes, exempting the result from standard size caps to ensure fidelity.

Working Memory Objects

In addition to compression recovery data, CCR stores typed working‑memory objects such as file observations, search results, and test results. As defined in engine/ccr/store.go, each object contains:

  • object_id – Deterministic ID derived from type, session, and source parameters.
  • type – Closed‑set ObjectType enum (FileObservation, SearchResult, etc.).
  • content_hash – SHA‑256 hash of the object’s raw data for integrity verification.
  • session_id, source, repository_state – Provenance fields tracking origin context.
  • currentness – Status flag indicating current, stale, or archived state.
  • lifecycle – Storage tier designation (hot, warm, cold).
  • original_byte_length / stored_byte_length – Accounting metrics comparing raw versus stored size.
  • data – The raw byte payload of the object.

Persistence and Retrieval Architecture

All CCR data persists to a local SQLite database at ~/.caveman/ccr.db, implemented in engine/ccr/store_sqlite.go with busy‑retry handling and strict schema enforcement. The RecoveryClient located in packages/pi-extension/src/recovery.ts manages the retrieval lifecycle: it spawns the caveman-mcp binary, sends JSON‑RPC requests containing the handle, and returns the original payload or an explicit cave_recovery_unavailable error if the handle is unknown (lines 76‑92).

Practical Examples

Generating a Recovery Handle via Compression

import { compress } from "caveman-sdk";

const text = "… a very large log …";
const result = await cave.tools.compress({ input: text });
/* result contains:
   {
     compressed: "...",          // shortened representation
     ratio: 0.23,
     tokens_before: 1200,
     tokens_after: 276,
     recovery_handle: "ccr_9f4a…"  // SHA-256 handle stored in CCR
   }
*/

The SDK forwards the call to the engine, which populates the CCR store via engine/ccr/store.go.

Retrieving Original Payloads

import { retrieve } from "caveman-sdk";

const handle = "ccr_9f4a…";
const { text, isError } = await cave.tools.retrieve({ recovery_handle: handle });
if (!isError) {
  console.log("Original payload:", text);  // Exact bytes from CCR
}

Under the hood, this invokes RecoveryClient.retrieve, which queries ~/.caveman/ccr.db through the caveman-mcp binary.

Direct RecoveryClient Usage

import { RecoveryClient } from "./recovery.ts";

const client = new RecoveryClient();           // Resolves binary automatically
await client.ensure();                         // Spawns caveman-mcp if needed
const { text } = await client.retrieve(
  "ccr_9f4a…", 
  undefined, 
  undefined
);
console.log(text);                             // Byte-exact original content

Summary

  • Caveman Context Recovery stores byte‑exact original payloads indexed by SHA‑256 content‑addressed handles (lines 24‑27 in store.go).
  • Each entry tracks content type, compressor algorithm, and token accounting (tokens_before/tokens_after) for budget management.
  • Optional metadata BLOBs support auxiliary provenance and timestamp information.
  • Working‑memory objects extend storage to typed data with ObjectType enums, lifecycle states (hot/warm/cold), and currentness flags.
  • All data persists to ~/.caveman/ccr.db via the SQLite backend (store_sqlite.go).
  • Retrieval uses the RecoveryClient (packages/pi-extension/src/recovery.ts) to communicate with the caveman-mcp binary and return exact bytes or explicit errors.

Frequently Asked Questions

What is a recovery handle in Caveman Context Recovery?

A recovery handle is a content‑addressed identifier prefixed with ccr_ that represents the SHA‑256 hash of the original payload. It functions as the primary key for retrieving the exact original bytes from the CCR store, ensuring that identical content always resolves to the same handle for natural deduplication.

How does Caveman Context Recovery ensure data integrity?

CCR employs SHA‑256 hashing for both recovery handles and working‑memory object content hashes. The original payload is stored as an immutable []byte slice in the SQLite database, and the RecoveryClient returns this data without transformation, guaranteeing bit‑for‑bit exact retrieval or an explicit cave_recovery_unavailable error.

Where is Caveman Context Recovery data physically stored?

All CCR data persists to a local SQLite database located at ~/.caveman/ccr.db. The schema definitions and persistence logic reside in engine/ccr/store_sqlite.go, while the core data structures are defined in engine/ccr/store.go.

What types of objects can be stored besides compressed payloads?

In addition to compression recovery entries, CCR stores typed working‑memory objects such as FileObservation, SearchResult, and TestResult. These objects include provenance fields (session_id, source, repository_state), lifecycle status flags (hot, warm, cold), and comprehensive size accounting metrics (original_byte_length, stored_byte_length).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →