# What Information Is Stored by Caveman Context Recovery?

> Caveman Context Recovery stores byte-exact payloads, MIME types, compressor metadata, token accounting, and auxiliary data indexed by a SHA-256 handle. Discover what CCR preserves.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-09-06

---

**Caveman Context Recovery (CCR) persists the byte-exact original payload, MIME type, compressor metadata, token accounting metrics, and optional auxiliary data, all indexed by a deterministic SHA‑256 recovery handle, alongside typed working‑memory objects containing provenance and lifecycle information.**

Caveman Context Recovery (CCR) is the durable storage layer in the JuliusBrussee/caveman engine that guarantees lossless retrieval of payloads processed through the engine’s lossy “S4” compression. When content is compressed, CCR generates a content‑addressed recovery handle and stores comprehensive metadata alongside the original bytes in a local SQLite database, enabling exact reconstruction via the `caveman_retrieve` tool.

## Recovery Handle and Payload Metadata

### Content‑Addressed Identifier

In [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go) (lines 24‑27), the **recovery handle** is generated as a SHA‑256 hash of the original payload bytes. Because the handle is derived from the content itself, identical payloads automatically map to the same handle, ensuring natural deduplication across sessions. The handle appears externally with a `ccr_` prefix (e.g., `ccr_9f4a…`) and serves as the primary key for all retrieval operations.

### Stored Metadata Fields

According to [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go) (lines 82‑90), each recovery entry persists the following fields:

- **handle** – The SHA‑256 identifier used as the database key.
- **content_type** – MIME‑type or logical type of the payload (e.g., `text`, `json`, `toon`).
- **compressor** – Name of the compression algorithm employed (`caveman_compress`, `caveman_head`, etc.).
- **tokens_before** – Inferred token count of the original payload, used for budget accounting.
- **tokens_after** – Inferred token count of the compressed representation.
- **original** – The full, byte‑exact original payload stored as a `[]byte` slice.
- **metadata** – Optional auxiliary BLOB for provenance data, timestamps, or custom tags.

When a client invokes `caveman_retrieve` (defined in [`mcp/engine_tools.go`](https://github.com/JuliusBrussee/caveman/blob/main/mcp/engine_tools.go)), the engine queries these fields and returns the exact `original` bytes, exempting the result from standard size caps to ensure fidelity.

## Working Memory Objects

In addition to compression recovery data, CCR stores **typed working‑memory objects** such as file observations, search results, and test results. As defined in [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go), each object contains:

- **object_id** – Deterministic ID derived from type, session, and source parameters.
- **type** – Closed‑set `ObjectType` enum (`FileObservation`, `SearchResult`, etc.).
- **content_hash** – SHA‑256 hash of the object’s raw data for integrity verification.
- **session_id**, **source**, **repository_state** – Provenance fields tracking origin context.
- **currentness** – Status flag indicating `current`, `stale`, or `archived` state.
- **lifecycle** – Storage tier designation (`hot`, `warm`, `cold`).
- **original_byte_length** / **stored_byte_length** – Accounting metrics comparing raw versus stored size.
- **data** – The raw byte payload of the object.

## Persistence and Retrieval Architecture

All CCR data persists to a local SQLite database at `~/.caveman/ccr.db`, implemented in [`engine/ccr/store_sqlite.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store_sqlite.go) with busy‑retry handling and strict schema enforcement. The **RecoveryClient** located in [`packages/pi-extension/src/recovery.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/pi-extension/src/recovery.ts) manages the retrieval lifecycle: it spawns the `caveman-mcp` binary, sends JSON‑RPC requests containing the handle, and returns the original payload or an explicit `cave_recovery_unavailable` error if the handle is unknown (lines 76‑92).

## Practical Examples

### Generating a Recovery Handle via Compression

```typescript
import { compress } from "caveman-sdk";

const text = "… a very large log …";
const result = await cave.tools.compress({ input: text });
/* result contains:
   {
     compressed: "...",          // shortened representation
     ratio: 0.23,
     tokens_before: 1200,
     tokens_after: 276,
     recovery_handle: "ccr_9f4a…"  // SHA-256 handle stored in CCR
   }
*/

```

The SDK forwards the call to the engine, which populates the CCR store via [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go).

### Retrieving Original Payloads

```typescript
import { retrieve } from "caveman-sdk";

const handle = "ccr_9f4a…";
const { text, isError } = await cave.tools.retrieve({ recovery_handle: handle });
if (!isError) {
  console.log("Original payload:", text);  // Exact bytes from CCR
}

```

Under the hood, this invokes `RecoveryClient.retrieve`, which queries `~/.caveman/ccr.db` through the `caveman-mcp` binary.

### Direct RecoveryClient Usage

```typescript
import { RecoveryClient } from "./recovery.ts";

const client = new RecoveryClient();           // Resolves binary automatically
await client.ensure();                         // Spawns caveman-mcp if needed
const { text } = await client.retrieve(
  "ccr_9f4a…", 
  undefined, 
  undefined
);
console.log(text);                             // Byte-exact original content

```

## Summary

- **Caveman Context Recovery** stores byte‑exact original payloads indexed by SHA‑256 content‑addressed handles (lines 24‑27 in [`store.go`](https://github.com/JuliusBrussee/caveman/blob/main/store.go)).
- Each entry tracks **content type**, **compressor algorithm**, and **token accounting** (`tokens_before`/`tokens_after`) for budget management.
- Optional **metadata BLOBs** support auxiliary provenance and timestamp information.
- **Working‑memory objects** extend storage to typed data with `ObjectType` enums, lifecycle states (`hot`/`warm`/`cold`), and currentness flags.
- All data persists to `~/.caveman/ccr.db` via the SQLite backend ([`store_sqlite.go`](https://github.com/JuliusBrussee/caveman/blob/main/store_sqlite.go)).
- Retrieval uses the **RecoveryClient** ([`packages/pi-extension/src/recovery.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/pi-extension/src/recovery.ts)) to communicate with the `caveman-mcp` binary and return exact bytes or explicit errors.

## Frequently Asked Questions

### What is a recovery handle in Caveman Context Recovery?

A recovery handle is a content‑addressed identifier prefixed with `ccr_` that represents the SHA‑256 hash of the original payload. It functions as the primary key for retrieving the exact original bytes from the CCR store, ensuring that identical content always resolves to the same handle for natural deduplication.

### How does Caveman Context Recovery ensure data integrity?

CCR employs SHA‑256 hashing for both recovery handles and working‑memory object content hashes. The original payload is stored as an immutable `[]byte` slice in the SQLite database, and the `RecoveryClient` returns this data without transformation, guaranteeing bit‑for‑bit exact retrieval or an explicit `cave_recovery_unavailable` error.

### Where is Caveman Context Recovery data physically stored?

All CCR data persists to a local SQLite database located at `~/.caveman/ccr.db`. The schema definitions and persistence logic reside in [`engine/ccr/store_sqlite.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store_sqlite.go), while the core data structures are defined in [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go).

### What types of objects can be stored besides compressed payloads?

In addition to compression recovery entries, CCR stores typed working‑memory objects such as `FileObservation`, `SearchResult`, and `TestResult`. These objects include provenance fields (`session_id`, `source`, `repository_state`), lifecycle status flags (`hot`, `warm`, `cold`), and comprehensive size accounting metrics (`original_byte_length`, `stored_byte_length`).