# How Caveman Generates Context-Addressed Handles: SHA-256 Content Hashing Explained

> Discover how Caveman generates context-addressed handles using SHA-256 content hashing for idempotent storage and byte-exact retrieval. Learn about deterministic handle generation.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-09-06

---

**Caveman generates deterministic context-addressed handles by computing a SHA-256 hash of the raw payload bytes and prefixing it with a namespace identifier, ensuring idempotent storage and byte-exact retrieval.**

The JuliusBrussee/caveman repository implements a deterministic context recovery system where every piece of stored data receives a unique, reproducible identifier. Understanding how Caveman context-addressed handles are generated is essential for developers integrating with the Caveman Context Recovery (CCR) engine, as these handles enable duplicate-free storage and reliable data retrieval across sessions.

## SHA-256 Foundation of Context-Addressed Handles

Caveman uses **SHA-256 cryptographic hashing** to create deterministic identifiers. When a payload is stored—whether compressed tool output or file observation—the engine computes the hash directly from the raw bytes.

According to the source code in [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go) (lines 6-8), the implementation calculates the content hash using Go's standard library:

```go
// The handle is content-addressed (sha256 of the original), which makes the
// engine idempotent — compressing the same payload twice yields the same handle
// and stores it once.
sum := sha256.Sum256(payload)
contentHash := "sha256:" + hex.EncodeToString(sum[:])

```

The resulting **content hash** follows the format `sha256:<hex-digest>` and serves as the canonical identifier for the payload bytes.

## Handle Composition and Object ID Generation

While the raw content hash identifies the bytes, Caveman wraps this in a **context-addressed handle** with additional metadata. The system generates a unique Object ID by combining the content hash with session metadata.

In [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go) (lines 141-145), the code constructs the Object ID as follows:

```go
if obj.ID == "" {
    identity := strings.Join([]string{string(obj.Type), obj.SessionID, obj.Source, obj.RepositoryState, obj.ContentHash}, "\x00")
    sum := sha256.Sum256([]byte(identity))
    obj.ID = "ccr_obj_" + hex.EncodeToString(sum[:16])
}

```

This process:

1. Joins the **Object Type**, **SessionID**, **Source**, **RepositoryState**, and **ContentHash** using null byte separators (`\x00`)
2. Computes a secondary SHA-256 hash of this identity string
3. Prefixes the first 16 bytes with `ccr_obj_` to create the final handle

## Idempotent Storage Guarantees

The deterministic nature of Caveman context-addressed handles ensures **idempotent storage**. Because the handle derives purely from the content bytes, storing identical data twice produces the same handle without creating duplicate entries.

This design eliminates redundant storage in the CCR engine. When the `PutObject` method encounters an existing handle, it recognizes the content already exists and avoids writing duplicate bytes to the store.

## Byte-Exact Retrieval by Handle

Retrieval operates through the `Get(handle)` method implemented in [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go). When provided with a context-addressed handle, the store performs a lookup using the embedded SHA-256 hash and returns the exact byte payload used during storage.

This guarantees **byte-exact round-tripping**—the retrieved data matches the originally stored bytes exactly, with no transformation or compression artifacts affecting the content address.

## Practical Implementation Examples

### Storing a Payload and Obtaining Its Handle

To store data and receive its context-addressed handle, use the `PrepareObject` and `PutObject` functions:

```go
import (
    "github.com/JuliusBrussee/caveman/engine/ccr"
)

func storeExample() (string, error) {
    // Raw bytes to store
    data := []byte(`{"type":"FileObservation","content":"example"}`)
    
    // Construct the object with metadata
    obj := ccr.Object{
        Type:            ccr.ObjectFileObservation,
        SessionID:       "sess_123",
        Source:          "my-plugin",
        RepositoryState: "main",
        Data:            data,
    }
    
    // Prepare computes the content hash
    prepared, err := ccr.PrepareObject(obj)
    if err != nil {
        return "", err
    }
    
    // Store returns the deterministic handle
    handle, err := ccrStore.PutObject(prepared)
    if err != nil {
        return "", err
    }
    // Returns: "sha256:ab12cd34..." or similar
    return handle, nil
}

```

### Retrieving Original Bytes

Use the handle to fetch the exact payload:

```go
func retrieveExample(handle string) ([]byte, error) {
    payload, err := ccrStore.Get(handle)
    if err != nil {
        return nil, err // Returns ccr.ErrNotFound if handle unknown
    }
    return payload, nil
}

```

### Verifying Idempotency

Demonstrating that identical inputs produce identical handles:

```go
h1, _ := storeExample()
h2, _ := storeExample() // Same data, same metadata
fmt.Println(h1 == h2)   // true - duplicate storage yields same handle

```

## Summary

- Caveman context-addressed handles are **deterministic SHA-256 hashes** of raw payload bytes, formatted as `sha256:<hex-digest>`.
- Object IDs combine content hashes with session metadata (Type, SessionID, Source, RepositoryState) joined by null bytes, prefixed with `ccr_obj_`.
- The implementation resides in [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go), specifically in the `prepareObject` logic and `PutObject` methods.
- **Idempotent storage** prevents duplicate entries—storing the same bytes twice yields the same handle without data duplication.
- **Byte-exact retrieval** guarantees that the `Get(handle)` method returns the identical bytes originally stored.

## Frequently Asked Questions

### What hash algorithm does Caveman use for context-addressed handles?

Caveman uses **SHA-256** (Secure Hash Algorithm 256-bit) to generate context-addressed handles. The implementation calls `sha256.Sum256` from Go's standard library to compute the hash of raw payload bytes and metadata identity strings.

### Why does Caveman use content-addressed storage instead of UUIDs?

Content-addressed storage provides **deterministic deduplication** and data integrity verification. Unlike UUIDs, which are random, Caveman context-addressed handles derived from SHA-256 ensure that identical content always maps to the same identifier, eliminating redundant storage and enabling verification that retrieved data matches the expected content.

### How does Caveman prevent duplicate context entries?

The CCR engine prevents duplicates through **cryptographic content addressing**. When storing data, the `PutObject` method computes the SHA-256 hash of the payload. If that handle already exists in the store, the engine recognizes the content as previously stored and skips writing duplicate bytes, making storage operations idempotent.

### Where is the context-addressed handle generation logic located?

The core logic resides in [`engine/ccr/store.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/ccr/store.go) in the JuliusBrussee/caveman repository. Lines 6-8 handle the initial content hash calculation, while lines 141-145 generate the final Object ID by hashing concatenated metadata fields with the content hash.