How Caveman Generates Context-Addressed Handles: SHA-256 Content Hashing Explained

Caveman generates deterministic context-addressed handles by computing a SHA-256 hash of the raw payload bytes and prefixing it with a namespace identifier, ensuring idempotent storage and byte-exact retrieval.

The JuliusBrussee/caveman repository implements a deterministic context recovery system where every piece of stored data receives a unique, reproducible identifier. Understanding how Caveman context-addressed handles are generated is essential for developers integrating with the Caveman Context Recovery (CCR) engine, as these handles enable duplicate-free storage and reliable data retrieval across sessions.

SHA-256 Foundation of Context-Addressed Handles

Caveman uses SHA-256 cryptographic hashing to create deterministic identifiers. When a payload is stored—whether compressed tool output or file observation—the engine computes the hash directly from the raw bytes.

According to the source code in engine/ccr/store.go (lines 6-8), the implementation calculates the content hash using Go's standard library:

// The handle is content-addressed (sha256 of the original), which makes the
// engine idempotent — compressing the same payload twice yields the same handle
// and stores it once.
sum := sha256.Sum256(payload)
contentHash := "sha256:" + hex.EncodeToString(sum[:])

The resulting content hash follows the format sha256:<hex-digest> and serves as the canonical identifier for the payload bytes.

Handle Composition and Object ID Generation

While the raw content hash identifies the bytes, Caveman wraps this in a context-addressed handle with additional metadata. The system generates a unique Object ID by combining the content hash with session metadata.

In engine/ccr/store.go (lines 141-145), the code constructs the Object ID as follows:

if obj.ID == "" {
    identity := strings.Join([]string{string(obj.Type), obj.SessionID, obj.Source, obj.RepositoryState, obj.ContentHash}, "\x00")
    sum := sha256.Sum256([]byte(identity))
    obj.ID = "ccr_obj_" + hex.EncodeToString(sum[:16])
}

This process:

  1. Joins the Object Type, SessionID, Source, RepositoryState, and ContentHash using null byte separators (\x00)
  2. Computes a secondary SHA-256 hash of this identity string
  3. Prefixes the first 16 bytes with ccr_obj_ to create the final handle

Idempotent Storage Guarantees

The deterministic nature of Caveman context-addressed handles ensures idempotent storage. Because the handle derives purely from the content bytes, storing identical data twice produces the same handle without creating duplicate entries.

This design eliminates redundant storage in the CCR engine. When the PutObject method encounters an existing handle, it recognizes the content already exists and avoids writing duplicate bytes to the store.

Byte-Exact Retrieval by Handle

Retrieval operates through the Get(handle) method implemented in engine/ccr/store.go. When provided with a context-addressed handle, the store performs a lookup using the embedded SHA-256 hash and returns the exact byte payload used during storage.

This guarantees byte-exact round-tripping—the retrieved data matches the originally stored bytes exactly, with no transformation or compression artifacts affecting the content address.

Practical Implementation Examples

Storing a Payload and Obtaining Its Handle

To store data and receive its context-addressed handle, use the PrepareObject and PutObject functions:

import (
    "github.com/JuliusBrussee/caveman/engine/ccr"
)

func storeExample() (string, error) {
    // Raw bytes to store
    data := []byte(`{"type":"FileObservation","content":"example"}`)
    
    // Construct the object with metadata
    obj := ccr.Object{
        Type:            ccr.ObjectFileObservation,
        SessionID:       "sess_123",
        Source:          "my-plugin",
        RepositoryState: "main",
        Data:            data,
    }
    
    // Prepare computes the content hash
    prepared, err := ccr.PrepareObject(obj)
    if err != nil {
        return "", err
    }
    
    // Store returns the deterministic handle
    handle, err := ccrStore.PutObject(prepared)
    if err != nil {
        return "", err
    }
    // Returns: "sha256:ab12cd34..." or similar
    return handle, nil
}

Retrieving Original Bytes

Use the handle to fetch the exact payload:

func retrieveExample(handle string) ([]byte, error) {
    payload, err := ccrStore.Get(handle)
    if err != nil {
        return nil, err // Returns ccr.ErrNotFound if handle unknown
    }
    return payload, nil
}

Verifying Idempotency

Demonstrating that identical inputs produce identical handles:

h1, _ := storeExample()
h2, _ := storeExample() // Same data, same metadata
fmt.Println(h1 == h2)   // true - duplicate storage yields same handle

Summary

  • Caveman context-addressed handles are deterministic SHA-256 hashes of raw payload bytes, formatted as sha256:<hex-digest>.
  • Object IDs combine content hashes with session metadata (Type, SessionID, Source, RepositoryState) joined by null bytes, prefixed with ccr_obj_.
  • The implementation resides in engine/ccr/store.go, specifically in the prepareObject logic and PutObject methods.
  • Idempotent storage prevents duplicate entries—storing the same bytes twice yields the same handle without data duplication.
  • Byte-exact retrieval guarantees that the Get(handle) method returns the identical bytes originally stored.

Frequently Asked Questions

What hash algorithm does Caveman use for context-addressed handles?

Caveman uses SHA-256 (Secure Hash Algorithm 256-bit) to generate context-addressed handles. The implementation calls sha256.Sum256 from Go's standard library to compute the hash of raw payload bytes and metadata identity strings.

Why does Caveman use content-addressed storage instead of UUIDs?

Content-addressed storage provides deterministic deduplication and data integrity verification. Unlike UUIDs, which are random, Caveman context-addressed handles derived from SHA-256 ensure that identical content always maps to the same identifier, eliminating redundant storage and enabling verification that retrieved data matches the expected content.

How does Caveman prevent duplicate context entries?

The CCR engine prevents duplicates through cryptographic content addressing. When storing data, the PutObject method computes the SHA-256 hash of the payload. If that handle already exists in the store, the engine recognizes the content as previously stored and skips writing duplicate bytes, making storage operations idempotent.

Where is the context-addressed handle generation logic located?

The core logic resides in engine/ccr/store.go in the JuliusBrussee/caveman repository. Lines 6-8 handle the initial content hash calculation, while lines 141-145 generate the final Object ID by hashing concatenated metadata fields with the content hash.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →