# How Claude-Obsidian Implements Content-Addressed Storage for Captured Sources

> Discover how Claude-Obsidian leverages content-addressed storage with SHA-256 hashes for efficient knowledge graph management. Ensure deduplication, stable referencing, and integrity.

- Repository: [Agrici.Daniel/claude-obsidian](https://github.com/AgriciDaniel/claude-obsidian)
- Tags: internals
- Published: 2026-08-26

---

**Claude-Obsidian uses SHA-256 hashes to generate deterministic source IDs, storing every captured source in a content-addressed ledger that enables deduplication, stable referencing, and integrity verification across the knowledge graph.**

Claude-Obsidian, an open-source knowledge management system by AgriciDaniel, treats every captured source as immutable content identified by its cryptographic hash. According to the source code, the system maintains a **source ledger** at [`wiki/meta/ledgers/source-ledger.json`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/wiki/meta/ledgers/source-ledger.json) that maps each source to a unique identifier derived from the SHA-256 hash of its raw bytes. This content-addressed storage architecture ensures that identical content always resolves to the same identifier, regardless of its original path or ingestion time.

## The Source Ledger Architecture

Each entry in the source ledger stores an `origin` object describing the source type (`file`, `url`, or `manual`) alongside a `content_sha256` field containing the hash of the source's raw bytes. This combination creates a tamper-evident record where the source ID is cryptographically bound to the content itself.

### Ledger Schema and Location

The system persists the ledger at [`wiki/meta/ledgers/source-ledger.json`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/wiki/meta/ledgers/source-ledger.json) within the vault. Each source record follows a strict schema that includes the origin kind, canonical locator, content hash, and metadata such as `review_status` and `authority`.

## Generating Content-Addressed Source IDs

The core of the addressing scheme is the `stable_source_id` function implemented in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py). This function constructs a deterministic identifier from three components: the origin kind, a canonical locator, and the content SHA-256 hash.

```python
def stable_source_id(origin_kind: str, locator: str, content_sha256: str | None) -> str:
    normalized_locator = _canonical_locator(origin_kind, locator)
    digest = hashlib.sha256(
        f"{origin_kind.casefold()}\0{normalized_locator}\0{(content_sha256 or '').casefold()}"
        .encode("utf-8", errors="surrogatepass")
    ).hexdigest()
    return f"src-{digest[:20]}"

```

*Source: [`ledgers.py:L1078‑L1085`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py#L1078-L1085)*

The function normalizes the locator using `_canonical_locator`, then computes a SHA-256 hash of the concatenated origin kind, normalized locator, and content hash. It returns a truncated 20-character hexadecimal digest prefixed with `src-`, producing identifiers like `src-1b2c3d4e5f6a7b8c9d0e`.

### Canonical Locator Normalization

Before hashing, the system normalizes locators to ensure consistency. File paths are converted to POSIX format using `PurePosixPath`, while URLs undergo `_canonical_url` processing to standardize percent-encoding and case. This normalization prevents variations in spelling or encoding from generating different IDs for the same logical source.

## Ingestion and Hash Computation

When ingesting a file source, the system computes the content hash using `_safe_hash` (with path safety enforced by [`claude_obsidian/paths.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/paths.py)), then generates the stable ID:

```python
digest = _safe_hash(root, locator)          # Compute SHA-256 of file contents

origin_kind = "file" if digest is not None else "manual"
source_id = stable_source_id(origin_kind, locator, digest)

```

*Source: [`ledgers.py:L1244‑L1250`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py#L1244-L1250)*

The `_safe_hash` function safely reads the file contents within the vault boundary and returns the hexadecimal SHA-256 digest. This digest is stored in the `content_sha256` field of the ledger entry. Additionally, [`claude_obsidian/transaction.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/transaction.py) handles bundling of writes and records the `content_sha256` for each write in transaction metadata.

### Practical Example: Adding a File Source

```python
from pathlib import Path
from claude_obsidian.ledgers import stable_source_id, _safe_hash

vault_root = Path("/path/to/vault")
locator = "wiki/raw/article.md"

# Compute hash of the file contents

content_hash = _safe_hash(vault_root, locator)          # e.g. "a3f5…"

# Build a stable, content-addressed source ID

source_id = stable_source_id("file", locator, content_hash)  # e.g. "src-1b2c3d4e5f6a7b8c9d0e"

# Example ledger entry

source_entry = {
    "origin": {"kind": "file", "locator": locator},
    "content_kind": "document",
    "title": Path(locator).stem,
    "authority": "unknown",
    "content_sha256": content_hash,
    "review_status": "active",
    "pages": ["wiki/article.md"],
}

```

## Integrity Verification

The `validate_source_ledger` function ensures ledger integrity by comparing stored hashes against current file bytes. If a source file has been modified since ingestion, the validation catches the discrepancy:

```python
if actual_hash != content_hash.lower():
    _error(errors, f"{prefix}.content_sha256", "does not match current file bytes")

```

*Source: [`ledgers.py:L1295‑L1300`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py#L1295-L1300)*

This verification prevents accidental drift between the ledger and the actual source content, prompting users to re-ingest modified files.

### Validation Example

```python
from claude_obsidian.ledgers import validate_source_ledger

ledger = {"schema": "claude-obsidian.source-ledger.v1", "sources": {source_id: source_entry}}
errors = validate_source_ledger(ledger, vault_root=vault_root)
assert not errors  # Passes only if content hash matches the file on disk

```

## Benefits of Content-Addressed Storage

Claude-Obsidian's content-addressed model provides three critical capabilities for knowledge management:

- **Deduplication**: Identical content from different paths generates the same `src-` ID, avoiding redundant evidence storage in the ledger.
- **Stable Referencing**: Claims and evidence link to sources via immutable content hashes, ensuring links remain valid as long as the content remains unchanged.
- **Integrity Verification**: Cryptographic hashing detects any modification to underlying files, maintaining the evidentiary chain and preventing silent data corruption.

## Summary

- Claude-Obsidian stores captured sources in [`wiki/meta/ledgers/source-ledger.json`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/wiki/meta/ledgers/source-ledger.json) with SHA-256 content hashes.
- The `stable_source_id` function in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) generates deterministic IDs by hashing the origin kind, canonical locator, and content hash.
- The system normalizes locators using `_canonical_locator` to ensure consistent ID generation regardless of path variations.
- `validate_source_ledger` enforces integrity by verifying that stored `content_sha256` values match current file contents.
- Content-addressed storage enables deduplication, stable cross-referencing, and tamper detection across the knowledge graph.

## Frequently Asked Questions

### How does content-addressed storage prevent duplicate sources?

When identical content is ingested from different locations, the SHA-256 hash remains constant. Because the `stable_source_id` function incorporates this hash into the source identifier, both ingestion attempts generate the same `src-` ID. The ledger treats these as the same source, preventing duplicate entries while still tracking multiple origin locators if needed.

### What happens if a source file changes after ingestion?

The `validate_source_ledger` function detects hash mismatches between the stored `content_sha256` and the current file bytes. When validation fails with an error indicating the content "does not match current file bytes," users must explicitly re-ingest the modified source to update the ledger, ensuring the knowledge graph maintains accurate provenance.

### How is the source ID format structured?

Source IDs follow the format `src-{digest}` where the digest consists of the first 20 characters of a SHA-256 hash. This hash is computed from the case-folded origin kind, a null byte separator, the canonical locator, another null byte separator, and the lowercase content hash or empty string.

### Where is the source ledger stored in the vault?

The source ledger resides at [`wiki/meta/ledgers/source-ledger.json`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/wiki/meta/ledgers/source-ledger.json) relative to the vault root. This location is reserved for metadata about captured sources, with each entry containing the origin details, content SHA-256 hash, and review status necessary for content-addressed storage.