How Claude-Obsidian Implements Content-Addressed Storage for Captured Sources
Claude-Obsidian uses SHA-256 hashes to generate deterministic source IDs, storing every captured source in a content-addressed ledger that enables deduplication, stable referencing, and integrity verification across the knowledge graph.
Claude-Obsidian, an open-source knowledge management system by AgriciDaniel, treats every captured source as immutable content identified by its cryptographic hash. According to the source code, the system maintains a source ledger at wiki/meta/ledgers/source-ledger.json that maps each source to a unique identifier derived from the SHA-256 hash of its raw bytes. This content-addressed storage architecture ensures that identical content always resolves to the same identifier, regardless of its original path or ingestion time.
The Source Ledger Architecture
Each entry in the source ledger stores an origin object describing the source type (file, url, or manual) alongside a content_sha256 field containing the hash of the source's raw bytes. This combination creates a tamper-evident record where the source ID is cryptographically bound to the content itself.
Ledger Schema and Location
The system persists the ledger at wiki/meta/ledgers/source-ledger.json within the vault. Each source record follows a strict schema that includes the origin kind, canonical locator, content hash, and metadata such as review_status and authority.
Generating Content-Addressed Source IDs
The core of the addressing scheme is the stable_source_id function implemented in claude_obsidian/ledgers.py. This function constructs a deterministic identifier from three components: the origin kind, a canonical locator, and the content SHA-256 hash.
def stable_source_id(origin_kind: str, locator: str, content_sha256: str | None) -> str:
normalized_locator = _canonical_locator(origin_kind, locator)
digest = hashlib.sha256(
f"{origin_kind.casefold()}\0{normalized_locator}\0{(content_sha256 or '').casefold()}"
.encode("utf-8", errors="surrogatepass")
).hexdigest()
return f"src-{digest[:20]}"
Source: ledgers.py:L1078‑L1085
The function normalizes the locator using _canonical_locator, then computes a SHA-256 hash of the concatenated origin kind, normalized locator, and content hash. It returns a truncated 20-character hexadecimal digest prefixed with src-, producing identifiers like src-1b2c3d4e5f6a7b8c9d0e.
Canonical Locator Normalization
Before hashing, the system normalizes locators to ensure consistency. File paths are converted to POSIX format using PurePosixPath, while URLs undergo _canonical_url processing to standardize percent-encoding and case. This normalization prevents variations in spelling or encoding from generating different IDs for the same logical source.
Ingestion and Hash Computation
When ingesting a file source, the system computes the content hash using _safe_hash (with path safety enforced by claude_obsidian/paths.py), then generates the stable ID:
digest = _safe_hash(root, locator) # Compute SHA-256 of file contents
origin_kind = "file" if digest is not None else "manual"
source_id = stable_source_id(origin_kind, locator, digest)
Source: ledgers.py:L1244‑L1250
The _safe_hash function safely reads the file contents within the vault boundary and returns the hexadecimal SHA-256 digest. This digest is stored in the content_sha256 field of the ledger entry. Additionally, claude_obsidian/transaction.py handles bundling of writes and records the content_sha256 for each write in transaction metadata.
Practical Example: Adding a File Source
from pathlib import Path
from claude_obsidian.ledgers import stable_source_id, _safe_hash
vault_root = Path("/path/to/vault")
locator = "wiki/raw/article.md"
# Compute hash of the file contents
content_hash = _safe_hash(vault_root, locator) # e.g. "a3f5…"
# Build a stable, content-addressed source ID
source_id = stable_source_id("file", locator, content_hash) # e.g. "src-1b2c3d4e5f6a7b8c9d0e"
# Example ledger entry
source_entry = {
"origin": {"kind": "file", "locator": locator},
"content_kind": "document",
"title": Path(locator).stem,
"authority": "unknown",
"content_sha256": content_hash,
"review_status": "active",
"pages": ["wiki/article.md"],
}
Integrity Verification
The validate_source_ledger function ensures ledger integrity by comparing stored hashes against current file bytes. If a source file has been modified since ingestion, the validation catches the discrepancy:
if actual_hash != content_hash.lower():
_error(errors, f"{prefix}.content_sha256", "does not match current file bytes")
Source: ledgers.py:L1295‑L1300
This verification prevents accidental drift between the ledger and the actual source content, prompting users to re-ingest modified files.
Validation Example
from claude_obsidian.ledgers import validate_source_ledger
ledger = {"schema": "claude-obsidian.source-ledger.v1", "sources": {source_id: source_entry}}
errors = validate_source_ledger(ledger, vault_root=vault_root)
assert not errors # Passes only if content hash matches the file on disk
Benefits of Content-Addressed Storage
Claude-Obsidian's content-addressed model provides three critical capabilities for knowledge management:
- Deduplication: Identical content from different paths generates the same
src-ID, avoiding redundant evidence storage in the ledger. - Stable Referencing: Claims and evidence link to sources via immutable content hashes, ensuring links remain valid as long as the content remains unchanged.
- Integrity Verification: Cryptographic hashing detects any modification to underlying files, maintaining the evidentiary chain and preventing silent data corruption.
Summary
- Claude-Obsidian stores captured sources in
wiki/meta/ledgers/source-ledger.jsonwith SHA-256 content hashes. - The
stable_source_idfunction inclaude_obsidian/ledgers.pygenerates deterministic IDs by hashing the origin kind, canonical locator, and content hash. - The system normalizes locators using
_canonical_locatorto ensure consistent ID generation regardless of path variations. validate_source_ledgerenforces integrity by verifying that storedcontent_sha256values match current file contents.- Content-addressed storage enables deduplication, stable cross-referencing, and tamper detection across the knowledge graph.
Frequently Asked Questions
How does content-addressed storage prevent duplicate sources?
When identical content is ingested from different locations, the SHA-256 hash remains constant. Because the stable_source_id function incorporates this hash into the source identifier, both ingestion attempts generate the same src- ID. The ledger treats these as the same source, preventing duplicate entries while still tracking multiple origin locators if needed.
What happens if a source file changes after ingestion?
The validate_source_ledger function detects hash mismatches between the stored content_sha256 and the current file bytes. When validation fails with an error indicating the content "does not match current file bytes," users must explicitly re-ingest the modified source to update the ledger, ensuring the knowledge graph maintains accurate provenance.
How is the source ID format structured?
Source IDs follow the format src-{digest} where the digest consists of the first 20 characters of a SHA-256 hash. This hash is computed from the case-folded origin kind, a null byte separator, the canonical locator, another null byte separator, and the lowercase content hash or empty string.
Where is the source ledger stored in the vault?
The source ledger resides at wiki/meta/ledgers/source-ledger.json relative to the vault root. This location is reserved for metadata about captured sources, with each entry containing the origin details, content SHA-256 hash, and review status necessary for content-addressed storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →