# How Claude-Obsidian Ensures Provenance Awareness in Notes: A Technical Deep Dive

> Discover how Claude Obsidian ensures provenance awareness with immutable ledgers, stable identifiers, and transaction validation for verifiable evidence. Learn more.

- Repository: [Agrici.Daniel/claude-obsidian](https://github.com/AgriciDaniel/claude-obsidian)
- Tags: deep-dive
- Published: 2026-08-25

---

**Claude-Obsidian ensures provenance awareness through immutable source and claim ledgers, stable content-based identifiers, and transaction-level validation that binds every assertion to cryptographically verifiable evidence.**

In the AgriciDaniel/claude-obsidian repository, provenance awareness is architected as a constraint system rather than optional metadata. The implementation uses dual ledgers living inside each vault to create a provable chain of custody from raw evidence to written claims, enforced by cryptographic hashing and strict schema validation.

## Dual-Ledger Architecture for Immutable Provenance

The provenance system rests on two canonical JSON schemas defined in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) that separate evidence from assertions.

**Source Ledger** records the evidentiary basis for knowledge. It uses the schema constant `SOURCE_SCHEMA = "claude-obsidian.source-ledger.v1"` at lines 21–25【source-ledger schema】 to validate entries describing files, URLs, or manual inputs. Each source entry includes `origin.kind` (file, url, or manual), a canonicalized locator, content SHA-256 hashes, and temporal metadata.

**Claim Ledger** captures assertions made in your notes. Using `CLAIM_SCHEMA = "claude-obsidian.claim-ledger.v1"`【source-ledger schema】, this ledger stores the claim text, risk level, confidence assessment, and—critically—a list of `evidence` objects linking back to source IDs. This separation ensures that claims cannot exist without explicit pointers to underlying evidence.

## Stable Content-Based Source Identification

To guarantee referential integrity across vaults and time, Claude-Obsidian generates deterministic source identifiers based on content rather than random assignment. The `stable_source_id` function in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) (lines 707–714)【stable-source-id】 implements this as follows:

```python
def stable_source_id(origin_kind, locator, content_sha256):
    normalized_locator = _canonical_locator(origin_kind, locator)
    digest = hashlib.sha256(
        f"{origin_kind.casefold()}\0{normalized_locator}\0{(content_sha256 or '').casefold()}".encode()
    ).hexdigest()
    return f"src-{digest[:20]}"

```

By hashing the origin type, normalized locator, and optional content SHA-256, the same external fact always receives the identical `src-` ID. This prevents link rot and allows cross-vault synchronization of provenance records without UUID collisions.

## Multi-Stage Validation Pipeline

Before any ledger modification is committed, Claude-Obsidian validates provenance through two strict stages defined in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py).

### Source-Ledger Validation

The `validate_source_ledger` function (lines 67–78)【source-validation】 enforces that every entry contains:

- A valid `origin.kind` from the allowed set (`file`, `url`, `manual`)
- A well-formed locator appropriate to the kind
- Proper 64-character SHA-256 hashes when content is provided
- A computed source ID that matches the canonical `stable_source_id` calculation

### Claim-Ledger Validation

The `validate_claim_ledger` function (lines 1080–1120)【claim-validation】 performs the critical work of evidence verification. It confirms that every claim points to existing source IDs, that referenced wiki pages and anchors exist, and that evidence satisfies **freshness** and **independence** rules.

**Freshness checks** compare the source record's `retrieved_at`, `ingested_at`, and `refresh_due` dates against the user-supplied audit date (`as_of`), flagging stale sources that should not support new claims【staleness check】.

**Independence verification** uses the `_independent_group_count` routine (lines 97–104)【independence】 to group sources sharing URLs, content hashes, or declared `independence_key` values. The validator merges overlapping groups and counts distinct independent sources, enforcing the "two-independent-sources" rule for high-risk claims.

## Transaction-Level Enforcement and Guardrails

The [`claude_obsidian/transaction.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/transaction.py) module provides the final safety layer. Before applying any write bundle, it invokes `_validate_provenance_writes` (lines 3510–3520)【transaction guard】 to reject transactions that would introduce:

- Duplicate JSON keys within ledger objects
- Malformed provenance structures violating schema constraints
- Mismatched IDs between claims and their cited sources

This enforces single-source-of-truth semantics, ensuring that once a source or claim is committed, its provenance record remains immutable and internally consistent.

## Practical Implementation Examples

Creating a provenance-aware note requires populating both ledgers. First, define a source entry and compute its stable ID:

```python
from claude_obsidian.ledgers import stable_source_id

source = {
    "origin": {"kind": "file", "locator": "docs/research.pdf"},
    "content_kind": "document",
    "title": "Research PDF",
    "authority": "official",
    "review_status": "active",
    "content_sha256": "a3f5…e9c2",   # 64-char SHA-256

    "retrieved_at": "2024-05-01",
    "refresh_due": "2025-05-01",
    "pages": ["wiki/research_summary.md"]
}

# Compute stable ID (identical across all vaults)

src_id = stable_source_id("file", "docs/research.pdf", "a3f5…e9c2")
print(src_id)   # → src-1b2c3d4e5f6a7b8c9d0e

```

Next, record a claim that depends on this source, then validate before committing:

```python
from claude_obsidian.ledgers import validate_claim_ledger
from datetime import date

claim = {
    "text": "The algorithm achieves 95% accuracy.",
    "risk": "high",
    "assessment": "accepted",
    "confidence": "high",
    "location": {"path": "wiki/algorithm.md"},
    "evidence": [
        {"source_id": src_id, "relation": "supports"}
    ],
    "reviewed_at": "2024-06-15"
}

errors = validate_claim_ledger(
    {"schema": "claude-obsidian.claim-ledger.v1", "generated_at": "2024-06-16T12:00:00Z", "claims": {"clm-001": claim}},
    source_ledger={"schema": "claude-obsidian.source-ledger.v1", "generated_at": "2024-06-10T09:30:00Z", "sources": {src_id: source}},
    as_of=date(2024, 6, 20)
)

assert not errors, errors

```

To audit provenance across your entire vault, run the built-in linter:

```bash
$ python -m claude_obsidian lint --as-of 2024-06-20

# → prints any provenance_errors, duplicate JSON keys, or stale sources

```

## Summary

Claude-Obsidian implements provenance awareness through a tightly integrated system that:

- **Separates evidence from claims** using immutable dual ledgers ([`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py))
- **Generates stable identifiers** via content-addressed hashing in `stable_source_id`
- **Validates integrity** through `validate_source_ledger` and `validate_claim_ledger`, enforcing freshness and independence rules
- **Guards transactions** via `_validate_provenance_writes` in [`claude_obsidian/transaction.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/transaction.py) to block malformed writes
- **Maintains backward compatibility** by migrating legacy manifests while preserving provenance IDs

## Frequently Asked Questions

### How does Claude-Obsidian prevent stale evidence from supporting new claims?

The `validate_claim_ledger` function compares each source's `refresh_due` date against the user-supplied `as_of` audit date. Sources past their refresh deadline trigger validation errors, forcing users to update or re-verify evidence before it can support new assertions【staleness check】.

### Can provenance records be shared across different vaults?

Yes. Because `stable_source_id` generates deterministic identifiers based on content hashes and canonical locators rather than random UUIDs, the same source document produces the identical `src-` ID regardless of which vault ingests it. This enables cross-vault provenance verification and federated knowledge bases.

### What prevents duplicate evidence from artificially inflating claim confidence?

The `_independent_group_count` routine clusters sources by shared URLs, content SHA-256, or explicit `independence_key` values. When validating high-risk claims, the system counts only distinct groups, ensuring that the same PDF hosted on two different mirrors does not count as two independent sources【independence】.

### Where does the system enforce schema compliance during edits?

Schema enforcement happens at the transaction layer in [`claude_obsidian/transaction.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/transaction.py). The `_validate_provenance_writes` method inspects every proposed ledger modification before commit, rejecting writes with duplicate JSON keys, schema version mismatches, or orphaned source references【transaction guard】.