How Claude-Obsidian Determines Independence for High-Risk Claims
Claude-Obsidian validates high-risk claims by requiring at least two independent sources, determining independence through identity set analysis and union-find grouping in the ledger validation engine.
The AgriciDaniel/claude-obsidian repository enforces strict source verification for high-risk claims through a multi-layered independence check implemented in claude_obsidian/ledgers.py. When a claim is flagged as high-risk, the system mandates supporting evidence from distinct, non-overlapping origins to prevent circular verification. Understanding how claude-obsidian determines independence for high-risk claims requires examining the ledger validation logic and its sophisticated source deduplication algorithm.
The Two-Source Rule for High-Risk Claims
High-risk claims in claude-obsidian require a minimum of two independent supporting sources. This requirement is enforced during ledger validation to ensure critical statements cannot be verified through a single point of failure or mutually dependent references. According to the CLAIM_SCHEMA defined in claude_obsidian/contracts.py, claims marked with "high_risk": true trigger this stricter validation path during the validate_claim_ledger execution.
How Independence Is Calculated in claude_obsidian/ledgers.py
The independence calculation occurs during the validation pipeline, specifically within the validate_claim_ledger function. The system constructs identity fingerprints for each source and applies union-find algorithms to detect overlaps and group related sources together.
Building Identity Sets (Lines 1265–1282)
For every source supporting a claim, claude-obsidian constructs a set of identifiers representing its origin. This logic appears in the loop starting at line 1265 of claude_obsidian/ledgers.py. Each source receives up to four canonical identifiers:
source:<source_id>— the unique source identifierorigin:<kind>:<canonical-locator>— normalized kind (url,file, etc.) with its canonical locatorcontent:<sha256>— the SHA-256 hash of the contentdeclared:<independence_key>— optional user-provided independence key after NFC-normalisation
Sources sharing any of these identifiers are considered related rather than independent.
Union-Find Grouping with _independent_group_count (Lines 997–1015)
The helper function _independent_group_count (lines 997–1015) implements union-find logic to merge sources sharing any identifier. If two sources intersect on origin, content hash, or declared independence key, they join the same group. The function returns the count of distinct groups, representing truly independent verification paths. Two sources from the same origin URL or with identical content hashes automatically merge into a single group, failing the independence requirement.
Validation Logic for High-Risk Claims (Line 1135)
When processing a claim marked with "high_risk": true, the validator checks the group count returned by _independent_group_count. At line 1135 in validate_claim_ledger, the system raises an error if fewer than two independent groups exist:
"high-risk acceptance requires two independent sources"
The claim only passes validation when at least two distinct groups provide support, ensuring no single source or related set of sources can unilaterally verify high-risk information.
Practical Implementation Examples
The following example demonstrates a valid high-risk claim with two independent sources:
from claude_obsidian.ledgers import validate_claim_ledger
# Example claim requiring two independent sources
claim = {
"claim_id": "c1",
"description": "Some critical statement",
"high_risk": True,
"supporting_sources": [
{"source_id": "src-1", "origin": {"kind": "url", "locator": "https://example.com/a"},
"content_sha256": "aaa…", "independence_key": None},
{"source_id": "src-2", "origin": {"kind": "url", "locator": "https://example.org/b"},
"content_sha256": "bbb…", "independence_key": None},
],
}
errors = validate_claim_ledger({"claims": {"c1": claim}}, vault_root=Path("/my/vault"))
print(errors) # → [] (claim accepted because two independent groups)
Conversely, this example fails validation because both sources share the same origin locator:
# Failing case: two sources sharing the same origin
claim["supporting_sources"][1]["origin"]["locator"] = "https://example.com/a"
errors = validate_claim_ledger({"claims": {"c1": claim}}, vault_root=Path("/my/vault"))
print(errors)
# → [{'path': 'c1', 'message': 'high-risk acceptance requires two independent sources'}]
Summary
- High-risk claims require two independent source groups as enforced in
claude_obsidian/ledgers.py - Independence is determined by analyzing identity sets for overlapping origins, content hashes, or declared keys
- The
_independent_group_countfunction uses union-find logic (lines 997–1015) to merge related sources - Validation fails at line 1135 when fewer than two independent groups support a high-risk claim
- Source identity construction occurs at lines 1265–1282 using four canonical identifier types
Frequently Asked Questions
What constitutes an independent source in claude-obsidian?
A source is considered independent when it belongs to a distinct identity group with no overlapping identifiers—origin, content hash, or declared independence key—with other sources supporting the same claim. The union-find algorithm in _independent_group_count automatically merges sources sharing any of these attributes into the same validation group.
Where is the high-risk claim validation logic implemented?
The core validation resides in claude_obsidian/ledgers.py, specifically in the validate_claim_ledger function at line 1135, with group calculation handled by _independent_group_count at lines 997–1015. The schema definition including the high_risk boolean flag exists in claude_obsidian/contracts.py.
Can two URLs from the same domain be considered independent?
No. If both URLs share the same canonical locator in their origin metadata (e.g., both point to https://example.com/a), the union-find algorithm merges them into the same group, making them non-independent for high-risk validation. They must differ in either the locator, content hash, or declared independence key to pass as separate groups.
How does the independence_key parameter work?
Users can optionally declare an independence_key in source metadata (NFC-normalized) to explicitly link or separate sources. Sources sharing the same declared key are grouped together as non-independent, while distinct keys preserve separation even if origins differ. This allows manual override of automatic independence detection for complex verification scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →