# How Claude-Obsidian Determines Independence for High-Risk Claims

> Discover how Claude-Obsidian validates high-risk claims using identity set analysis and union-find grouping for ledger validation. Learn about determining independence for critical data.

- Repository: [Agrici.Daniel/claude-obsidian](https://github.com/AgriciDaniel/claude-obsidian)
- Tags: how-to-guide
- Published: 2026-08-29

---

**Claude-Obsidian validates high-risk claims by requiring at least two independent sources, determining independence through identity set analysis and union-find grouping in the ledger validation engine.**

The AgriciDaniel/claude-obsidian repository enforces strict source verification for high-risk claims through a multi-layered independence check implemented in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py). When a claim is flagged as high-risk, the system mandates supporting evidence from distinct, non-overlapping origins to prevent circular verification. Understanding how claude-obsidian determines independence for high-risk claims requires examining the ledger validation logic and its sophisticated source deduplication algorithm.

## The Two-Source Rule for High-Risk Claims

High-risk claims in claude-obsidian require a minimum of **two independent supporting sources**. This requirement is enforced during ledger validation to ensure critical statements cannot be verified through a single point of failure or mutually dependent references. According to the `CLAIM_SCHEMA` defined in [`claude_obsidian/contracts.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/contracts.py), claims marked with `"high_risk": true` trigger this stricter validation path during the `validate_claim_ledger` execution.

## How Independence Is Calculated in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py)

The independence calculation occurs during the validation pipeline, specifically within the `validate_claim_ledger` function. The system constructs identity fingerprints for each source and applies **union-find algorithms** to detect overlaps and group related sources together.

### Building Identity Sets (Lines 1265–1282)

For every source supporting a claim, claude-obsidian constructs a set of identifiers representing its origin. This logic appears in the loop starting at **line 1265** of [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py). Each source receives up to four canonical identifiers:

- `source:<source_id>` — the unique source identifier
- `origin:<kind>:<canonical-locator>` — normalized kind (`url`, `file`, etc.) with its canonical locator
- `content:<sha256>` — the SHA-256 hash of the content
- `declared:<independence_key>` — optional user-provided independence key after **NFC-normalisation**

Sources sharing any of these identifiers are considered related rather than independent.

### Union-Find Grouping with `_independent_group_count` (Lines 997–1015)

The helper function **`_independent_group_count`** (lines 997–1015) implements union-find logic to merge sources sharing any identifier. If two sources intersect on origin, content hash, or declared independence key, they join the same group. The function returns the count of distinct groups, representing truly independent verification paths. Two sources from the same origin URL or with identical content hashes automatically merge into a single group, failing the independence requirement.

## Validation Logic for High-Risk Claims (Line 1135)

When processing a claim marked with `"high_risk": true`, the validator checks the group count returned by `_independent_group_count`. At **line 1135** in `validate_claim_ledger`, the system raises an error if fewer than two independent groups exist:

```python
"high-risk acceptance requires two independent sources"

```

The claim only passes validation when **at least two distinct groups** provide support, ensuring no single source or related set of sources can unilaterally verify high-risk information.

## Practical Implementation Examples

The following example demonstrates a valid high-risk claim with two independent sources:

```python
from claude_obsidian.ledgers import validate_claim_ledger

# Example claim requiring two independent sources

claim = {
    "claim_id": "c1",
    "description": "Some critical statement",
    "high_risk": True,
    "supporting_sources": [
        {"source_id": "src-1", "origin": {"kind": "url", "locator": "https://example.com/a"},
         "content_sha256": "aaa…", "independence_key": None},
        {"source_id": "src-2", "origin": {"kind": "url", "locator": "https://example.org/b"},
         "content_sha256": "bbb…", "independence_key": None},
    ],
}
errors = validate_claim_ledger({"claims": {"c1": claim}}, vault_root=Path("/my/vault"))
print(errors)          # → []  (claim accepted because two independent groups)

```

Conversely, this example fails validation because both sources share the same origin locator:

```python

# Failing case: two sources sharing the same origin

claim["supporting_sources"][1]["origin"]["locator"] = "https://example.com/a"
errors = validate_claim_ledger({"claims": {"c1": claim}}, vault_root=Path("/my/vault"))
print(errors)

# → [{'path': 'c1', 'message': 'high-risk acceptance requires two independent sources'}]

```

## Summary

- High-risk claims require **two independent source groups** as enforced in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py)
- Independence is determined by analyzing identity sets for overlapping **origins, content hashes, or declared keys**
- The `_independent_group_count` function uses union-find logic (**lines 997–1015**) to merge related sources
- Validation fails at **line 1135** when fewer than two independent groups support a high-risk claim
- Source identity construction occurs at **lines 1265–1282** using four canonical identifier types

## Frequently Asked Questions

### What constitutes an independent source in claude-obsidian?

A source is considered independent when it belongs to a distinct identity group with no overlapping identifiers—origin, content hash, or declared independence key—with other sources supporting the same claim. The union-find algorithm in `_independent_group_count` automatically merges sources sharing any of these attributes into the same validation group.

### Where is the high-risk claim validation logic implemented?

The core validation resides in **[`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py)**, specifically in the `validate_claim_ledger` function at **line 1135**, with group calculation handled by `_independent_group_count` at **lines 997–1015**. The schema definition including the `high_risk` boolean flag exists in [`claude_obsidian/contracts.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/contracts.py).

### Can two URLs from the same domain be considered independent?

No. If both URLs share the same **canonical locator** in their origin metadata (e.g., both point to `https://example.com/a`), the union-find algorithm merges them into the same group, making them non-independent for high-risk validation. They must differ in either the locator, content hash, or declared independence key to pass as separate groups.

### How does the `independence_key` parameter work?

Users can optionally declare an `independence_key` in source metadata (**NFC-normalized**) to explicitly link or separate sources. Sources sharing the same declared key are grouped together as non-independent, while distinct keys preserve separation even if origins differ. This allows manual override of automatic independence detection for complex verification scenarios.