How Claude-Obsidian Performs Deterministic Ledger Validation for a Given Audit Date
Claude-Obsidian validates source and claim ledgers through a pure functional pipeline that produces identical results for any given as_of audit date, ensuring bit-for-bit reproducible verification without external state dependencies.
The AgriciDaniel/claude-obsidian repository implements a fully deterministic validation system for knowledge management ledgers. By treating the audit date as the sole variable input and applying strict canonicalization rules at every step, the codebase guarantees that validation outcomes remain stable across executions, environments, and time.
The Deterministic Validation Pipeline
The validation process follows a fixed nine-step pipeline defined in claude_obsidian/ledgers.py. Each step uses pure functions that depend exclusively on the ledger data and the supplied audit date.
1. Audit Date Resolution via _audit_date
The helper function _audit_date normalizes the caller-provided as_of parameter (defaulting to the current date when omitted) and guarantees a datetime.date object. This normalization ensures that all subsequent temporal comparisons operate against a consistent, timezone-aware baseline.
Located at lines 97-103 in claude_obsidian/ledgers.py, this function acts as the single source of truth for the temporal boundary used throughout validation.
2. Strict JSON Parsing with strict_json_loads
Both validate_source_ledger and validate_claim_ledger begin by parsing the JSON payload using strict_json_loads (lines 75-84). This utility rejects duplicate keys and non-finite numbers, ensuring a canonical representation that eliminates parser-specific ambiguity.
3. Ledger Timestamp Validation
The validator checks the generated_at field, which must conform to the UTC timestamp format YYYY-MM-DDTHH:MM:SSZ. If generated_at is later than the audit date, the validator raises an immediate error. This check appears at lines 80-89 for source ledgers and lines 17-23 for claim ledgers.
4. Temporal Consistency Enforcement
The system enforces distinct temporal constraints based on ledger type:
- Source records:
retrieved_atandingested_attimestamps must not exceed the audit date (lines 60-66) - Claim records: The
reviewed_attimestamp must precede the audit date, and the claim cannot cite sources with retrieval dates newer than the audit date (lines 30-41)
5. Deterministic Source ID Generation
The stable_source_id function (lines 107-115) builds a SHA-256 digest from three canonical components: the lower-cased origin kind, the canonicalized locator, and an optional content hash. The function returns a deterministic identifier formatted as src-<20-character-prefix>, ensuring that identical sources always resolve to identical IDs regardless of ingestion order or storage location.
6. Source ID Correspondence Verification
Each source record's id field must exactly equal the value returned by stable_source_id for that record's content (lines 118-124). This verification ties ledger entries to their canonical identities, preventing ID drift or manual tampering.
7. Fresh-Source Determination
A source qualifies as fresh only when four conditions are met simultaneously: the source is active, its retrieved_at/ingested_at timestamps are less than or equal to the audit date, its refresh_due date is greater than or equal to the observed date, and the source's own timestamp is present (lines 111-122). This logic operates purely on ledger data and the audit date, never querying external state.
8. Independence Counting for High-Risk Claims
For accepted claims with risk == "high", the validator applies _independent_group_count (lines 97-138) to count supporting sources. The algorithm merges sources sharing any of the following canonical attributes: source ID, origin, content hash, or declared independence key. Because this count depends only on immutable ledger data and the audit date, it produces deterministic results essential for high-stakes verification.
9. Deterministic Error Reporting
All validation errors accumulate as dictionaries containing path and message fields. The final return statement sorts these errors by path and message (lines 49-58), producing a reproducible ordering that enables deterministic testing and diff-based auditing.
Implementing Deterministic Validation in Python
The following example demonstrates how to invoke the validation functions with a specific audit date, ensuring reproducible results:
from datetime import date
from pathlib import Path
from claude_obsidian.ledgers import (
validate_source_ledger,
validate_claim_ledger,
stable_source_id,
)
# Load a source ledger (already parsed JSON)
source_ledger = {
"schema": "claude-obsidian.source-ledger.v1",
"generated_at": "2024-01-01T12:00:00Z",
"sources": {
"src-01a2b3c4d5e6f7g8h9i0": {
"origin": {"kind": "url", "locator": "https://example.com/data.json"},
"content_kind": "document",
"title": "Example Data",
"authority": "official",
"content_sha256": "e3b0c44298fc1c149afbf4c8996fb924...",
"retrieved_at": "2023-12-15",
"refresh_due": "2024-06-01",
"review_status": "active",
"pages": ["wiki/example.md"],
}
},
}
# Validate the source ledger as of 2024-03-01
errors = validate_source_ledger(
source_ledger,
as_of=date(2024, 3, 1), # deterministic audit date
vault_root=Path("/path/to/vault") # optional, for file-hash checks
)
print("Source-ledger errors:", errors) # → [] if everything passes
# Validate a claim ledger referencing the source above
claim_ledger = {
"schema": "claude-obsidian.claim-ledger.v1",
"generated_at": "2024-01-02T08:00:00Z",
"claims": {
"clm-01a2b3c4d5e6f7g8h9i0": {
"text": "The dataset is up-to-date.",
"risk": "high",
"assessment": "accepted",
"confidence": "high",
"location": {"path": "wiki/analysis.md"},
"reviewed_at": "2024-02-28",
"evidence": [
{
"source_id": "src-01a2b3c4d5e6f7g8h9i0",
"relation": "supports",
}
],
}
},
}
claim_errors = validate_claim_ledger(
claim_ledger,
source_ledger,
as_of=date(2024, 3, 1)
)
print("Claim-ledger errors:", claim_errors) # → [] if high-risk independence passes
Key implementation details:
- The same
as_ofdate propagates to both validators, guaranteeing identical temporal boundaries stable_source_idcreates reproducible identifiers from canonical URLs and content hashes- Validation requires no network access or system state, enabling offline reproducibility
Core Validation Files and Functions
The deterministic behavior relies on specific modules within the repository:
-
claude_obsidian/ledgers.py: Containsvalidate_source_ledger,validate_claim_ledger,_audit_date, andstable_source_id. This file implements the pure functional pipeline that ensures deterministic validation. -
claude_obsidian/paths.py: Providescanonical()and other path-normalization utilities used when validating file-based sources, ensuring consistent locator representations across operating systems. -
claude_obsidian/url_safety.py: Suppliesurl_credential_issuefor deterministic URL validation, rejecting unsafe URLs through canonical rules rather than heuristics. -
tests/test_ledgers.py: Houses exhaustive unit tests asserting deterministic behavior across edge cases including leap years, timezone boundaries, and hash collisions.
Summary
- Pure functional design: Every validation step uses pure functions with no side effects or external dependencies
- Canonical identifiers:
stable_source_idgenerates reproducible source IDs using SHA-256 digests of normalized content - Temporal boundaries: The
as_ofaudit date acts as the sole variable, with all timestamps evaluated against this fixed point - High-risk claim handling: Independence counting for high-risk claims uses deterministic merging algorithms based on immutable attributes
- Reproducible output: Sorted error lists ensure identical output ordering across runs
Frequently Asked Questions
What makes Claude-Obsidian's ledger validation deterministic?
The validation is deterministic because it uses pure functions that depend exclusively on two inputs: the ledger JSON data and the as_of audit date. The system contains no random number generation, no external API calls, and no reliance on system state such as current time (unless explicitly used as the default as_of value). Every operation—from JSON parsing to ID generation to error sorting—follows canonical algorithms that produce identical outputs for identical inputs.
How does stable_source_id ensure reproducible identifiers?
The stable_source_id function in claude_obsidian/ledgers.py constructs identifiers by computing a SHA-256 hash of three normalized components: the lower-cased origin kind (e.g., "url"), the canonicalized locator, and the optional content hash. By lower-casing strings and canonicalizing paths before hashing, the function eliminates case-sensitivity and path-formatting differences across platforms. The resulting 20-character prefix appended to src- creates a stable, content-addressed identifier.
What temporal checks prevent future-dated data from affecting validation?
The validator performs three layers of temporal checks: first, it validates that the ledger's own generated_at timestamp precedes the audit date; second, it enforces that source retrieved_at/ingested_at timestamps and claim reviewed_at timestamps are not after the audit date; third, for claim ledgers, it verifies that cited sources were retrieved before the audit date. Any violation produces a validation error that prevents future information from contaminating historical audits.
How are high-risk claims validated differently from standard claims?
Claims marked with risk: "high" undergo additional scrutiny through the _independent_group_count function. While standard claims only require evidence existence, high-risk claims must demonstrate support from multiple independent sources. The validator groups sources by shared canonical attributes (ID, origin, content hash, or declared independence key) and requires that the count of distinct groups meets the independence threshold. This prevents a single source cited multiple times from artificially inflating confidence in high-stakes assertions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →