Claude-Obsidian Source Ledger Schemas and Kinds: A Complete Technical Reference
The claude-obsidian source ledger uses the fixed schema identifier "claude-obsidian.source-ledger.v1" and defines enumerated types including three source kinds (file, url, manual), nine content kinds, six authority levels, and four review statuses to validate knowledge provenance.
The claude-obsidian project provides a structured framework for tracking the provenance of knowledge within Obsidian vaults. Understanding the claude-obsidian source ledger schemas and kinds is essential for developers extending the system or debugging validation errors. This guide examines the concrete schema definitions enforced in claude_obsidian/ledgers.py and the specific enumerated values that govern how source metadata is stored and validated.
Schema Identifier
The source ledger must declare its compliance version through a strict schema string. In claude_obsidian/ledgers.py at line 21, the SOURCE_SCHEMA constant defines the required identifier:
SOURCE_SCHEMA = "claude-obsidian.source-ledger.v1"
Every source ledger JSON document must include this exact string in its top-level "schema" field. The validate_source_ledger function checks this value first and rejects any ledger using an unrecognized schema version.
Source Kinds (origin.kind)
The origin.kind field describes how a source enters the system. Defined in ledgers.py lines 25-26, three enumerated values are permitted:
File Sources
file indicates a vault-relative filesystem path. The locator field contains the relative path from the vault root, such as "raw/data/sensor.csv". This kind is used for local documents, datasets, or media stored within the Obsidian vault.
URL Sources
url represents external web resources accessed via HTTPS. The locator field stores the complete URL, enabling the system to track external documentation, web pages, or downloadable assets referenced by vault pages.
Manual Sources
manual designates hand-entered entries without external locators. These sources represent user-provided knowledge, AI-generated summaries, or institutional knowledge that exists only within the ledger itself rather than pointing to an external file or URL.
Content Kinds (content_kind)
Beyond acquisition method, the system categorizes the payload type through content_kind. Implemented in ledgers.py lines 26-37, nine distinct types classify source material:
document– Textual documents and PDFswebpage– Rendered HTML pages captured from URLsdataset– Structured tabular or JSON dataimage,audio,video– Binary media assetscode– Source code files and repositoriesconversation– Dialogue transcripts or chat logssynthetic– AI-generated content including summaries and analysisother– Fallback category for unclassified formats
Authority Levels and Review Statuses
The ledger tracks source reliability and lifecycle state through two additional enumerated fields defined in ledgers.py.
Authority Levels
The authority field accepts one of six trust tiers defined at line 38:
official– Authoritative institutional sourcesprimary– Direct evidence or original researchsecondary– Analysis or interpretation of primary sourcescommunity– Crowdsourced or informal knowledgesynthetic– AI-generated or algorithmically derived contentunknown– Unclassified authority level
Review Statuses
Source lifecycle management uses four states defined at line 39:
unreviewed– Newly ingested, pending validationactive– Verified and currently reliablesuperseded– Replaced by newer informationrejected– Invalidated or deemed unreliable
Implementation in claude_obsidian/ledgers.py
The constants governing these enumerations appear as module-level definitions in the ledger validation module:
# claude_obsidian/ledgers.py
SOURCE_SCHEMA = "claude-obsidian.source-ledger.v1"
SOURCE_KINDS = {"file", "url", "manual"}
CONTENT_KINDS = {
"document", "webpage", "dataset", "image",
"audio", "video", "code", "conversation",
"synthetic", "other"
}
AUTHORITY_LEVELS = {
"official", "primary", "secondary",
"community", "synthetic", "unknown"
}
REVIEW_STATUSES = {
"unreviewed", "active", "superseded", "rejected"
}
The validate_source_ledger function references these sets to enforce contract compliance, raising descriptive errors when a ledger deviates from the defined schemas.
Complete Source Ledger Structure
A valid source ledger JSON document demonstrates the interaction of these schema elements:
{
"schema": "claude-obsidian.source-ledger.v1",
"generated_at": "2024-10-01T12:00:00Z",
"sources": {
"src-1a2b3c4d5e6f7g8h9i0j": {
"origin": {
"kind": "url",
"locator": "https://example.com/report.pdf"
},
"content_kind": "document",
"title": "Example Report",
"authority": "official",
"review_status": "active",
"content_sha256": "a3f5c8d9e0b1c2d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7g8h9i0j",
"ingested_at": "2024-09-28",
"retrieved_at": "2024-09-28",
"refresh_due": "2025-09-28",
"pages": ["wiki/reports/example-report.md"]
},
"src-9j8i7h6g5f4e3d2c1b0a": {
"origin": {
"kind": "file",
"locator": "raw/data/sensor.csv"
},
"content_kind": "dataset",
"title": "Sensor Readings",
"authority": "primary",
"review_status": "unreviewed",
"content_sha256": null,
"ingested_at": null,
"retrieved_at": null,
"refresh_due": null,
"pages": []
},
"src-0a1b2c3d4e5f6g7h8i9j": {
"origin": {
"kind": "manual",
"locator": "User‑provided summary"
},
"content_kind": "synthetic",
"title": "AI‑generated Summary",
"authority": "synthetic",
"review_status": "active",
"content_sha256": null,
"ingested_at": "2024-09-30",
"retrieved_at": null,
"refresh_due": "2025-09-30",
"pages": []
}
}
}
Note that the top-level schema matches the SOURCE_SCHEMA constant, while each source entry uses valid enumerations for origin.kind, content_kind, authority, and review_status.
Summary
- The claude-obsidian source ledger requires the schema identifier
"claude-obsidian.source-ledger.v1"defined inclaude_obsidian/ledgers.pyline 21. - Three source kinds (
file,url,manual) control how sources are obtained and referenced through theoriginobject. - Nine content kinds classify the payload type, from
documentanddatasettosyntheticAI-generated content. - Six authority levels and four review statuses provide metadata for trust assessment and lifecycle management.
- The
validate_source_ledgerfunction enforces these constraints using the constant definitions inledgers.pylines 25-39.
Frequently Asked Questions
What is the required schema string for a claude-obsidian source ledger?
Every source ledger must include "schema": "claude-obsidian.source-ledger.v1" at the root level. This constant is defined in claude_obsidian/ledgers.py at line 21, and the validate_source_ledger function rejects any document using a different identifier.
How do I reference a local file versus a URL in the source ledger?
Use the file kind for vault-relative paths and the url kind for HTTPS addresses. Both are set in the origin.kind field, with the specific path or URL stored in origin.locator as implemented in ledgers.py lines 25-26.
What content kind should I use for AI-generated summaries?
Use content_kind: "synthetic" for AI-generated content, and set authority: "synthetic" to indicate the source type. This classification distinguishes generated content from primary documents or official sources according to the enumerations in ledgers.py lines 26-38.
Where are the validation rules for these schemas enforced?
The validate_source_ledger function in claude_obsidian/ledgers.py enforces all schema constraints, checking the schema identifier first, then validating that origin.kind, content_kind, authority, and review_status match their respective enumerated sets defined in lines 25-39.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →