# Claude-Obsidian Source Ledger Schemas and Kinds: A Complete Technical Reference

> Discover the claude-obsidian source ledger schemas and kinds. Learn about file, url, and manual sources, content types, authority levels, and review statuses for knowledge provenance.

- Repository: [Agrici.Daniel/claude-obsidian](https://github.com/AgriciDaniel/claude-obsidian)
- Tags: api-reference
- Published: 2026-08-29

---

**The claude-obsidian source ledger uses the fixed schema identifier `"claude-obsidian.source-ledger.v1"` and defines enumerated types including three source kinds (`file`, `url`, `manual`), nine content kinds, six authority levels, and four review statuses to validate knowledge provenance.**

The claude-obsidian project provides a structured framework for tracking the provenance of knowledge within Obsidian vaults. Understanding the **claude-obsidian source ledger schemas and kinds** is essential for developers extending the system or debugging validation errors. This guide examines the concrete schema definitions enforced in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) and the specific enumerated values that govern how source metadata is stored and validated.

## Schema Identifier

The source ledger must declare its compliance version through a strict schema string. In [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) at line 21, the `SOURCE_SCHEMA` constant defines the required identifier:

```python
SOURCE_SCHEMA = "claude-obsidian.source-ledger.v1"

```

Every source ledger JSON document must include this exact string in its top-level `"schema"` field. The `validate_source_ledger` function checks this value first and rejects any ledger using an unrecognized schema version.

## Source Kinds (origin.kind)

The `origin.kind` field describes how a source enters the system. Defined in [`ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/ledgers.py) lines 25-26, three enumerated values are permitted:

### File Sources

**`file`** indicates a vault-relative filesystem path. The `locator` field contains the relative path from the vault root, such as `"raw/data/sensor.csv"`. This kind is used for local documents, datasets, or media stored within the Obsidian vault.

### URL Sources

**`url`** represents external web resources accessed via HTTPS. The `locator` field stores the complete URL, enabling the system to track external documentation, web pages, or downloadable assets referenced by vault pages.

### Manual Sources

**`manual`** designates hand-entered entries without external locators. These sources represent user-provided knowledge, AI-generated summaries, or institutional knowledge that exists only within the ledger itself rather than pointing to an external file or URL.

## Content Kinds (content_kind)

Beyond acquisition method, the system categorizes the payload type through `content_kind`. Implemented in [`ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/ledgers.py) lines 26-37, nine distinct types classify source material:

- **`document`** – Textual documents and PDFs
- **`webpage`** – Rendered HTML pages captured from URLs
- **`dataset`** – Structured tabular or JSON data
- **`image`**, **`audio`**, **`video`** – Binary media assets
- **`code`** – Source code files and repositories
- **`conversation`** – Dialogue transcripts or chat logs
- **`synthetic`** – AI-generated content including summaries and analysis
- **`other`** – Fallback category for unclassified formats

## Authority Levels and Review Statuses

The ledger tracks source reliability and lifecycle state through two additional enumerated fields defined in [`ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/ledgers.py).

### Authority Levels

The `authority` field accepts one of six trust tiers defined at line 38:

- **`official`** – Authoritative institutional sources
- **`primary`** – Direct evidence or original research
- **`secondary`** – Analysis or interpretation of primary sources
- **`community`** – Crowdsourced or informal knowledge
- **`synthetic`** – AI-generated or algorithmically derived content
- **`unknown`** – Unclassified authority level

### Review Statuses

Source lifecycle management uses four states defined at line 39:

- **`unreviewed`** – Newly ingested, pending validation
- **`active`** – Verified and currently reliable
- **`superseded`** – Replaced by newer information
- **`rejected`** – Invalidated or deemed unreliable

## Implementation in claude_obsidian/ledgers.py

The constants governing these enumerations appear as module-level definitions in the ledger validation module:

```python

# claude_obsidian/ledgers.py

SOURCE_SCHEMA = "claude-obsidian.source-ledger.v1"

SOURCE_KINDS = {"file", "url", "manual"}
CONTENT_KINDS = {
    "document", "webpage", "dataset", "image", 
    "audio", "video", "code", "conversation", 
    "synthetic", "other"
}
AUTHORITY_LEVELS = {
    "official", "primary", "secondary", 
    "community", "synthetic", "unknown"
}
REVIEW_STATUSES = {
    "unreviewed", "active", "superseded", "rejected"
}

```

The `validate_source_ledger` function references these sets to enforce contract compliance, raising descriptive errors when a ledger deviates from the defined schemas.

## Complete Source Ledger Structure

A valid source ledger JSON document demonstrates the interaction of these schema elements:

```json
{
  "schema": "claude-obsidian.source-ledger.v1",
  "generated_at": "2024-10-01T12:00:00Z",
  "sources": {
    "src-1a2b3c4d5e6f7g8h9i0j": {
      "origin": { 
        "kind": "url", 
        "locator": "https://example.com/report.pdf" 
      },
      "content_kind": "document",
      "title": "Example Report",
      "authority": "official",
      "review_status": "active",
      "content_sha256": "a3f5c8d9e0b1c2d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7g8h9i0j",
      "ingested_at": "2024-09-28",
      "retrieved_at": "2024-09-28",
      "refresh_due": "2025-09-28",
      "pages": ["wiki/reports/example-report.md"]
    },
    "src-9j8i7h6g5f4e3d2c1b0a": {
      "origin": { 
        "kind": "file", 
        "locator": "raw/data/sensor.csv" 
      },
      "content_kind": "dataset",
      "title": "Sensor Readings",
      "authority": "primary",
      "review_status": "unreviewed",
      "content_sha256": null,
      "ingested_at": null,
      "retrieved_at": null,
      "refresh_due": null,
      "pages": []
    },
    "src-0a1b2c3d4e5f6g7h8i9j": {
      "origin": { 
        "kind": "manual", 
        "locator": "User‑provided summary" 
      },
      "content_kind": "synthetic",
      "title": "AI‑generated Summary",
      "authority": "synthetic",
      "review_status": "active",
      "content_sha256": null,
      "ingested_at": "2024-09-30",
      "retrieved_at": null,
      "refresh_due": "2025-09-30",
      "pages": []
    }
  }
}

```

Note that the top-level `schema` matches the `SOURCE_SCHEMA` constant, while each source entry uses valid enumerations for `origin.kind`, `content_kind`, `authority`, and `review_status`.

## Summary

- The **claude-obsidian source ledger** requires the schema identifier `"claude-obsidian.source-ledger.v1"` defined in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) line 21.
- Three **source kinds** (`file`, `url`, `manual`) control how sources are obtained and referenced through the `origin` object.
- Nine **content kinds** classify the payload type, from `document` and `dataset` to `synthetic` AI-generated content.
- Six **authority levels** and four **review statuses** provide metadata for trust assessment and lifecycle management.
- The `validate_source_ledger` function enforces these constraints using the constant definitions in [`ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/ledgers.py) lines 25-39.

## Frequently Asked Questions

### What is the required schema string for a claude-obsidian source ledger?

Every source ledger must include `"schema": "claude-obsidian.source-ledger.v1"` at the root level. This constant is defined in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) at line 21, and the `validate_source_ledger` function rejects any document using a different identifier.

### How do I reference a local file versus a URL in the source ledger?

Use the `file` kind for vault-relative paths and the `url` kind for HTTPS addresses. Both are set in the `origin.kind` field, with the specific path or URL stored in `origin.locator` as implemented in [`ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/ledgers.py) lines 25-26.

### What content kind should I use for AI-generated summaries?

Use `content_kind: "synthetic"` for AI-generated content, and set `authority: "synthetic"` to indicate the source type. This classification distinguishes generated content from primary documents or official sources according to the enumerations in [`ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/ledgers.py) lines 26-38.

### Where are the validation rules for these schemas enforced?

The `validate_source_ledger` function in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) enforces all schema constraints, checking the schema identifier first, then validating that `origin.kind`, `content_kind`, `authority`, and `review_status` match their respective enumerated sets defined in lines 25-39.