# Content Kinds in the claude-obsidian Source Ledger: Complete Enumeration and Validation Guide

> Discover the ten content kinds tracked in the claude-obsidian source ledger document webpage dataset image audio video code conversation synthetic other. Learn validation methods.

- Repository: [Agrici.Daniel/claude-obsidian](https://github.com/AgriciDaniel/claude-obsidian)
- Tags: deep-dive
- Published: 2026-08-26

---

**The claude-obsidian source ledger recognizes ten distinct content kinds—document, webpage, dataset, image, audio, video, code, conversation, synthetic, and other—enforced through the `CONTENT_KINDS` constant in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) and validated by the `validate_source_ledger` function.**

The claude-obsidian source ledger system maintains rigorous provenance tracking by categorizing every source material through a controlled vocabulary of content types. This classification schema, defined as the `CONTENT_KINDS` enumeration in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py), ensures that wiki entries linked to external sources maintain consistent metadata regardless of whether the origin is a PDF document, web page, or AI-generated synthetic text.

## The Ten Content Kinds Defined in CONTENT_KINDS

In [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) (lines 25-38), the ledger defines a fixed enumeration of permissible content kinds. Each source entry must specify exactly one of the following values in its `content_kind` field:

- **document**: Generic textual documents including PDFs, Word files, and plain text files.
- **webpage**: Public web pages retrieved via HTTPS protocols.
- **dataset**: Structured data collections such as CSV, JSON, or database exports.
- **image**: Raster graphics (PNG, JPEG) and vector formats (SVG).
- **audio**: Audio recordings, podcasts, and sound files.
- **video**: Video files or streaming media content.
- **code**: Source code files, scripts, or code snippets from version control.
- **conversation**: Chat transcripts, dialogue logs, or messaging history.
- **synthetic**: Machine-generated content produced by AI models or automated systems.
- **other**: Fallback category for source materials that do not match the above classifications.

## How Validation Enforces Content Kind Constraints

When the `validate_source_ledger` function processes a ledger, it performs strict type checking on the `content_kind` field. The validator confirms that the value is a string and exists within the `CONTENT_KINDS` enumeration. Any deviation—whether an unrecognized category or a non-string type—triggers a ledger validation error, preventing malformed entries from entering the knowledge base.

This validation step is critical for maintaining data integrity across the claude-obsidian ecosystem, as downstream processors rely on these explicit categories to apply appropriate rendering and linking strategies.

## Creating Source Ledger Entries with Content Kinds

To instantiate a new source record, use the `empty_source_ledger` and `stable_source_id` utilities:

```python
from claude_obsidian.ledgers import empty_source_ledger, stable_source_id

ledger = empty_source_ledger()
source_id = stable_source_id("url", "https://example.com/article", None)

ledger["sources"][source_id] = {
    "origin": {"kind": "url", "locator": "https://example.com/article"},
    "content_kind": "webpage",
    "title": "Example Article",
    "authority": "official",
    "review_status": "unreviewed",
    "content_sha256": None,
    "pages": [],
}

```

Notice that `content_kind` is set to `"webpage"`, which must match one of the ten allowed values in `CONTENT_KINDS`.

## Validating Mixed-Content Ledgers

Complex ledgers containing multiple content types require validation before persistence:

```python
from claude_obsidian.ledgers import validate_source_ledger

source_ledger = {
    "schema": "claude-obsidian.source-ledger.v1",
    "generated_at": "2026-08-26T12:00:00Z",
    "sources": {
        "src-abcdef1234567890abcd": {
            "origin": {"kind": "file", "locator": "data/dataset.csv"},
            "content_kind": "dataset",
            "title": "Dataset",
            "authority": "primary",
            "review_status": "active",
            "content_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
            "pages": ["wiki/dataset-overview.md"],
        },
        "src-1234567890abcdefabcd": {
            "origin": {"kind": "url", "locator": "https://example.com/video.mp4"},
            "content_kind": "video",
            "title": "Tutorial Video",
            "authority": "official",
            "review_status": "active",
            "content_sha256": None,
            "pages": ["wiki/video-summary.md"],
        },
    },
}

errors = validate_source_ledger(source_ledger)
assert not errors, f"Ledger validation failed: {errors}"

```

The validation routine iterates through all entries, verifying that both `"dataset"` and `"video"` are legitimate members of the `CONTENT_KINDS` set.

## Summary

- The claude-obsidian source ledger recognizes **ten content kinds**: document, webpage, dataset, image, audio, video, code, conversation, synthetic, and other.
- Valid values are defined in the `CONTENT_KINDS` constant within [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) (lines 25-38).
- The `validate_source_ledger` function enforces strict adherence to this enumeration, rejecting any undefined content types.
- Each source entry must specify a `content_kind` string that matches the predefined vocabulary to pass validation.
- The system supports heterogeneous ledgers containing multiple content types, provided all entries conform to the schema.

## Frequently Asked Questions

### What happens if I use an undefined content kind in the source ledger?

The `validate_source_ledger` function will return a validation error indicating that the `content_kind` value is not recognized. According to the implementation in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py), only the ten predefined categories in `CONTENT_KINDS` are acceptable; any other string will cause the ledger to fail validation.

### Can I extend the content kinds to support custom media types?

Currently, the `CONTENT_KINDS` enumeration is fixed in the source code. To add new categories, you would need to modify the constant definition in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) and update the validation logic accordingly. The `other` category exists specifically to handle edge cases without requiring code changes.

### How does the ledger distinguish between a webpage and a document source?

Both types use different `content_kind` values—`webpage` for HTTPS-retrieved content versus `document` for static files like PDFs or Word documents. While both may contain text, the distinction allows the system to apply appropriate caching, rendering, and authority verification strategies based on the source's nature.

### Is the content kind used when generating wiki page links?

Yes, the `content_kind` field influences how sources are indexed and displayed in the wiki. While the `pages` array determines which markdown files reference the source, the content kind helps the Obsidian integration apply suitable preview templates and metadata formatting for different media types.