Content Kinds in the claude-obsidian Source Ledger: Complete Enumeration and Validation Guide
The claude-obsidian source ledger recognizes ten distinct content kinds—document, webpage, dataset, image, audio, video, code, conversation, synthetic, and other—enforced through the CONTENT_KINDS constant in claude_obsidian/ledgers.py and validated by the validate_source_ledger function.
The claude-obsidian source ledger system maintains rigorous provenance tracking by categorizing every source material through a controlled vocabulary of content types. This classification schema, defined as the CONTENT_KINDS enumeration in claude_obsidian/ledgers.py, ensures that wiki entries linked to external sources maintain consistent metadata regardless of whether the origin is a PDF document, web page, or AI-generated synthetic text.
The Ten Content Kinds Defined in CONTENT_KINDS
In claude_obsidian/ledgers.py (lines 25-38), the ledger defines a fixed enumeration of permissible content kinds. Each source entry must specify exactly one of the following values in its content_kind field:
- document: Generic textual documents including PDFs, Word files, and plain text files.
- webpage: Public web pages retrieved via HTTPS protocols.
- dataset: Structured data collections such as CSV, JSON, or database exports.
- image: Raster graphics (PNG, JPEG) and vector formats (SVG).
- audio: Audio recordings, podcasts, and sound files.
- video: Video files or streaming media content.
- code: Source code files, scripts, or code snippets from version control.
- conversation: Chat transcripts, dialogue logs, or messaging history.
- synthetic: Machine-generated content produced by AI models or automated systems.
- other: Fallback category for source materials that do not match the above classifications.
How Validation Enforces Content Kind Constraints
When the validate_source_ledger function processes a ledger, it performs strict type checking on the content_kind field. The validator confirms that the value is a string and exists within the CONTENT_KINDS enumeration. Any deviation—whether an unrecognized category or a non-string type—triggers a ledger validation error, preventing malformed entries from entering the knowledge base.
This validation step is critical for maintaining data integrity across the claude-obsidian ecosystem, as downstream processors rely on these explicit categories to apply appropriate rendering and linking strategies.
Creating Source Ledger Entries with Content Kinds
To instantiate a new source record, use the empty_source_ledger and stable_source_id utilities:
from claude_obsidian.ledgers import empty_source_ledger, stable_source_id
ledger = empty_source_ledger()
source_id = stable_source_id("url", "https://example.com/article", None)
ledger["sources"][source_id] = {
"origin": {"kind": "url", "locator": "https://example.com/article"},
"content_kind": "webpage",
"title": "Example Article",
"authority": "official",
"review_status": "unreviewed",
"content_sha256": None,
"pages": [],
}
Notice that content_kind is set to "webpage", which must match one of the ten allowed values in CONTENT_KINDS.
Validating Mixed-Content Ledgers
Complex ledgers containing multiple content types require validation before persistence:
from claude_obsidian.ledgers import validate_source_ledger
source_ledger = {
"schema": "claude-obsidian.source-ledger.v1",
"generated_at": "2026-08-26T12:00:00Z",
"sources": {
"src-abcdef1234567890abcd": {
"origin": {"kind": "file", "locator": "data/dataset.csv"},
"content_kind": "dataset",
"title": "Dataset",
"authority": "primary",
"review_status": "active",
"content_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"pages": ["wiki/dataset-overview.md"],
},
"src-1234567890abcdefabcd": {
"origin": {"kind": "url", "locator": "https://example.com/video.mp4"},
"content_kind": "video",
"title": "Tutorial Video",
"authority": "official",
"review_status": "active",
"content_sha256": None,
"pages": ["wiki/video-summary.md"],
},
},
}
errors = validate_source_ledger(source_ledger)
assert not errors, f"Ledger validation failed: {errors}"
The validation routine iterates through all entries, verifying that both "dataset" and "video" are legitimate members of the CONTENT_KINDS set.
Summary
- The claude-obsidian source ledger recognizes ten content kinds: document, webpage, dataset, image, audio, video, code, conversation, synthetic, and other.
- Valid values are defined in the
CONTENT_KINDSconstant withinclaude_obsidian/ledgers.py(lines 25-38). - The
validate_source_ledgerfunction enforces strict adherence to this enumeration, rejecting any undefined content types. - Each source entry must specify a
content_kindstring that matches the predefined vocabulary to pass validation. - The system supports heterogeneous ledgers containing multiple content types, provided all entries conform to the schema.
Frequently Asked Questions
What happens if I use an undefined content kind in the source ledger?
The validate_source_ledger function will return a validation error indicating that the content_kind value is not recognized. According to the implementation in claude_obsidian/ledgers.py, only the ten predefined categories in CONTENT_KINDS are acceptable; any other string will cause the ledger to fail validation.
Can I extend the content kinds to support custom media types?
Currently, the CONTENT_KINDS enumeration is fixed in the source code. To add new categories, you would need to modify the constant definition in claude_obsidian/ledgers.py and update the validation logic accordingly. The other category exists specifically to handle edge cases without requiring code changes.
How does the ledger distinguish between a webpage and a document source?
Both types use different content_kind values—webpage for HTTPS-retrieved content versus document for static files like PDFs or Word documents. While both may contain text, the distinction allows the system to apply appropriate caching, rendering, and authority verification strategies based on the source's nature.
Is the content kind used when generating wiki page links?
Yes, the content_kind field influences how sources are indexed and displayed in the wiki. While the pages array determines which markdown files reference the source, the content kind helps the Obsidian integration apply suitable preview templates and metadata formatting for different media types.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →