How Claude-Obsidian Handles Synthetic Content and Authority Declaration: Validation Rules and Implementation
Claude-Obsidian treats synthetic material as a first-class content kind but enforces strict bidirectional validation that requires synthetic content to always pair with synthetic authority, raising LedgerValidationError when mismatched.
Managing artificial knowledge requires rigorous provenance tracking to maintain knowledge base integrity. In the AgriciDaniel/claude-obsidian repository, synthetic content and authority declaration are tightly coupled through ledger validation rules that prevent mislabeling and ensure auditability across BM25 indexing and retrieval pipelines.
The Ledger Schema Definition
Content Kinds and Authority Sets
The foundation resides in claude_obsidian/ledgers.py, where the system defines valid taxonomies for knowledge classification. The CONTENT_KINDS set includes "synthetic" alongside traditional sources like documents and webpages (lines 35-38):
CONTENT_KINDS = {
"document",
"webpage",
"dataset",
"synthetic",
# ...
}
Similarly, the AUTHORITIES enumeration recognizes "synthetic" as a valid authority value (lines 38-39). This dual classification allows the system to tag artificial content distinctly from organic sources while maintaining schema consistency for downstream processing.
Bidirectional Validation Rules
The Synthetic Authority Invariant
The critical enforcement mechanism lives in the validate_source_ledger function (lines 583-587 of claude_obsidian/ledgers.py). This function implements an XOR-style invariant that mandates exact pairing between content kind and authority:
if (content_kind == "synthetic") != (authority == "synthetic"):
raise LedgerValidationError([
{
"path": path,
"message": "synthetic content and synthetic authority must be declared together",
}
])
This validation ensures three strict constraints:
- Synthetic content cannot claim real-world authority
- Real content cannot be mislabeled as synthetic
- Every synthetic entry carries explicit provenance marking required for retrieval filtering
Synthetic Prefix Generation
Forced Synthetic Mode
When generating artificial chunks for testing or offline operation, scripts/contextual-prefix.py manages the synthetic tier selection through the pick_prefix_tier helper (lines 672-689). When force_synthetic is enabled or LLM egress is disabled, the system bypasses environment variable checks and returns metadata with source = "synthetic":
# Conceptual implementation from pick_prefix_tier
if force_synthetic or not allow_llm_calls:
return {"tier": "synthetic", "source": "synthetic"}
This guarantees that programmatically generated synthetic chunks automatically satisfy the ledger validation requirements without manual metadata configuration, ensuring seamless integration with the source ledger.
Practical Implementation Examples
Creating valid synthetic entries requires simultaneous declaration of both fields to pass validation:
from claude_obsidian.ledgers import validate_source_ledger
from pathlib import Path
source_entry = {
"content_kind": "synthetic",
"authority": "synthetic",
"address": "synthetic://example-001",
"title": "Demo synthetic chunk",
}
# Validation passes only when both fields match
validate_source_ledger([source_entry], vault_root=Path("/my/vault"))
For command-line generation of synthetic prefixes without LLM calls:
python scripts/contextual-prefix.py path/to/page.md --no-llm
The script outputs prefixes with source = "synthetic", ensuring immediate compatibility with the ledger schema defined in claude_obsidian/ledgers.py.
Summary
- Synthetic taxonomy:
CONTENT_KINDSandAUTHORITIESinclaude_obsidian/ledgers.pyboth include"synthetic"as valid enumeration values - Bidirectional validation: The
validate_source_ledgerfunction enforces thatcontent_kind == "synthetic"if and only ifauthority == "synthetic" - Error handling: Mismatched pairs immediately raise
LedgerValidationErrorwith the message "synthetic content and synthetic authority must be declared together" and explicit path tracking - Automated compliance:
pick_prefix_tierinscripts/contextual-prefix.pyautomatically assigns synthetic metadata when forcing synthetic mode via--no-llmflags - Provenance integrity: This coupling prevents accidental mislabeling that could compromise retrieval accuracy or contaminate the knowledge base with unmarked artificial data
Frequently Asked Questions
What error occurs if I declare synthetic content without synthetic authority?
The validate_source_ledger function raises a LedgerValidationError containing the specific message "synthetic content and synthetic authority must be declared together" and the file path where the mismatch occurred. This prevents the entry from entering the knowledge base with inconsistent provenance metadata that could corrupt downstream retrieval pipelines.
How do I generate synthetic chunks for testing purposes?
Use the contextual prefix script with the --no-llm flag: python scripts/contextual-prefix.py <file> --no-llm. This invokes pick_prefix_tier with force_synthetic enabled, automatically setting source = "synthetic" and ensuring the generated metadata passes ledger validation without requiring manual authority assignment.
Why does Claude-Obsidian require synthetic content to declare synthetic authority?
This bidirectional requirement creates an immutable audit trail that distinguishes artificial knowledge from real-world sources. It prevents synthetic training data or test fixtures from being incorrectly indexed as authoritative documents, which is essential for maintaining retrieval accuracy and provenance tracking in downstream BM25 indexing and RAG pipelines.
Can real-world content be assigned synthetic authority?
No. The validation logic uses an equivalence check (content_kind == "synthetic") != (authority == "synthetic") that raises errors in both directions. Real content with non-synthetic kinds must use non-synthetic authorities, ensuring the authority declaration accurately reflects the content's actual origin and preventing synthetic pollution of legitimate knowledge bases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →