How Stable Source IDs Are Generated in Claude-Obsidian: A 3-Step Canonicalization Guide
Claude-Obsidian generates deterministic source identifiers by canonicalizing input locators, concatenating normalized components with NUL separators, and computing a truncated SHA-256 hash prefixed with src-.
Claude-Obsidian relies on reproducible identifiers to maintain data integrity across ledger operations. The stable_source_id function implemented in claude_obsidian/ledgers.py creates stable source IDs from origin metadata, ensuring that identical sources always resolve to the same identifier regardless of when or where they are processed.
The Stable Source ID Generation Pipeline
The generation process follows a strict three-step pipeline defined in claude_obsidian/ledgers.py (lines 407-415). Because the identifier derives solely from normalized inputs, the function is idempotent and collision-resistant for practical ledger operations.
Step 1: Locator Normalization
First, the raw locator string undergoes origin-specific canonicalization based on the origin_kind parameter. The helper function _canonical_locator processes:
- URLs: Passed through
_canonical_urlto normalize encoding, case variations, and redundant parameters - File paths: Converted to POSIX format to eliminate platform-specific path differences
- Other types: Left unchanged but treated as literal strings
This ensures semantically equivalent locators map to identical representations before hashing.
Step 2: Component Hashing
The function constructs a single normalized string combining three elements separated by NUL characters (\0):
- Lower-cased
origin_kind(usingcasefold()) - Canonicalized locator from Step 1
- Optional content SHA-256 (empty string if
None, also casefolded)
The implementation encodes this string as UTF-8 with surrogatepass error handling, then computes the SHA-256 digest as implemented in lines 410-414:
digest = hashlib.sha256(
f"{origin_kind.casefold()}\0{normalized_locator}\0{(content_sha256 or '').casefold()}"
.encode("utf-8", errors="surrogatepass")
).hexdigest()
Step 3: Identifier Derivation
Finally, the function extracts the first 20 hexadecimal characters from the hash and prefixes them with "src-" to create the final identifier at line 415:
return f"src-{digest[:20]}"
This truncation provides sufficient entropy for collision resistance while keeping identifiers concise for ledger storage and transaction references.
Implementation Reference
The core logic resides in claude_obsidian/ledgers.py at lines 407-415. The stable_source_id function signature accepts three parameters:
origin_kind: String indicating the source type ("url","file", or custom values)locator: The raw location string (URL, filepath, or manual identifier)content_sha256: Optional hex digest of the content for integrity verification
Because the identifier derives deterministically from these inputs, regenerating the stable source ID with identical parameters always yields the same src- prefixed result.
Practical Code Examples
The following examples demonstrate how different source types generate consistent identifiers:
from claude_obsidian.ledgers import stable_source_id
# URL source without content verification
src_id = stable_source_id(
origin_kind="url",
locator="https://example.com/articles/intro?ref=home",
content_sha256=None,
)
print(src_id) # → src-a1b2c3d4e5f6a7b8c9d0
# File source with known content hash
src_id = stable_source_id(
origin_kind="file",
locator=".raw/notes/meeting.md",
content_sha256="e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
)
print(src_id) # → src-5f4e3d2c1b0a9f8e7d6c5b4a3
# Manual batch identifier
src_id = stable_source_id("manual", "batch-2023-07", None)
print(src_id) # → src-9c8b7a6d5e4f3a2b1c0d
Each call produces the same output when invoked with identical parameters, enabling reliable cross-referencing in ledger validation and evidence linking logic.
System Integration
Generated stable source IDs serve as primary keys throughout the Claude-Obsidian ecosystem. The claude_obsidian/transaction.py module uses these identifiers when constructing transaction drafts, while tests/test_ledgers.py validates their stability through unit tests like test_stable_source_ids. During linting operations, tests/test_lint_engine.py verifies that all source IDs conform to the required src- format.
Summary
- Canonicalization eliminates platform and encoding variations before hashing via
_canonical_locator - Deterministic hashing uses SHA-256 on a NUL-separated string of normalized components (origin kind, locator, and content hash)
- Truncated output takes the first 20 hex characters prefixed with
src-for practical identifier length - Idempotent generation ensures identical inputs always produce the same stable source ID
- Integration points span ledger validation, transaction handling, and evidence linking throughout the codebase
Frequently Asked Questions
What makes a source ID "stable" in Claude-Obsidian?
A stable source ID is deterministic: given the same origin_kind, canonical locator, and content_sha256, the stable_source_id function always returns an identical src- prefixed identifier. This reproducibility ensures that sources can be referenced consistently across different ledger operations and validation checks without relying on mutable timestamps or random generation.
How does Claude-Obsidian prevent collisions between different sources?
The implementation uses SHA-256 hashing on normalized inputs followed by truncation to 20 hexadecimal characters. While theoretically possible, SHA-256's 160-bit entropy (in the truncated output) makes accidental collisions computationally infeasible for typical document management use cases. The canonicalization step further reduces collision risk by ensuring that semantically equivalent URLs or file paths always hash to the same value.
Can I manually generate a stable source ID for external integration?
Yes. You can import and call the stable_source_id function directly from claude_obsidian.ledgers. Provide the appropriate origin_kind (e.g., "url" or "file"), the normalized locator string, and optionally the content SHA-256 hash. The function returns the same src- identifier that Claude-Obsidian would generate internally, enabling external tools to reference ledger entries unambiguously.
Why does the implementation use NUL separators and UTF-8 with surrogatepass?
The NUL character (\0) serves as an unambiguous delimiter that cannot appear in valid URLs, file paths, or hexadecimal hash strings, preventing parsing ambiguities during string reconstruction. The surrogatepass error handling ensures that even malformed UTF-16 surrogates in file paths are encoded deterministically without raising exceptions, maintaining stability across different operating system encodings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →