# How Stable Source IDs Are Generated in Claude-Obsidian: A 3-Step Canonicalization Guide

> Learn how Claude-Obsidian generates stable source IDs with this 3-step canonicalization guide. Discover the process of creating deterministic identifiers for your Claude Obsidian data.

- Repository: [Agrici.Daniel/claude-obsidian](https://github.com/AgriciDaniel/claude-obsidian)
- Tags: how-to-guide
- Published: 2026-08-29

---

**Claude-Obsidian generates deterministic source identifiers by canonicalizing input locators, concatenating normalized components with NUL separators, and computing a truncated SHA-256 hash prefixed with `src-`.**

Claude-Obsidian relies on reproducible identifiers to maintain data integrity across ledger operations. The `stable_source_id` function implemented in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) creates **stable source IDs** from origin metadata, ensuring that identical sources always resolve to the same identifier regardless of when or where they are processed.

## The Stable Source ID Generation Pipeline

The generation process follows a strict three-step pipeline defined in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) (lines 407-415). Because the identifier derives solely from normalized inputs, the function is idempotent and collision-resistant for practical ledger operations.

### Step 1: Locator Normalization

First, the raw locator string undergoes origin-specific canonicalization based on the `origin_kind` parameter. The helper function `_canonical_locator` processes:

- **URLs**: Passed through `_canonical_url` to normalize encoding, case variations, and redundant parameters
- **File paths**: Converted to POSIX format to eliminate platform-specific path differences
- **Other types**: Left unchanged but treated as literal strings

This ensures semantically equivalent locators map to identical representations before hashing.

### Step 2: Component Hashing

The function constructs a single normalized string combining three elements separated by NUL characters (`\0`):

1. Lower-cased `origin_kind` (using `casefold()`)
2. Canonicalized locator from Step 1
3. Optional content SHA-256 (empty string if `None`, also casefolded)

The implementation encodes this string as UTF-8 with `surrogatepass` error handling, then computes the SHA-256 digest as implemented in lines 410-414:

```python
digest = hashlib.sha256(
    f"{origin_kind.casefold()}\0{normalized_locator}\0{(content_sha256 or '').casefold()}"
    .encode("utf-8", errors="surrogatepass")
).hexdigest()

```

### Step 3: Identifier Derivation

Finally, the function extracts the first 20 hexadecimal characters from the hash and prefixes them with `"src-"` to create the final identifier at line 415:

```python
return f"src-{digest[:20]}"

```

This truncation provides sufficient entropy for collision resistance while keeping identifiers concise for ledger storage and transaction references.

## Implementation Reference

The core logic resides in [`claude_obsidian/ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/ledgers.py) at lines 407-415. The `stable_source_id` function signature accepts three parameters:

- `origin_kind`: String indicating the source type (`"url"`, `"file"`, or custom values)
- `locator`: The raw location string (URL, filepath, or manual identifier)
- `content_sha256`: Optional hex digest of the content for integrity verification

Because the identifier derives deterministically from these inputs, regenerating the **stable source ID** with identical parameters always yields the same `src-` prefixed result.

## Practical Code Examples

The following examples demonstrate how different source types generate consistent identifiers:

```python
from claude_obsidian.ledgers import stable_source_id

# URL source without content verification

src_id = stable_source_id(
    origin_kind="url",
    locator="https://example.com/articles/intro?ref=home",
    content_sha256=None,
)
print(src_id)   # → src-a1b2c3d4e5f6a7b8c9d0

# File source with known content hash

src_id = stable_source_id(
    origin_kind="file",
    locator=".raw/notes/meeting.md",
    content_sha256="e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
)
print(src_id)   # → src-5f4e3d2c1b0a9f8e7d6c5b4a3

# Manual batch identifier

src_id = stable_source_id("manual", "batch-2023-07", None)
print(src_id)   # → src-9c8b7a6d5e4f3a2b1c0d

```

Each call produces the same output when invoked with identical parameters, enabling reliable cross-referencing in ledger validation and evidence linking logic.

## System Integration

Generated **stable source IDs** serve as primary keys throughout the Claude-Obsidian ecosystem. The [`claude_obsidian/transaction.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/transaction.py) module uses these identifiers when constructing transaction drafts, while [`tests/test_ledgers.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/tests/test_ledgers.py) validates their stability through unit tests like `test_stable_source_ids`. During linting operations, [`tests/test_lint_engine.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/tests/test_lint_engine.py) verifies that all source IDs conform to the required `src-` format.

## Summary

- **Canonicalization** eliminates platform and encoding variations before hashing via `_canonical_locator`
- **Deterministic hashing** uses SHA-256 on a NUL-separated string of normalized components (origin kind, locator, and content hash)
- **Truncated output** takes the first 20 hex characters prefixed with `src-` for practical identifier length
- **Idempotent generation** ensures identical inputs always produce the same stable source ID
- **Integration points** span ledger validation, transaction handling, and evidence linking throughout the codebase

## Frequently Asked Questions

### What makes a source ID "stable" in Claude-Obsidian?

A **stable source ID** is deterministic: given the same `origin_kind`, canonical `locator`, and `content_sha256`, the `stable_source_id` function always returns an identical `src-` prefixed identifier. This reproducibility ensures that sources can be referenced consistently across different ledger operations and validation checks without relying on mutable timestamps or random generation.

### How does Claude-Obsidian prevent collisions between different sources?

The implementation uses SHA-256 hashing on normalized inputs followed by truncation to 20 hexadecimal characters. While theoretically possible, SHA-256's 160-bit entropy (in the truncated output) makes accidental collisions computationally infeasible for typical document management use cases. The canonicalization step further reduces collision risk by ensuring that semantically equivalent URLs or file paths always hash to the same value.

### Can I manually generate a stable source ID for external integration?

Yes. You can import and call the `stable_source_id` function directly from `claude_obsidian.ledgers`. Provide the appropriate `origin_kind` (e.g., `"url"` or `"file"`), the normalized locator string, and optionally the content SHA-256 hash. The function returns the same `src-` identifier that Claude-Obsidian would generate internally, enabling external tools to reference ledger entries unambiguously.

### Why does the implementation use NUL separators and UTF-8 with surrogatepass?

The NUL character (`\0`) serves as an unambiguous delimiter that cannot appear in valid URLs, file paths, or hexadecimal hash strings, preventing parsing ambiguities during string reconstruction. The `surrogatepass` error handling ensures that even malformed UTF-16 surrogates in file paths are encoded deterministically without raising exceptions, maintaining stability across different operating system encodings.