# How Content-Addressing Works in the Claude-Obsidian Capture System

> Learn how the Claude-Obsidian capture system uses content-addressing with SHA-256 hashes for deterministic paths, immutable storage, and automatic data deduplication.

- Repository: [Agrici.Daniel/claude-obsidian](https://github.com/AgriciDaniel/claude-obsidian)
- Tags: internals
- Published: 2026-08-29

---

**The Claude-Obsidian capture system implements content-addressed storage (CAS) by computing SHA-256 hashes of source files to generate deterministic paths, ensuring identical content always maps to the same immutable location while automatically deduplicating redundant data.**

The capture subsystem in the `AgriciDaniel/claude-obsidian` repository provides a robust mechanism for ingesting external files into an Obsidian vault using content-addressing principles. This approach treats file content as the sole source of identity, creating an immutable, deduplicated storage layer that prevents duplicate copies and guarantees data integrity. Understanding how this content-addressing implementation works is essential for developers extending the capture pipeline or integrating with the raw storage backend.

## The Three-Step Content-Addressing Pipeline

The implementation in [`claude_obsidian/capture.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/capture.py) enforces content-addressed storage through three tightly coupled operations that transform arbitrary source files into deterministic vault locations.

### Step 1: Deterministic Identity Generation

The `source_identity` function establishes the foundation of the CAS model by computing a cryptographically secure hash of the file's exact byte stream. Located at lines 38-52, this function opens source files with `O_NOFOLLOW` to reject symbolic links, preventing path traversal attacks and ensuring the hash represents actual file content rather than link targets. During the read operation, it captures pre- and post-`stat` signatures to detect concurrent modifications; if the file changes during the hashing process, the capture aborts to preserve integrity.

### Step 2: Destination Planning

Once the SHA-256 digest is computed, the `_planned_destination` function (lines 66-74) constructs the content-addressed path within `<vault>/.raw/captured/`. The target filename consists solely of the hexadecimal digest plus a safe extension, eliminating user-controlled naming that could introduce collisions or ambiguity. This function also validates that the raw store directory is not a symlink and confirms the final destination remains strictly within the vault boundaries using path canonicalization helpers from [`claude_obsidian/paths.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/paths.py).

### Step 3: Duplicate Detection and Immutability

The deduplication logic resides in `plan_filesystem_batch` (lines 215-232) and `_find_existing_capture` (lines 94-101). Before writing any payload, the system probes the `.raw/captured/` directory for an existing file matching the computed digest. If found, the implementation performs a verification read to confirm the stored content matches the source digest. When contents align, the capture marks the operation as "unchanged" and skips the write entirely. This guarantees that identical content always resolves to the same address, maintains a single source of truth, and prevents storage bloat through automatic deduplication.

## Practical Example: Capturing Files with Content-Addressing

The public `capture_filesystem` API exposes these internals through a simple interface that returns the content-derived identity and storage path:

```python
from pathlib import Path
from claude_obsidian.capture import capture_filesystem

# Source file to ingest

src = Path("/home/user/documents/reference.md")

# Execute capture against target vault

vault = Path("/home/user/obsidian-vault")
result = capture_filesystem(vault, src)

print("SHA-256 Identity:", result["source_identity"])
print("Content-Addressed Path:", result["stored_path"])
print("New Write Required:", result["changed"])

```

Executing this against new content produces output similar to:

```text
SHA-256 Identity: 3e2f9a7b5c4d8e1f...
Content-Addressed Path: .raw/captured/3e2f9a7b5c4d8e1f.bin
New Write Required: True

```

Re-running the same capture returns `changed: False` because the system detects the existing content-addressed file and refrains from redundant I/O operations.

## Core Implementation Files

The content-addressing architecture spans several modules:

- **[`claude_obsidian/capture.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/capture.py)**: Contains the primary CAS implementation including `source_identity`, `_planned_destination`, and `plan_filesystem_batch`, handling SHA-256 computation, path generation, and deduplication logic.
- **[`claude_obsidian/paths.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/paths.py)**: Provides `canonical` and `is_relative_to` utilities that ensure all content-addressed destinations remain confined within the vault perimeter.
- **[`claude_obsidian/transaction.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/claude_obsidian/transaction.py)**: Implements `apply_bundle` for atomic commits of capture batches, ensuring that verified content-addressed files move to their final locations only after successful integrity checks.
- **[`tests/test_capture.py`](https://github.com/AgriciDaniel/claude-obsidian/blob/main/tests/test_capture.py)**: Validates CAS behavior through unit tests covering duplicate detection, symlink rejection, and immutability guarantees.

## Summary

- **SHA-256 hashes** serve as the sole addressing mechanism, computed over exact byte streams with `O_NOFOLLOW` protection against symlink attacks.
- **Immutable paths** are constructed as `.raw/captured/<digest>.<ext>`, ensuring content-derived locations that never change regardless of external context.
- **Automatic deduplication** occurs when `_find_existing_capture` discovers matching digests, eliminating redundant storage operations and maintaining single-instance semantics.
- **Integrity verification** through pre- and post-read stat comparisons prevents capturing files that modify during the ingestion window.

## Frequently Asked Questions

### Why does the capture system reject symbolic links?

The implementation explicitly opens files with `O_NOFOLLOW` to prevent symlink traversal attacks and ensure the SHA-256 hash reflects actual file content rather than link metadata. This security measure guarantees that content-addressing remains deterministic and that vault boundaries cannot be escaped through path manipulation.

### How does the system handle concurrent file modifications during capture?

The `source_identity` function captures file metadata before and after reading the byte stream. If timestamps or sizes differ between these measurements, indicating concurrent modification, the capture aborts before computing the final digest. This prevents storing inconsistent or partial content at the addressed location.

### Can two different files ever receive the same content-addressed path?

No. Because paths derive directly from SHA-256 digests of the exact byte content, distinct files with different content will always produce different hashes. In the astronomically unlikely event of a SHA-256 collision, the deduplication logic in `_find_existing_capture` would still verify byte-for-byte equality before skipping a write, ensuring correctness.

### What happens to the content-addressed files after capture?

Files stored in `.raw/captured/` are treated as immutable and read-only after initial write. Downstream processes such as wiki generation or linking systems can safely reference these stable addresses knowing the content will never change, enabling reliable cache strategies and persistent cross-references within the Obsidian vault.