How Content-Addressing Works in the Claude-Obsidian Capture System
The Claude-Obsidian capture system implements content-addressed storage (CAS) by computing SHA-256 hashes of source files to generate deterministic paths, ensuring identical content always maps to the same immutable location while automatically deduplicating redundant data.
The capture subsystem in the AgriciDaniel/claude-obsidian repository provides a robust mechanism for ingesting external files into an Obsidian vault using content-addressing principles. This approach treats file content as the sole source of identity, creating an immutable, deduplicated storage layer that prevents duplicate copies and guarantees data integrity. Understanding how this content-addressing implementation works is essential for developers extending the capture pipeline or integrating with the raw storage backend.
The Three-Step Content-Addressing Pipeline
The implementation in claude_obsidian/capture.py enforces content-addressed storage through three tightly coupled operations that transform arbitrary source files into deterministic vault locations.
Step 1: Deterministic Identity Generation
The source_identity function establishes the foundation of the CAS model by computing a cryptographically secure hash of the file's exact byte stream. Located at lines 38-52, this function opens source files with O_NOFOLLOW to reject symbolic links, preventing path traversal attacks and ensuring the hash represents actual file content rather than link targets. During the read operation, it captures pre- and post-stat signatures to detect concurrent modifications; if the file changes during the hashing process, the capture aborts to preserve integrity.
Step 2: Destination Planning
Once the SHA-256 digest is computed, the _planned_destination function (lines 66-74) constructs the content-addressed path within <vault>/.raw/captured/. The target filename consists solely of the hexadecimal digest plus a safe extension, eliminating user-controlled naming that could introduce collisions or ambiguity. This function also validates that the raw store directory is not a symlink and confirms the final destination remains strictly within the vault boundaries using path canonicalization helpers from claude_obsidian/paths.py.
Step 3: Duplicate Detection and Immutability
The deduplication logic resides in plan_filesystem_batch (lines 215-232) and _find_existing_capture (lines 94-101). Before writing any payload, the system probes the .raw/captured/ directory for an existing file matching the computed digest. If found, the implementation performs a verification read to confirm the stored content matches the source digest. When contents align, the capture marks the operation as "unchanged" and skips the write entirely. This guarantees that identical content always resolves to the same address, maintains a single source of truth, and prevents storage bloat through automatic deduplication.
Practical Example: Capturing Files with Content-Addressing
The public capture_filesystem API exposes these internals through a simple interface that returns the content-derived identity and storage path:
from pathlib import Path
from claude_obsidian.capture import capture_filesystem
# Source file to ingest
src = Path("/home/user/documents/reference.md")
# Execute capture against target vault
vault = Path("/home/user/obsidian-vault")
result = capture_filesystem(vault, src)
print("SHA-256 Identity:", result["source_identity"])
print("Content-Addressed Path:", result["stored_path"])
print("New Write Required:", result["changed"])
Executing this against new content produces output similar to:
SHA-256 Identity: 3e2f9a7b5c4d8e1f...
Content-Addressed Path: .raw/captured/3e2f9a7b5c4d8e1f.bin
New Write Required: True
Re-running the same capture returns changed: False because the system detects the existing content-addressed file and refrains from redundant I/O operations.
Core Implementation Files
The content-addressing architecture spans several modules:
claude_obsidian/capture.py: Contains the primary CAS implementation includingsource_identity,_planned_destination, andplan_filesystem_batch, handling SHA-256 computation, path generation, and deduplication logic.claude_obsidian/paths.py: Providescanonicalandis_relative_toutilities that ensure all content-addressed destinations remain confined within the vault perimeter.claude_obsidian/transaction.py: Implementsapply_bundlefor atomic commits of capture batches, ensuring that verified content-addressed files move to their final locations only after successful integrity checks.tests/test_capture.py: Validates CAS behavior through unit tests covering duplicate detection, symlink rejection, and immutability guarantees.
Summary
- SHA-256 hashes serve as the sole addressing mechanism, computed over exact byte streams with
O_NOFOLLOWprotection against symlink attacks. - Immutable paths are constructed as
.raw/captured/<digest>.<ext>, ensuring content-derived locations that never change regardless of external context. - Automatic deduplication occurs when
_find_existing_capturediscovers matching digests, eliminating redundant storage operations and maintaining single-instance semantics. - Integrity verification through pre- and post-read stat comparisons prevents capturing files that modify during the ingestion window.
Frequently Asked Questions
Why does the capture system reject symbolic links?
The implementation explicitly opens files with O_NOFOLLOW to prevent symlink traversal attacks and ensure the SHA-256 hash reflects actual file content rather than link metadata. This security measure guarantees that content-addressing remains deterministic and that vault boundaries cannot be escaped through path manipulation.
How does the system handle concurrent file modifications during capture?
The source_identity function captures file metadata before and after reading the byte stream. If timestamps or sizes differ between these measurements, indicating concurrent modification, the capture aborts before computing the final digest. This prevents storing inconsistent or partial content at the addressed location.
Can two different files ever receive the same content-addressed path?
No. Because paths derive directly from SHA-256 digests of the exact byte content, distinct files with different content will always produce different hashes. In the astronomically unlikely event of a SHA-256 collision, the deduplication logic in _find_existing_capture would still verify byte-for-byte equality before skipping a write, ensuring correctness.
What happens to the content-addressed files after capture?
Files stored in .raw/captured/ are treated as immutable and read-only after initial write. Downstream processes such as wiki generation or linking systems can safely reference these stable addresses knowing the content will never change, enabling reliable cache strategies and persistent cross-references within the Obsidian vault.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →