How Captures Are Stored and Prevented from Duplication in Claude‑Obsidian
Claude‑Obsidian stores every capture as an immutable SHA‑256‑named file in .raw/ and uses transaction‑level hash checks to ensure identical content is never written twice.
The open‑source knowledge management system AgriciDaniel/claude‑obsidian implements a content‑addressable storage layer that treats every user capture as a first‑class object. By combining cryptographic hashing with transactional commits, the system guarantees that duplicate information never pollutes the vault, even when ingestion scripts are run repeatedly.
The Capture Pipeline
Every piece of information entering the system follows a strict, immutable pipeline from transient input to permanent storage. According to the vault conventions documented in AGENTS.md, raw captures land in the user‑visible inbox/ directory before being promoted to the immutable .raw/ store. This two‑stage process ensures that users can review and edit content before it becomes part of the canonical knowledge base.
Immutable Storage with SHA‑256 Hashing
The core deduplication mechanism relies on content‑addressable storage implemented in claude_obsidian/capture.py. When a capture is ready for permanent storage, the system computes the SHA‑256 hash of its exact content and uses that hash as the filename inside the .raw/ directory.
Content‑Addressable File Naming
Because the filename is deterministically derived from the file’s contents, any attempt to store identical data resolves to the same filesystem path. The capture.py module materialises captures as plain files under .raw/, making the storage layer inherently idempotent. If a file with the computed hash already exists, the write operation is skipped entirely, preventing physical duplication on disk.
# Example: Adding a capture programmatically
from claude_obsidian.capture import Capture
text = "Important insight about AI alignment."
capture = Capture(text=text, source="user_manual")
capture.save() # Writes to inbox/
capture.commit() # Moves to .raw/ with SHA‑256 filename
Transaction‑Level Deduplication
Before any capture reaches the .raw/ directory, it passes through a transactional safety layer defined in claude_obsidian/transaction.py. The transaction wrapper enforces deduplication at the commit stage by checking target SHA‑256 values against the current vault state.
The Commit Phase
When capture.commit() is invoked, the transaction object queries the vault for existing hashes. If a file with the identical SHA‑256 hash is already present, the transaction treats that capture as a no‑op and aborts the write for that specific object. This prevents race conditions and ensures that even concurrent capture operations cannot create duplicates.
Provenance Tracking with Ledgers
After successful storage, claude_obsidian/ledgers.py records an entry in the provenance ledger located at wiki/meta/ledgers/. The ledger stores the hash, source location, and timestamp for every capture, creating an audit trail that makes accidental re‑ingestion easy to detect and reject.
# Inspecting the ledger for a given capture
from claude_obsidian.ledgers import Ledger
ledger = Ledger()
entries = ledger.find_by_hash("<sha256-of-capture>")
print(entries) # Shows provenance, timestamp, and source
Idempotent Vault Setup
The helper scripts bin/setup‑vault.sh and scripts/claude‑obsidian.py invoke the same capture‑to‑.raw pipeline used by the Python API. Because they rely on the hash‑based naming scheme, running these scripts multiple times is safe—the vault will always end up with exactly one copy of each unique capture, regardless of how many times the setup command is executed.
# Example: Running the idempotent vault‑setup script
$ bin/setup-vault.sh my-vault
# The script will invoke capture → .raw conversion.
# Re‑running the command will not duplicate existing captures.
Summary
- Immutable payloads: Captures are stored in
.raw/using SHA‑256 hashes as filenames, ensuring identical content maps to a single file path. - Transactional safety: The
transaction.pylayer validates hashes before committing, turning duplicate writes into no‑ops. - Audit trail: The ledger system in
ledgers.pyrecords every capture’s provenance, providing a secondary check against duplication. - Idempotent operations: Setup scripts and API calls can be retried safely without risking duplicate data.
Frequently Asked Questions
What happens if I try to capture the same text twice?
The second attempt resolves to the same SHA‑256 filename in .raw/. The transaction layer detects the existing file and skips the write, returning success without creating a duplicate. The ledger entry remains unchanged.
Where are captures stored before processing?
New captures are written to the inbox/ directory, a visible folder that is never pruned automatically. Users can edit or delete files here before calling commit() to move them into the immutable .raw/ store.
How does the ledger help with deduplication?
The ledger maintains a permanent record of every hash stored in the vault. While the filesystem check in capture.py provides immediate deduplication, the ledger allows administrators to audit when a specific capture was first added and verify that no hash collisions or manual copies exist.
Can I safely run setup scripts multiple times?
Yes. The bin/setup‑vault.sh script uses the same hash‑based pipeline as the Python API. Re‑running it against an existing vault will only process new captures, making it safe to execute in CI/CD pipelines or during incremental backups.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →