How the Claude-Obsidian Capture System Implements Offline-First Primitives
The claude-obsidian capture system implements offline-first primitives by confining all operations to the local filesystem, using atomic writes and SHA-256 identity checks, while expressing network-dependent operations as inert command plans that require explicit external execution.
The AgriciDaniel/claude-obsidian repository provides a knowledge management pipeline designed to function entirely without network connectivity. Its capture module enforces strict offline-first primitives by ensuring that all core operations interact solely with the local filesystem through confined directories, advisory locking, and atomic transactions. Every function in claude_obsidian/capture.py—from path validation to batch processing—is architected to avoid network I/O, subprocess execution, and external service dependencies unless the user explicitly consents via a separate runner.
Core Architecture of Offline-First Capture
Filesystem Confinement and Safety Primitives
The system establishes a security boundary using _safe_runtime_dir (lines 92-108 in claude_obsidian/capture.py), which creates a confined .vault-meta/capture directory using openat-style file descriptors that never follow symlinks. This prevents directory traversal attacks while ensuring all capture data remains locally scoped.
CaptureConfig (lines 31-48) validates that the inbox is a visible directory and the raw store uses safe relative paths, while CaptureBudget (lines 7-14) enforces hard limits on batch size. Before any file enters the pipeline, allowed_source_path (lines 85-108) guarantees the source is a regular file residing inside configured inbox directories, contains no symlinks, and uses safe filenames (NFC normalized, ≤240 bytes, no reserved Windows names).
Atomic Write Guarantees
To prevent corruption during power loss or concurrent access, the system implements _atomic_runtime_write (lines 53-81) using a write-then-rename pattern. The function creates a temporary file, calls fsync to flush data to disk, then uses os.replace to atomically move the entry into place. This ensures queue entries in .vault-meta/capture are never partially written, making the system resilient on laptops with unreliable power.
Content Identity and Metadata Without Network
Before processing, source_identity (lines 38-58) computes a stable SHA-256 hash without following leaf symlinks, verifying the file did not change during hashing. For type detection, sniff_metadata (lines 221-250) reads at most 64 KiB locally to infer MIME types (text, image, PDF, EPUB, media) without performing any content extraction or network requests.
The Two-Phase Capture Pipeline
Read-Only Planning with plan_filesystem_batch
The offline-first workflow separates planning from execution. plan_filesystem_batch (lines 145-176) builds a read-only capture plan that validates every source, checks budgets, computes SHA-256 identities, and determines whether new immutable copies are required—performing no mutations. This function operates entirely within the confined runtime directory and returns an in-memory plan safe to inspect without side effects.
from pathlib import Path
from claude_obsidian.capture import plan_filesystem_batch, CaptureConfig
vault = Path("/path/to/vault")
sources = [Path("inbox/doc1.md"), Path("inbox/image.png")]
# Build a read‑only plan – no files are written yet
plan = plan_filesystem_batch(vault, sources, config=CaptureConfig.load(vault))
for entry in plan:
print(f"{entry['source']} → {entry['stored_path']} (change? {entry['would_change']})")
The function validates each source, computes SHA‑256 identities, and decides whether a new immutable copy is required. No I/O beyond reading the source files occurs.
Immutable Execution with capture_filesystem_batch
Actual mutation occurs only through capture_filesystem_batch (lines 224-260), which turns plans into recoverable transactions via apply_bundle (implemented in claude_obsidian/transaction.py). This function writes only files whose content differs from existing copies, never deletes inbox files, and executes under a vault-wide advisory lock. The result is a concise, deterministic record of what changed.
from claude_obsidian.capture import capture_filesystem_batch
vault = Path("/path/to/vault")
sources = [Path("inbox/report.pdf"), Path("inbox/notes.txt")]
results = capture_filesystem_batch(vault, sources)
for r in results:
print(r["source"], "captured as", r["stored_path"], "changed:", r["changed"])
Only the files whose content differs from any existing copy are written into .vault-meta/capture/captured/<sha256>.
Isolating External Dependencies
Inert Command Plans via plan_external_action
Operations requiring network, OCR, transcription, or other external services are never executed directly. Instead, plan_external_action (lines 511-540) creates an inert argv plan containing the adapter ID, normalized source, and runner path—but never launches any process. These plans are stored as JSON in the capture queue and executed only after explicit user consent by a separate runner.
from claude_obsidian.capture import plan_external_action
action = plan_external_action(
adapter_id="ocr",
source="https://example.com/document.png",
approved_hosts=["example.com"],
runner=["/usr/local/bin/ocr-runner"]
)
print(action["command"])
# ['/usr/local/bin/ocr-runner', 'ocr', '--source', 'https://example.com/document.png']
The returned dict is an inert plan; the runner is not started here. The plan can be stored in the queue and later executed after the user gives explicit consent.
Strict URL Validation Without DNS Lookups
When external plans reference URLs, validate_https_url (lines 78-104) enforces strict HTTPS-only schemes, blocks private hosts, disallows credentials, and canonicalizes URLs without any DNS lookup. This prevents information leakage and ensures the validation remains purely computational and offline-safe.
Summary
- Filesystem confinement: All operations use
_safe_runtime_dirwithopenat-style descriptors that never follow symlinks, keeping data within.vault-meta/capture. - Atomic durability: Writes use temporary files with
fsyncandos.replaceto guarantee crash-safe queue entries. - Identity-based deduplication: SHA-256 hashes computed by
source_identityensure only changed content triggers new writes. - Plan-execution separation:
plan_filesystem_batchperforms zero side effects, whilecapture_filesystem_batchexecutes under advisory locks viatransaction.py. - Explicit external consent: Network-dependent operations use
plan_external_actionto create inert argv plans that require separate runner execution.
Frequently Asked Questions
How does claude-obsidian ensure data integrity during offline captures?
The system uses atomic write operations via _atomic_runtime_write (lines 53-81), which creates temporary files, flushes them to disk with fsync, and renames them into place using os.replace. Additionally, source_identity (lines 38-58) computes SHA-256 hashes to detect file modifications during processing, and the apply_bundle mechanism in transaction.py provides vault-wide advisory locking to prevent concurrent corruption.
What happens if a file changes during the capture process?
The source_identity function detects mid-read modifications by comparing file statistics before and after hashing. If a file changes between validation and capture, the hash mismatch triggers a CaptureValidationError, forcing the system to abort the specific entry and maintain store consistency without partial writes.
Why does the system use inert command plans for external actions?
Inert command plans created by plan_external_action (lines 511-540) ensure the core capture system never executes untrusted code, makes network requests, or spawns subprocesses without explicit user consent. This design maintains the offline-first guarantee while allowing extensibility—external adapters (OCR, transcription) receive structured argv plans that can be audited and executed only after the user approves the operation.
Can the capture system run on a machine with no internet connection?
Yes. All primitive operations in claude_obsidian/capture.py—including batch planning, content sniffing, SHA-256 identity calculation, and atomic writes—function entirely on the local filesystem. Network-dependent adapters are expressed as inert JSON plans that sit in the queue until connectivity returns or the user manually runs the external runner.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →