Ingest Workflow in Claude-Obsidian: 7-Stage Pipeline Explained
The ingest workflow in Claude-Obsidian is a seven-stage pipeline orchestrated by the wiki-ingest agent that atomically transforms raw web pages, PDFs, and notes into structured, source-cited knowledge within your vault.
The ingest workflow is the foundational mechanism of the AgriciDaniel/claude-obsidian repository that converts external source material into structured Obsidian-compatible knowledge. This workflow ensures atomicity, immutability, and provenance tracking throughout the entire process, from initial capture to final vault integration.
The Seven Stages of the Ingest Workflow
The ingest workflow operates through a sequence of distinct stages defined in agents/wiki-ingest.md and implemented across the core Python modules. Each stage handles a specific transformation from raw input to structured knowledge.
1. Capture
Raw content is collected via the capture module in claude_obsidian/capture.py. This stage stores immutable source payloads in the .raw/ directory, preserving original metadata and ensuring a permanent audit trail. The capture module handles URLs, PDFs, and local files, creating SHA-256 checksums for integrity verification.
2. Transaction Creation
A transaction is opened using claude_obsidian/transaction.py. The system records SHA-256 hashes of all target files, builds drafts for each change, and assembles them into a claude-obsidian.transaction.v1 bundle. This transaction wrapper ensures that the entire ingest operation succeeds or fails atomically.
3. Parsing and Normalization
The captured payload is parsed using language-model-driven think skills defined in the ingest agent. Content is normalized to Obsidian-flavoured Markdown, with front-matter YAML added for provenance tracking. This stage handles HTML-to-Markdown conversion, PDF text extraction, and structural normalization.
4. Ledger Update
Provenance information—source URL, timestamp, author, and checksum—is written to the ledger via claude_obsidian/ledgers.py. This creates a searchable claim record that enables source-cited knowledge bases and maintains citation chains for every ingested document.
5. Commit to Vault
The transaction bundle is applied atomically by the vault-ops module in claude_obsidian/vault_ops.py. Files are written to the wiki/ tree, indexes are refreshed, and a checkpoint is recorded. This stage ensures that partial writes cannot occur—either the complete ingest succeeds or the vault remains unchanged.
6. Post-Processing
After the commit, the lint engine (claude_obsidian/lint_engine.py) runs automatically to enforce formatting rules, generate backlinks, and update the hot-context (wiki/hot.md). This stage maintains markdown consistency and updates dynamic indexes that track recently ingested knowledge.
7. Finalization
A summary of the ingest operation is written to wiki/log.md, and the transaction is closed. The user can now query the new knowledge via wiki-query, wiki-mode, or the CLI. This stage completes the lifecycle and releases transaction locks.
Key Design Principles
The ingest workflow is built on four core architectural principles that distinguish it from simple file import tools:
-
Atomicity: All changes are bundled in a transaction managed by
claude_obsidian/transaction.py. Either the entire ingest succeeds or nothing is written to the vault, preventing partial or corrupted states. -
Immutability: Raw source files in the
.raw/directory are write-once. They never change after capture, ensuring a permanent audit trail and allowing re-processing from original sources if needed. -
Provenance: Ledgers in
claude_obsidian/ledgers.pycapture full citation data including URLs, timestamps, and checksums. This enables source-cited knowledge bases where every claim can be traced to its origin. -
Extensibility: New parsers or enrichers can be added as separate "skills" without modifying the core transaction logic in
claude_obsidian/transaction.py, allowing the workflow to adapt to new content types.
Implementation Examples
You can invoke the ingest workflow programmatically using the Python API or through the command-line interface.
Python API
The following example demonstrates the complete workflow using the core modules:
from claude_obsidian.capture import capture_url
from claude_obsidian.transaction import Transaction
from claude_obsidian.vault_ops import apply_transaction
# 1. Capture a web page
raw_path = capture_url("https://example.com/article")
# 2. Start a transaction
with Transaction() as txn:
# 3. Add the captured payload as a draft
txn.add_raw(raw_path)
# 4. Let the ingest skill parse & normalise
# (the skill is invoked internally via transaction hooks)
# 5. Apply the transaction atomically
apply_transaction(txn)
The Transaction context manager handles the transaction lifecycle, while apply_transaction() commits the bundle to the vault via claude_obsidian/vault_ops.py.
Command-Line Interface
For direct usage without Python scripting, the CLI provides a single-command interface:
claude-obsidian wiki-ingest https://example.com/article
This command automatically:
- Captures the URL using
claude_obsidian/capture.py - Opens and manages a transaction
- Runs the ingest skill from
agents/wiki-ingest.md - Commits changes atomically
- Updates the hot-context and runs lint checks via
claude_obsidian/lint_engine.py
Summary
The ingest workflow in Claude-Obsidian provides a robust, transactional pipeline for knowledge management:
- Seven stages from capture to finalization ensure complete provenance and data integrity
- Atomic transactions prevent partial writes and maintain vault consistency
- Immutable raw storage in
.raw/preserves original sources for auditing - Ledger integration creates searchable, citeable knowledge records
- Multiple interfaces support both Python API usage and CLI workflows
This architecture allows users to build trusted knowledge bases where every piece of information can be traced back to its original source through the ledger system implemented in claude_obsidian/ledgers.py.
Frequently Asked Questions
How does the ingest workflow ensure data integrity?
The ingest workflow ensures data integrity through transactional atomicity implemented in claude_obsidian/transaction.py. All changes are bundled into a claude-obsidian.transaction.v1 bundle with SHA-256 hashes. The vault_ops.py module applies these bundles atomically—either the entire ingest succeeds or the vault remains unchanged. Additionally, the .raw/ directory stores immutable copies of source files, creating a permanent audit trail that cannot be altered after initial capture.
What file formats does the ingest workflow support?
According to the agents/wiki-ingest.md agent definition and claude_obsidian/capture.py implementation, the ingest workflow supports web pages (HTML), PDF documents, and plain text files. The parsing stage uses language-model-driven skills to normalize HTML to Obsidian-flavoured Markdown and extract text from PDFs. New formats can be added through the extensible "skills" architecture without modifying core transaction logic.
Can I use the ingest workflow programmatically without the CLI?
Yes, you can invoke the ingest workflow entirely through the Python API without using the command line. Import capture_url from claude_obsidian/capture.py, wrap operations in a Transaction context manager from claude_obsidian/transaction.py, and commit via apply_transaction() from claude_obsidian/vault_ops.py. This programmatic approach is documented in the methodology guide and allows integration with custom automation scripts.
Where are raw sources stored during the ingest process?
Raw sources are stored in the .raw/ directory within your vault structure, as implemented in claude_obsidian/capture.py. This directory operates on a write-once principle: files are immutable after capture, preserving original metadata and checksums. The ledger system in claude_obsidian/ledgers.py maintains references to these raw files, ensuring you can always trace processed knowledge back to its original source payload.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →