Memory Tree and Obsidian Wiki Synchronization in OpenHuman

OpenHuman achieves memory-tree and Obsidian wiki synchronization by running every ingested item through a deterministic, bucket-sealed pipeline that writes scored, vector-embedded markdown chunks to both a local SQLite database and a live Obsidian vault folder.

The tinyhumansai/openhuman repository implements this dual-write system in Rust, keeping the vault readable and editable while the database handles fast vector retrieval. The architecture is deliberately flavor-agnostic: the same generic tree engine that powers the Obsidian view can support future output formats without changing the core ingestion logic.

The Memory Tree Pipeline

The generic tree engine lives under src/openhuman/memory/tree/ and processes every source item—whether it originates from email, chat, documents, or integration syncs—through a fixed sequence of stages.

Ingest and Chunk

A source connector produces a stream of raw items. Each item is split into markdown leaves capped at roughly three thousand tokens. This chunking logic is implemented in src/openhuman/memory/tree/tree.rs. The resulting leaves preserve provenance metadata so that every chunk can be traced back to its origin.

Score, Embed, and Seal

A lightweight scorer decides whether a chunk is worth keeping. Retained chunks are passed to the embedding model configured under config.memory_tree.embedding_model, with inference handled in src/openhuman/inference/embeddings/ollama.rs.

Leaves are accumulated in an L0 buffer. Once the buffer crosses a size threshold, it is sealed and a summarization LLM generates a parent summary node. This operation is defined in src/openhuman/tree_summarizer/ops.rs. The summary becomes the next-level bucket, and the process repeats upward to build a summary forest spanning daily, topic, and global horizons.

Dual Persistence to SQLite and the Vault

When a bucket is sealed, the pipeline persists data in two places simultaneously. The operation is coordinated in src/openhuman/memory/tree/tree_runtime/ops.rs.

  • SQLite database at <workspace>/memory_tree/chunks.db stores raw chunk data, embedding vectors, scoring metadata, and the entity index. This supports fast vector search during agent retrieval.
  • Obsidian vault at <workspace>/wiki/ stores human-readable markdown files, one per chunk and per summary node. The file name encodes the node ID, and frontmatter stores provenance such as source, timestamp, and entity IDs.

Because both stores receive identical chunks, the system executes read-only retrieval against the database for speed while presenting full provenance to the user through vault files.

Obsidian Vault Registration and Launch

Opening the vault from the OpenHuman UI requires that the wiki folder be registered in Obsidian’s own vault registry. Host-side detection logic is implemented in src/openhuman/memory/obsidian_registry.rs.

The detector searches common configuration locations—including ~/.config/obsidian/obsidian.json and the Flatpak-equivalent path—for the list of registered vaults. It then checks whether <workspace>/wiki/ or any ancestor directory appears in that list. The result is returned as an ObsidianVaultStatusResponse via the read-RPC endpoint openhuman.memory_tree_obsidian_vault_status, defined in src/openhuman/memory/read_rpc/vault.rs.

If the folder is not yet registered, the UI displays an inline guide that instructs the user to add the folder as a vault inside Obsidian. Once registered, clicking “View vault in Obsidian” fires the deep-link obsidian://open?path=…, which the operating system hands off to the Obsidian application.

Auto-Fetch Background Synchronization

Every twenty minutes, a background auto-fetch job runs inside src/openhuman/memory/sync/mod.rs. For each enabled connector, the job pulls new items since the last cursor and feeds them into the Memory Tree pipeline. Because the pipeline writes markdown files into the vault as part of its standard persist step, the Obsidian vault remains a live mirror of the database without a separate export phase.

Progress and errors are exposed through the RPC openhuman.memory_tree_sync_status. You can also trigger the same pipeline manually through the openhuman.memory_tree_sync_now RPC, which invokes the scheduler logic directly.

Agent Retrieval and Vault Provenance

When the LLM requires context, the memory tools—recall, search, and drill_down—invoke the retrieval layer inside src/openhuman/memory/tree/retrieval/. The retrieval logic executes a vector similarity search against the SQLite database, falling back to exact-match search if embeddings are disabled.

After identifying relevant chunk IDs, the layer resolves each ID to its corresponding markdown file in the Obsidian vault and attaches the file path as a provenance citation. Consequently, every answer the agent produces can be traced to a concrete .md file that the user can open, edit, or delete.

Edits made inside Obsidian are not lost. On the next ingest run, the pipeline treats the vault as the source of truth for manual notes. Any file present in the vault that has not yet been indexed is ingested as a user note chunk through the same src/openhuman/memory/tree/tree_runtime/ops.rs path.

Code Examples

Trigger a Recall via RPC

The openhuman.memory_tree_recall RPC is defined in src/openhuman/memory/tree/retrieval/rpc.rs. The following Rust client requests the five most recent chunks about “Alice”:

use openhuman_core::client::CoreRpcClient;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let client = CoreRpcClient::new("http://127.0.0.1:3000/rpc")?;
    let result = client
        .call("openhuman.memory_tree_recall", json!({
            "query": "Alice",
            "limit": 5,
            "source_scoped": false
        }))
        .await?;
    println!("{:#?}", result);
    Ok(())
}

Check Vault Registration via RPC

The response type ObsidianVaultStatusResponse is declared in src/openhuman/memory/read_rpc/types.rs:

let status = client
    .call("openhuman.memory_tree_obsidian_vault_status", json!({}))
    .await?;
println!("Vault registered: {}", status["registered"]);

Manually Add a Markdown Note

Creating a file inside the vault folder causes the next sync run to ingest it automatically:

cd ~/OpenHuman/projects/my-assistant/wiki
cat > manual-note.md <<'EOF'
---
id: note-001
kind: manual
display_name: "My custom note"
---
This is a free-form note. The next sync run will ingest it as a chunk.
EOF

The ingest path in src/openhuman/memory/tree/tree_runtime/ops.rs indexes this file on the following auto-fetch cycle.

Force a Sync Manually

To trigger the background job immediately, call the RPC defined in src/openhuman/memory/sync/mod.rs:

client
    .call("openhuman.memory_tree_sync_now", json!({}))
    .await?;

Summary

  • The Memory Tree pipeline is a deterministic, bucket-sealed engine that converts raw items into scored, vector-embedded markdown chunks.
  • Sealed chunks are written simultaneously to <workspace>/memory_tree/chunks.db and the <workspace>/wiki/ Obsidian vault by src/openhuman/memory/tree/tree_runtime/ops.rs.
  • Vault registration is detected by scanning obsidian.json in src/openhuman/memory/obsidian_registry.rs, and launch is handled via the obsidian:// deep-link protocol.
  • The auto-fetch scheduler in src/openhuman/memory/sync/mod.rs runs every twenty minutes to pull new connector data and refresh both stores.
  • Agent retrieval in src/openhuman/memory/tree/retrieval/ queries SQLite for speed, then maps results back to vault file paths so every citation is user-inspectable.

Frequently Asked Questions

How does OpenHuman keep the SQLite database and Obsidian vault in sync?

The pipeline in src/openhuman/memory/tree/tree_runtime/ops.rs writes every sealed chunk to both destinations during the same commit phase. There is no secondary sync process; the vault is produced inline as the pipeline runs. Manual edits in the vault are re-ingested as user notes on the next auto-fetch cycle.

What happens if the Obsidian vault folder is not registered?

The host-side detector in src/openhuman/memory/obsidian_registry.rs returns an unregistered status through the openhuman.memory_tree_obsidian_vault_status RPC. The UI then displays instructions for adding <workspace>/wiki/ to Obsidian’s vault list before enabling the deep-link launch button.

Can I query the Memory Tree without opening Obsidian?

Yes. Read-only retrieval is performed against the SQLite database via the recall, search, and drill_down tools in src/openhuman/memory/tree/retrieval/. These tools execute vector similarity search against chunks.db and only resolve file paths when building citations, so Obsidian does not need to be running.

What embedding model does the Memory Tree use?

The embedding model is configurable via config.memory_tree.embedding_model. The default inference path is implemented in src/openhuman/inference/embeddings/ollama.rs, which generates vectors for chunks that pass the lightweight scorer before they enter the L0 buffer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →