# Memory Tree and Obsidian Wiki Synchronization in OpenHuman

> Discover how OpenHuman synchronizes memory-tree and Obsidian wikis. Learn about its deterministic pipeline, vector embeddings, and SQLite integration for seamless data management.

- Repository: [Tiny Humans/openhuman](https://github.com/tinyhumansai/openhuman)
- Tags: internals
- Published: 2026-08-28

---

**OpenHuman achieves memory-tree and Obsidian wiki synchronization by running every ingested item through a deterministic, bucket-sealed pipeline that writes scored, vector-embedded markdown chunks to both a local SQLite database and a live Obsidian vault folder.**

The `tinyhumansai/openhuman` repository implements this dual-write system in Rust, keeping the vault readable and editable while the database handles fast vector retrieval. The architecture is deliberately flavor-agnostic: the same generic tree engine that powers the Obsidian view can support future output formats without changing the core ingestion logic.

## The Memory Tree Pipeline

The generic tree engine lives under `src/openhuman/memory/tree/` and processes every source item—whether it originates from email, chat, documents, or integration syncs—through a fixed sequence of stages.

### Ingest and Chunk

A **source connector** produces a stream of raw items. Each item is split into markdown leaves capped at roughly three thousand tokens. This chunking logic is implemented in [`src/openhuman/memory/tree/tree.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/tree.rs). The resulting leaves preserve provenance metadata so that every chunk can be traced back to its origin.

### Score, Embed, and Seal

A lightweight scorer decides whether a chunk is worth keeping. Retained chunks are passed to the embedding model configured under `config.memory_tree.embedding_model`, with inference handled in [`src/openhuman/inference/embeddings/ollama.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/inference/embeddings/ollama.rs).

Leaves are accumulated in an **L0 buffer**. Once the buffer crosses a size threshold, it is sealed and a summarization LLM generates a parent summary node. This operation is defined in [`src/openhuman/tree_summarizer/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tree_summarizer/ops.rs). The summary becomes the next-level bucket, and the process repeats upward to build a **summary forest** spanning daily, topic, and global horizons.

### Dual Persistence to SQLite and the Vault

When a bucket is sealed, the pipeline persists data in two places simultaneously. The operation is coordinated in [`src/openhuman/memory/tree/tree_runtime/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/tree_runtime/ops.rs).

- **SQLite database** at `<workspace>/memory_tree/chunks.db` stores raw chunk data, embedding vectors, scoring metadata, and the entity index. This supports fast vector search during agent retrieval.
- **Obsidian vault** at `<workspace>/wiki/` stores human-readable markdown files, one per chunk and per summary node. The file name encodes the node ID, and frontmatter stores provenance such as source, timestamp, and entity IDs.

Because both stores receive identical chunks, the system executes read-only retrieval against the database for speed while presenting full provenance to the user through vault files.

## Obsidian Vault Registration and Launch

Opening the vault from the OpenHuman UI requires that the wiki folder be registered in Obsidian’s own vault registry. Host-side detection logic is implemented in [`src/openhuman/memory/obsidian_registry.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/obsidian_registry.rs).

The detector searches common configuration locations—including `~/.config/obsidian/obsidian.json` and the Flatpak-equivalent path—for the list of registered vaults. It then checks whether `<workspace>/wiki/` or any ancestor directory appears in that list. The result is returned as an `ObsidianVaultStatusResponse` via the read-RPC endpoint `openhuman.memory_tree_obsidian_vault_status`, defined in [`src/openhuman/memory/read_rpc/vault.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/read_rpc/vault.rs).

If the folder is not yet registered, the UI displays an inline guide that instructs the user to add the folder as a vault inside Obsidian. Once registered, clicking **“View vault in Obsidian”** fires the deep-link `obsidian://open?path=…`, which the operating system hands off to the Obsidian application.

## Auto-Fetch Background Synchronization

Every twenty minutes, a background auto-fetch job runs inside [`src/openhuman/memory/sync/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/sync/mod.rs). For each enabled connector, the job pulls new items since the last cursor and feeds them into the Memory Tree pipeline. Because the pipeline writes markdown files into the vault as part of its standard persist step, the Obsidian vault remains a live mirror of the database without a separate export phase.

Progress and errors are exposed through the RPC `openhuman.memory_tree_sync_status`. You can also trigger the same pipeline manually through the `openhuman.memory_tree_sync_now` RPC, which invokes the scheduler logic directly.

## Agent Retrieval and Vault Provenance

When the LLM requires context, the memory tools—`recall`, `search`, and `drill_down`—invoke the retrieval layer inside `src/openhuman/memory/tree/retrieval/`. The retrieval logic executes a **vector similarity search** against the SQLite database, falling back to exact-match search if embeddings are disabled.

After identifying relevant chunk IDs, the layer resolves each ID to its corresponding markdown file in the Obsidian vault and attaches the file path as a provenance citation. Consequently, every answer the agent produces can be traced to a concrete `.md` file that the user can open, edit, or delete.

Edits made inside Obsidian are not lost. On the next ingest run, the pipeline treats the vault as the source of truth for manual notes. Any file present in the vault that has not yet been indexed is ingested as a **user note** chunk through the same [`src/openhuman/memory/tree/tree_runtime/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/tree_runtime/ops.rs) path.

## Code Examples

### Trigger a Recall via RPC

The `openhuman.memory_tree_recall` RPC is defined in [`src/openhuman/memory/tree/retrieval/rpc.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/retrieval/rpc.rs). The following Rust client requests the five most recent chunks about “Alice”:

```rust
use openhuman_core::client::CoreRpcClient;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let client = CoreRpcClient::new("http://127.0.0.1:3000/rpc")?;
    let result = client
        .call("openhuman.memory_tree_recall", json!({
            "query": "Alice",
            "limit": 5,
            "source_scoped": false
        }))
        .await?;
    println!("{:#?}", result);
    Ok(())
}

```

### Check Vault Registration via RPC

The response type `ObsidianVaultStatusResponse` is declared in [`src/openhuman/memory/read_rpc/types.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/read_rpc/types.rs):

```rust
let status = client
    .call("openhuman.memory_tree_obsidian_vault_status", json!({}))
    .await?;
println!("Vault registered: {}", status["registered"]);

```

### Manually Add a Markdown Note

Creating a file inside the vault folder causes the next sync run to ingest it automatically:

```bash
cd ~/OpenHuman/projects/my-assistant/wiki
cat > manual-note.md <<'EOF'
---
id: note-001
kind: manual
display_name: "My custom note"
---
This is a free-form note. The next sync run will ingest it as a chunk.
EOF

```

The ingest path in [`src/openhuman/memory/tree/tree_runtime/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/tree_runtime/ops.rs) indexes this file on the following auto-fetch cycle.

### Force a Sync Manually

To trigger the background job immediately, call the RPC defined in [`src/openhuman/memory/sync/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/sync/mod.rs):

```rust
client
    .call("openhuman.memory_tree_sync_now", json!({}))
    .await?;

```

## Summary

- The **Memory Tree** pipeline is a deterministic, bucket-sealed engine that converts raw items into scored, vector-embedded markdown chunks.
- Sealed chunks are written simultaneously to `<workspace>/memory_tree/chunks.db` and the `<workspace>/wiki/` Obsidian vault by [`src/openhuman/memory/tree/tree_runtime/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/tree_runtime/ops.rs).
- Vault registration is detected by scanning [`obsidian.json`](https://github.com/tinyhumansai/openhuman/blob/main/obsidian.json) in [`src/openhuman/memory/obsidian_registry.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/obsidian_registry.rs), and launch is handled via the `obsidian://` deep-link protocol.
- The auto-fetch scheduler in [`src/openhuman/memory/sync/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/sync/mod.rs) runs every twenty minutes to pull new connector data and refresh both stores.
- Agent retrieval in `src/openhuman/memory/tree/retrieval/` queries SQLite for speed, then maps results back to vault file paths so every citation is user-inspectable.

## Frequently Asked Questions

### How does OpenHuman keep the SQLite database and Obsidian vault in sync?

The pipeline in [`src/openhuman/memory/tree/tree_runtime/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/tree_runtime/ops.rs) writes every sealed chunk to both destinations during the same commit phase. There is no secondary sync process; the vault is produced inline as the pipeline runs. Manual edits in the vault are re-ingested as user notes on the next auto-fetch cycle.

### What happens if the Obsidian vault folder is not registered?

The host-side detector in [`src/openhuman/memory/obsidian_registry.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/obsidian_registry.rs) returns an unregistered status through the `openhuman.memory_tree_obsidian_vault_status` RPC. The UI then displays instructions for adding `<workspace>/wiki/` to Obsidian’s vault list before enabling the deep-link launch button.

### Can I query the Memory Tree without opening Obsidian?

Yes. Read-only retrieval is performed against the SQLite database via the `recall`, `search`, and `drill_down` tools in `src/openhuman/memory/tree/retrieval/`. These tools execute vector similarity search against `chunks.db` and only resolve file paths when building citations, so Obsidian does not need to be running.

### What embedding model does the Memory Tree use?

The embedding model is configurable via `config.memory_tree.embedding_model`. The default inference path is implemented in [`src/openhuman/inference/embeddings/ollama.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/inference/embeddings/ollama.rs), which generates vectors for chunks that pass the lightweight scorer before they enter the L0 buffer.