# How Auto-Fetch and Memory Source Sync Integrate Data into OpenHuman's Memory Tree

> Discover how OpenHuman integrates data into its memory tree using auto-fetch and memory source sync. Learn about WebFetchTool and background ingestion for unified data.

- Repository: [Tiny Humans/openhuman](https://github.com/tinyhumansai/openhuman)
- Tags: internals
- Published: 2026-08-29

---

**OpenHuman ingests both ad-hoc web data and synchronized source content into a unified memory tree through two convergent pathways: the WebFetchTool for on-demand retrieval and the memory-source sync loop for periodic background ingestion.**

The `tinyhumansai/openhuman` repository implements a hierarchical **memory tree** architecture that indexes all user-generated and imported information as queryable chunks. Understanding how auto-fetch and memory source sync integrate data into OpenHuman's memory tree requires examining the distinct ingestion pipelines that ultimately converge on the same SQLite-backed storage layer.

## The Memory Tree Architecture

OpenHuman persists all contextual data in a **memory tree**—a hierarchical index of `Chunk` objects defined in `memory::types::Chunk`. This structure supports both real-time agent queries and UI components. Whether data arrives via an on-demand web fetch or a scheduled Gmail synchronization, it must pass through the ingestion pipeline in [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs) and ultimately materialize in the tree via [`src/openhuman/memory/tree/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/ops.rs).

## Auto-Fetch: On-Demand Data Ingestion

The **auto-fetch** mechanism operates through the `WebFetchTool` and related fetch-type tools, allowing agents to pull remote content dynamically without pre-configured source integration.

### Tool Invocation and Security

When an agent invokes a fetch operation, the system first constructs a tool instance with permission checks. In [`src/openhuman/tools/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/ops.rs), the `WebFetchTool::new` function initializes the tool with concurrency limits and URL allow-lists. The actual network request executes in [`src/openhuman/tools/impl/network/http_request.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/network/http_request.rs), which validates the URL and explicitly blocks loop-back or private hosts—security constraints verified in [`tests/raw_coverage/tools_approval_channels_raw_coverage_e2e.rs`](https://github.com/tinyhumansai/openhuman/blob/main/tests/raw_coverage/tools_approval_channels_raw_coverage_e2e.rs).

### Content Parsing and Chunk Creation

Fetched content flows into [`src/openhuman/tools/impl/web_fetch.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/tools/impl/web_fetch.rs), which extracts readable text, strips scripts, and truncates output to the tool’s `max_result_size_chars`. The purified text then moves to [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs), where the `create_chunk` function wraps it in a `Chunk` struct. Critically, auto-fetched content receives a **source-ID of `None`** because it operates outside the persistent source-sync registry.

### Tree Insertion

The final step calls `insert_chunk` in [`src/openhuman/memory/tree/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/ops.rs), which persists the chunk to the SQLite-backed memory store and updates the memory-tree index. Despite lacking a source ID, the chunk remains fully searchable via the standard query APIs.

## Memory-Source Sync: Background Synchronization

**Memory-source sync** provides incremental, scheduled ingestion for persistent data sources such as Gmail, Slack, or Composio integrations. This process runs independently of agent actions through a background job system.

### Scheduler and Triggering

The sync lifecycle begins in [`src/openhuman/cron/scheduler_gate.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/cron/scheduler_gate.rs), where the `cron::scheduler_gate` triggers jobs based on each source's configured `sync_interval` or explicit user requests. Each configured source follows the schema defined in [`openhuman/memory/sources/schemas.rs`](https://github.com/tinyhumansai/openhuman/blob/main/openhuman/memory/sources/schemas.rs).

### Driver Resolution and RPC Execution

For each scheduled source, the system invokes `driver_for_source` in [`src/openhuman/memory/sync/composio/bus.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/sync/composio/bus.rs) to resolve the appropriate memory driver from the `memory_driver_registry`. The driver executes RPC calls defined in [`src/openhuman/memory/sources/rpc.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/sources/rpc.rs), such as `sync_audit_log` or `estimate_sync_cost`, returning a list of new or updated chunks along with a concrete `source_id`. If a driver lacks sync support, the operation raises "the bound memory driver … does not serve source sync" errors as handled in [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs).

### Chunk Import and Tree Reconciliation

Returned payloads enter `import_chunks` in [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs), which converts driver data into `Chunk` objects with attached `source_id` values. These chunks then flow to `reconcile_tree` in [`src/openhuman/memory/tree/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/ops.rs), which merges new data into the existing hierarchy, updates freshness labels, and prunes stale entries. Completion triggers a `MemorySourceSyncCompleted` event in [`src/core/events.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/core/events.rs), allowing UI components to display real-time progress.

## Convergence in the Memory Tree

Both ingestion pathways ultimately execute `insert_chunk` or `import_chunks`, writing `Chunk` records to the SQLite database (`memory::store::chunks`). The tree index in `memory::tree` recomputes regardless of origin, ensuring **all content appears under a unified hierarchical view**.

Key integration points include:

- **Source-ID handling**: Auto-fetch uses `source_id = None` while source-sync sets concrete identifiers validated in [`memory/sync_events_bridge.rs`](https://github.com/tinyhumansai/openhuman/blob/main/memory/sync_events_bridge.rs).
- **Chunk metadata**: Both paths populate `ChunkMetadata` with timestamps, size, and provenance data, enabling UI labels like "last synced 5 minutes ago".
- **Event bus**: The `MemorySourceSyncCompleted` event published in [`src/core/events.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/core/events.rs) notifies listeners across the stack, including front-end components.

## Implementation Examples

### Triggering Auto-Fetch from an Agent

```rust
let fetch = openhuman_core::openhuman::tools::WebFetchTool::new(
    security,
    vec!["*".into()],   // allow list
    Some(0),            // no concurrency limit
    Some(50_000),      // max result size
);
let result = fetch.invoke(json!({ "url": "https://example.com/article" })).await?;

```

The fetched text becomes a new memory chunk without a source ID, immediately available for agent retrieval.

### Manually Starting a Source Sync

```rust
use openhuman_core::openhuman::memory::ops::sync_source;
let source_id = "gmail-account-1".to_string();
sync_source(&core, &source_id).await?;

```

This routine contacts the Gmail driver, imports new email chunks, and refreshes the memory tree via `reconcile_tree`.

### Listening for Sync Completion in the UI

```typescript
import { socketService } from "@/services/socketService";

socketService.on("MemorySourceSyncCompleted", (event) => {
  console.log(`Sync finished for ${event.sourceId}: ${event.newChunks} new chunks`);
});

```

The front-end receives events published by [`src/core/events.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/core/events.rs), enabling progress indicators and notification badges.

## Summary

- **Auto-fetch** provides ad-hoc data retrieval through `WebFetchTool`, bypassing persistent source configuration but still populating the memory tree via `insert_chunk`.
- **Memory-source sync** runs scheduled background jobs through [`src/openhuman/cron/scheduler_gate.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/cron/scheduler_gate.rs) and [`src/openhuman/modules/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/modules/ops.rs), using driver-specific RPCs to bulk-import data with attached source IDs.
- Both pathways converge in [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs) and [`src/openhuman/memory/tree/ops.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/tree/ops.rs), ensuring unified indexing and queryability regardless of ingestion method.
- The system emits `MemorySourceSyncCompleted` events for real-time UI updates, while SQLite persistence guarantees durability across both auto-fetched and synchronized content.

## Frequently Asked Questions

### What is the difference between auto-fetch and memory source sync in OpenHuman?

**Auto-fetch** operates on-demand through agent tools like `WebFetchTool`, pulling arbitrary URLs without persistent source configuration and assigning `source_id = None`. **Memory-source sync** runs periodic background jobs for configured integrations (Gmail, Slack, etc.), maintains persistent `source_id` associations, and supports incremental updates through driver-specific RPC calls.

### How does OpenHuman handle metadata for memory chunks?

All chunks populate a `ChunkMetadata` structure containing timestamps, content size, and provenance information. According to [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs), both auto-fetch and source-sync populate these fields, enabling the UI to display freshness indicators and source attribution regardless of ingestion pathway.

### Can auto-fetched content be distinguished from synchronized source data?

Yes. Auto-fetched chunks explicitly set `source_id = None` during creation in [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs), while source-sync operations attach concrete source identifiers validated through [`memory/sync_events_bridge.rs`](https://github.com/tinyhumansai/openhuman/blob/main/memory/sync_events_bridge.rs). This distinction allows filtering queries by source provenance while maintaining unified tree indexing.

### What happens when a memory source sync fails?

Errors during driver resolution or RPC execution—such as "the bound memory driver … does not serve source sync"—are raised in [`src/openhuman/memory/ops/sync.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/ops/sync.rs) and propagated through the composio bus in [`src/openhuman/memory/sync/composio/bus.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/memory/sync/composio/bus.rs). Failed syncs do not update the memory tree, and error states are emitted through the event bus for UI handling.