How Auto-Fetch and Memory Source Sync Integrate Data into OpenHuman's Memory Tree

OpenHuman ingests both ad-hoc web data and synchronized source content into a unified memory tree through two convergent pathways: the WebFetchTool for on-demand retrieval and the memory-source sync loop for periodic background ingestion.

The tinyhumansai/openhuman repository implements a hierarchical memory tree architecture that indexes all user-generated and imported information as queryable chunks. Understanding how auto-fetch and memory source sync integrate data into OpenHuman's memory tree requires examining the distinct ingestion pipelines that ultimately converge on the same SQLite-backed storage layer.

The Memory Tree Architecture

OpenHuman persists all contextual data in a memory tree—a hierarchical index of Chunk objects defined in memory::types::Chunk. This structure supports both real-time agent queries and UI components. Whether data arrives via an on-demand web fetch or a scheduled Gmail synchronization, it must pass through the ingestion pipeline in src/openhuman/memory/ops/sync.rs and ultimately materialize in the tree via src/openhuman/memory/tree/ops.rs.

Auto-Fetch: On-Demand Data Ingestion

The auto-fetch mechanism operates through the WebFetchTool and related fetch-type tools, allowing agents to pull remote content dynamically without pre-configured source integration.

Tool Invocation and Security

When an agent invokes a fetch operation, the system first constructs a tool instance with permission checks. In src/openhuman/tools/ops.rs, the WebFetchTool::new function initializes the tool with concurrency limits and URL allow-lists. The actual network request executes in src/openhuman/tools/impl/network/http_request.rs, which validates the URL and explicitly blocks loop-back or private hosts—security constraints verified in tests/raw_coverage/tools_approval_channels_raw_coverage_e2e.rs.

Content Parsing and Chunk Creation

Fetched content flows into src/openhuman/tools/impl/web_fetch.rs, which extracts readable text, strips scripts, and truncates output to the tool’s max_result_size_chars. The purified text then moves to src/openhuman/memory/ops/sync.rs, where the create_chunk function wraps it in a Chunk struct. Critically, auto-fetched content receives a source-ID of None because it operates outside the persistent source-sync registry.

Tree Insertion

The final step calls insert_chunk in src/openhuman/memory/tree/ops.rs, which persists the chunk to the SQLite-backed memory store and updates the memory-tree index. Despite lacking a source ID, the chunk remains fully searchable via the standard query APIs.

Memory-Source Sync: Background Synchronization

Memory-source sync provides incremental, scheduled ingestion for persistent data sources such as Gmail, Slack, or Composio integrations. This process runs independently of agent actions through a background job system.

Scheduler and Triggering

The sync lifecycle begins in src/openhuman/cron/scheduler_gate.rs, where the cron::scheduler_gate triggers jobs based on each source's configured sync_interval or explicit user requests. Each configured source follows the schema defined in openhuman/memory/sources/schemas.rs.

Driver Resolution and RPC Execution

For each scheduled source, the system invokes driver_for_source in src/openhuman/memory/sync/composio/bus.rs to resolve the appropriate memory driver from the memory_driver_registry. The driver executes RPC calls defined in src/openhuman/memory/sources/rpc.rs, such as sync_audit_log or estimate_sync_cost, returning a list of new or updated chunks along with a concrete source_id. If a driver lacks sync support, the operation raises "the bound memory driver … does not serve source sync" errors as handled in src/openhuman/memory/ops/sync.rs.

Chunk Import and Tree Reconciliation

Returned payloads enter import_chunks in src/openhuman/memory/ops/sync.rs, which converts driver data into Chunk objects with attached source_id values. These chunks then flow to reconcile_tree in src/openhuman/memory/tree/ops.rs, which merges new data into the existing hierarchy, updates freshness labels, and prunes stale entries. Completion triggers a MemorySourceSyncCompleted event in src/core/events.rs, allowing UI components to display real-time progress.

Convergence in the Memory Tree

Both ingestion pathways ultimately execute insert_chunk or import_chunks, writing Chunk records to the SQLite database (memory::store::chunks). The tree index in memory::tree recomputes regardless of origin, ensuring all content appears under a unified hierarchical view.

Key integration points include:

  • Source-ID handling: Auto-fetch uses source_id = None while source-sync sets concrete identifiers validated in memory/sync_events_bridge.rs.
  • Chunk metadata: Both paths populate ChunkMetadata with timestamps, size, and provenance data, enabling UI labels like "last synced 5 minutes ago".
  • Event bus: The MemorySourceSyncCompleted event published in src/core/events.rs notifies listeners across the stack, including front-end components.

Implementation Examples

Triggering Auto-Fetch from an Agent

let fetch = openhuman_core::openhuman::tools::WebFetchTool::new(
    security,
    vec!["*".into()],   // allow list
    Some(0),            // no concurrency limit
    Some(50_000),      // max result size
);
let result = fetch.invoke(json!({ "url": "https://example.com/article" })).await?;

The fetched text becomes a new memory chunk without a source ID, immediately available for agent retrieval.

Manually Starting a Source Sync

use openhuman_core::openhuman::memory::ops::sync_source;
let source_id = "gmail-account-1".to_string();
sync_source(&core, &source_id).await?;

This routine contacts the Gmail driver, imports new email chunks, and refreshes the memory tree via reconcile_tree.

Listening for Sync Completion in the UI

import { socketService } from "@/services/socketService";

socketService.on("MemorySourceSyncCompleted", (event) => {
  console.log(`Sync finished for ${event.sourceId}: ${event.newChunks} new chunks`);
});

The front-end receives events published by src/core/events.rs, enabling progress indicators and notification badges.

Summary

  • Auto-fetch provides ad-hoc data retrieval through WebFetchTool, bypassing persistent source configuration but still populating the memory tree via insert_chunk.
  • Memory-source sync runs scheduled background jobs through src/openhuman/cron/scheduler_gate.rs and src/openhuman/modules/ops.rs, using driver-specific RPCs to bulk-import data with attached source IDs.
  • Both pathways converge in src/openhuman/memory/ops/sync.rs and src/openhuman/memory/tree/ops.rs, ensuring unified indexing and queryability regardless of ingestion method.
  • The system emits MemorySourceSyncCompleted events for real-time UI updates, while SQLite persistence guarantees durability across both auto-fetched and synchronized content.

Frequently Asked Questions

What is the difference between auto-fetch and memory source sync in OpenHuman?

Auto-fetch operates on-demand through agent tools like WebFetchTool, pulling arbitrary URLs without persistent source configuration and assigning source_id = None. Memory-source sync runs periodic background jobs for configured integrations (Gmail, Slack, etc.), maintains persistent source_id associations, and supports incremental updates through driver-specific RPC calls.

How does OpenHuman handle metadata for memory chunks?

All chunks populate a ChunkMetadata structure containing timestamps, content size, and provenance information. According to src/openhuman/memory/ops/sync.rs, both auto-fetch and source-sync populate these fields, enabling the UI to display freshness indicators and source attribution regardless of ingestion pathway.

Can auto-fetched content be distinguished from synchronized source data?

Yes. Auto-fetched chunks explicitly set source_id = None during creation in src/openhuman/memory/ops/sync.rs, while source-sync operations attach concrete source identifiers validated through memory/sync_events_bridge.rs. This distinction allows filtering queries by source provenance while maintaining unified tree indexing.

What happens when a memory source sync fails?

Errors during driver resolution or RPC execution—such as "the bound memory driver … does not serve source sync"—are raised in src/openhuman/memory/ops/sync.rs and propagated through the composio bus in src/openhuman/memory/sync/composio/bus.rs. Failed syncs do not update the memory tree, and error states are emitted through the event bus for UI handling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →