How Auto-Fetch and Memory Source Sync Integrate Data into OpenHuman's Memory Tree
OpenHuman ingests both ad-hoc web data and synchronized source content into a unified memory tree through two convergent pathways: the WebFetchTool for on-demand retrieval and the memory-source sync loop for periodic background ingestion.
The tinyhumansai/openhuman repository implements a hierarchical memory tree architecture that indexes all user-generated and imported information as queryable chunks. Understanding how auto-fetch and memory source sync integrate data into OpenHuman's memory tree requires examining the distinct ingestion pipelines that ultimately converge on the same SQLite-backed storage layer.
The Memory Tree Architecture
OpenHuman persists all contextual data in a memory tree—a hierarchical index of Chunk objects defined in memory::types::Chunk. This structure supports both real-time agent queries and UI components. Whether data arrives via an on-demand web fetch or a scheduled Gmail synchronization, it must pass through the ingestion pipeline in src/openhuman/memory/ops/sync.rs and ultimately materialize in the tree via src/openhuman/memory/tree/ops.rs.
Auto-Fetch: On-Demand Data Ingestion
The auto-fetch mechanism operates through the WebFetchTool and related fetch-type tools, allowing agents to pull remote content dynamically without pre-configured source integration.
Tool Invocation and Security
When an agent invokes a fetch operation, the system first constructs a tool instance with permission checks. In src/openhuman/tools/ops.rs, the WebFetchTool::new function initializes the tool with concurrency limits and URL allow-lists. The actual network request executes in src/openhuman/tools/impl/network/http_request.rs, which validates the URL and explicitly blocks loop-back or private hosts—security constraints verified in tests/raw_coverage/tools_approval_channels_raw_coverage_e2e.rs.
Content Parsing and Chunk Creation
Fetched content flows into src/openhuman/tools/impl/web_fetch.rs, which extracts readable text, strips scripts, and truncates output to the tool’s max_result_size_chars. The purified text then moves to src/openhuman/memory/ops/sync.rs, where the create_chunk function wraps it in a Chunk struct. Critically, auto-fetched content receives a source-ID of None because it operates outside the persistent source-sync registry.
Tree Insertion
The final step calls insert_chunk in src/openhuman/memory/tree/ops.rs, which persists the chunk to the SQLite-backed memory store and updates the memory-tree index. Despite lacking a source ID, the chunk remains fully searchable via the standard query APIs.
Memory-Source Sync: Background Synchronization
Memory-source sync provides incremental, scheduled ingestion for persistent data sources such as Gmail, Slack, or Composio integrations. This process runs independently of agent actions through a background job system.
Scheduler and Triggering
The sync lifecycle begins in src/openhuman/cron/scheduler_gate.rs, where the cron::scheduler_gate triggers jobs based on each source's configured sync_interval or explicit user requests. Each configured source follows the schema defined in openhuman/memory/sources/schemas.rs.
Driver Resolution and RPC Execution
For each scheduled source, the system invokes driver_for_source in src/openhuman/memory/sync/composio/bus.rs to resolve the appropriate memory driver from the memory_driver_registry. The driver executes RPC calls defined in src/openhuman/memory/sources/rpc.rs, such as sync_audit_log or estimate_sync_cost, returning a list of new or updated chunks along with a concrete source_id. If a driver lacks sync support, the operation raises "the bound memory driver … does not serve source sync" errors as handled in src/openhuman/memory/ops/sync.rs.
Chunk Import and Tree Reconciliation
Returned payloads enter import_chunks in src/openhuman/memory/ops/sync.rs, which converts driver data into Chunk objects with attached source_id values. These chunks then flow to reconcile_tree in src/openhuman/memory/tree/ops.rs, which merges new data into the existing hierarchy, updates freshness labels, and prunes stale entries. Completion triggers a MemorySourceSyncCompleted event in src/core/events.rs, allowing UI components to display real-time progress.
Convergence in the Memory Tree
Both ingestion pathways ultimately execute insert_chunk or import_chunks, writing Chunk records to the SQLite database (memory::store::chunks). The tree index in memory::tree recomputes regardless of origin, ensuring all content appears under a unified hierarchical view.
Key integration points include:
- Source-ID handling: Auto-fetch uses
source_id = Nonewhile source-sync sets concrete identifiers validated inmemory/sync_events_bridge.rs. - Chunk metadata: Both paths populate
ChunkMetadatawith timestamps, size, and provenance data, enabling UI labels like "last synced 5 minutes ago". - Event bus: The
MemorySourceSyncCompletedevent published insrc/core/events.rsnotifies listeners across the stack, including front-end components.
Implementation Examples
Triggering Auto-Fetch from an Agent
let fetch = openhuman_core::openhuman::tools::WebFetchTool::new(
security,
vec!["*".into()], // allow list
Some(0), // no concurrency limit
Some(50_000), // max result size
);
let result = fetch.invoke(json!({ "url": "https://example.com/article" })).await?;
The fetched text becomes a new memory chunk without a source ID, immediately available for agent retrieval.
Manually Starting a Source Sync
use openhuman_core::openhuman::memory::ops::sync_source;
let source_id = "gmail-account-1".to_string();
sync_source(&core, &source_id).await?;
This routine contacts the Gmail driver, imports new email chunks, and refreshes the memory tree via reconcile_tree.
Listening for Sync Completion in the UI
import { socketService } from "@/services/socketService";
socketService.on("MemorySourceSyncCompleted", (event) => {
console.log(`Sync finished for ${event.sourceId}: ${event.newChunks} new chunks`);
});
The front-end receives events published by src/core/events.rs, enabling progress indicators and notification badges.
Summary
- Auto-fetch provides ad-hoc data retrieval through
WebFetchTool, bypassing persistent source configuration but still populating the memory tree viainsert_chunk. - Memory-source sync runs scheduled background jobs through
src/openhuman/cron/scheduler_gate.rsandsrc/openhuman/modules/ops.rs, using driver-specific RPCs to bulk-import data with attached source IDs. - Both pathways converge in
src/openhuman/memory/ops/sync.rsandsrc/openhuman/memory/tree/ops.rs, ensuring unified indexing and queryability regardless of ingestion method. - The system emits
MemorySourceSyncCompletedevents for real-time UI updates, while SQLite persistence guarantees durability across both auto-fetched and synchronized content.
Frequently Asked Questions
What is the difference between auto-fetch and memory source sync in OpenHuman?
Auto-fetch operates on-demand through agent tools like WebFetchTool, pulling arbitrary URLs without persistent source configuration and assigning source_id = None. Memory-source sync runs periodic background jobs for configured integrations (Gmail, Slack, etc.), maintains persistent source_id associations, and supports incremental updates through driver-specific RPC calls.
How does OpenHuman handle metadata for memory chunks?
All chunks populate a ChunkMetadata structure containing timestamps, content size, and provenance information. According to src/openhuman/memory/ops/sync.rs, both auto-fetch and source-sync populate these fields, enabling the UI to display freshness indicators and source attribution regardless of ingestion pathway.
Can auto-fetched content be distinguished from synchronized source data?
Yes. Auto-fetched chunks explicitly set source_id = None during creation in src/openhuman/memory/ops/sync.rs, while source-sync operations attach concrete source identifiers validated through memory/sync_events_bridge.rs. This distinction allows filtering queries by source provenance while maintaining unified tree indexing.
What happens when a memory source sync fails?
Errors during driver resolution or RPC execution—such as "the bound memory driver … does not serve source sync"—are raised in src/openhuman/memory/ops/sync.rs and propagated through the composio bus in src/openhuman/memory/sync/composio/bus.rs. Failed syncs do not update the memory tree, and error states are emitted through the event bus for UI handling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →