How Rowboat Builds and Updates Its Knowledge Graph from Emails and Meeting Notes

Rowboat incrementally builds and updates its knowledge graph by detecting new or changed markdown files across source folders, rebuilding a live index of existing entities, and invoking an LLM agent to extract entities and update notes in the knowledge/ directory.

The rowboatlabs/rowboat repository implements a file-centric knowledge graph build and update pipeline that transforms raw data sources into an Obsidian-style vault. By treating every email thread, meeting transcript, and voice memo as a markdown file, Rowboat enables incremental updates without full reprocessing.

Ingesting Raw Data into the Knowledge Graph Pipeline

Rowboat consumes heterogeneous data by first normalizing it to markdown. Each source writes to a dedicated folder under the workspace, monitored by the graph builder.

Gmail Synchronization via sync_gmail.ts

The Gmail integration polls the API and persists threads as markdown. In apps/x/packages/core/src/knowledge/sync_gmail.ts, the processThread function fetches thread data and writes it to gmail_sync/<threadId>.md:

// apps/x/packages/core/src/knowledge/sync_gmail.ts
const SYNC_DIR = path.join(WorkDir, 'gmail_sync');

async function processThread(auth, threadId, syncDir, attachmentsDir) {
  const gmail = google.gmail({ version: 'v1', auth });
  const res = await gmail.users.threads.get({ userId: 'me', id: threadId });
  const thread = res.data;
  // …build markdown with subject, messages, attachments…
  fs.writeFileSync(path.join(syncDir, `${threadId}.md`), mdContent);
}

Each execution writes or overwrites the file, preserving a stable identifier (the thread ID) so subsequent runs can detect changes via filesystem metadata.

Meeting Transcripts and Voice Memos

Other sources follow the same pattern:

  • Fireflies transcripts land in fireflies_transcripts/ as individual markdown files
  • Granola notes are written directly to granola_notes/ via UI components
  • Voice memos are saved to knowledge/Voice Memos/<date>/voice-memo-*.md by the Electron UI

All folders are registered in the SOURCE_FOLDERS array within apps/x/packages/core/src/knowledge/build_graph.ts, enabling the pipeline to monitor them collectively.

Tracking Changes for Incremental Knowledge Graph Updates

To avoid reprocessing unchanged files, Rowboat maintains a lightweight state database.

The State File: knowledge_graph_state.json

The system records the mtime and SHA-256 hash of every processed file in knowledge_graph_state.json. The state management logic resides in apps/x/packages/core/src/knowledge/graph_state.ts:

  • loadState() and saveState() handle persistence
  • hasFileChanged() performs a fast mtime check followed by hash verification
  • markFileAsProcessed() commits the new metadata after successful ingestion

Detecting Modified Files with graph_state.ts

The getFilesToProcess function recursively scans source directories and filters for markdown files that have changed since the last run:

// apps/x/packages/core/src/knowledge/graph_state.ts
export function getFilesToProcess(sourceDir: string, state: GraphState): string[] {
  // Recursively walk the directory …
  if (stat.isFile() && entry.endsWith('.md') && hasFileChanged(fullPath, state)) {
    filesToProcess.push(fullPath);
  }
}

This function is invoked for each entry in SOURCE_FOLDERS, producing a deduplicated list of files requiring processing.

Indexing Existing Entities

Before processing new data, Rowboat constructs a comprehensive index of the current knowledge vault to provide the agent with context for entity resolution.

Building the Knowledge Index with knowledge_index.ts

The buildKnowledgeIndex function in apps/x/packages/core/src/knowledge/knowledge_index.ts scans the knowledge/ directory recursively and parses markdown frontmatter and content to extract:

  • People (names, emails, aliases)
  • Organizations
  • Projects
  • Topics
  • Other entities

The parser uses helper functions like extractField() and extractList() to pull structured data from the markdown.

The index is then formatted for LLM consumption via formatIndexForPrompt(), which renders the entities as a markdown table. This index is injected into the agent's system prompt, ensuring the agent can resolve entities (e.g., determining that "John Doe" in a new email is the same person as "J. Doe" in an existing note) without performing expensive filesystem searches during inference.

The Graph Builder Loop

The orchestration logic runs continuously, coordinating detection, indexing, and agent invocation.

Orchestrating the Knowledge Graph Build with build_graph.ts

The main loop in apps/x/packages/core/src/knowledge/build_graph.ts executes every 30 seconds:

// apps/x/packages/core/src/knowledge/build_graph.ts
async function runBuilder() {
  const state = loadState();
  const files = SOURCE_FOLDERS.flatMap(dir => getFilesToProcess(path.join(WorkDir, dir), state));
  if (!files.length) return;            // nothing new

  const index = buildKnowledgeIndex();  // current view of all notes
  const fileContents = await readFileContents(files);
  const run = await createRun({ agentId: NOTE_CREATION_AGENT });

  // Build the system prompt – index + raw files
  const message = `
    Process the following ${files.length} source files and create/update obsidian notes.
    ---\n${formatIndexForPrompt(index)}\n---\n${fileContents.map(...).join('\n')}`;
  await createMessage({ runId: run.id, role: 'user', content: message });

  await waitForRunCompletion(run.id);
  files.forEach(fp => markFileAsProcessed(fp, state));
  saveState(state);
}

This function implements the complete knowledge graph build and update cycle: detecting changes, indexing the current vault, prompting the agent with context, and committing state updates.

The Note Creation Agent

The agent responsible for entity extraction is defined in apps/x/packages/core/src/pre_built/note_creation_high.ts. It receives the knowledge index and raw source content, then uses workspace tools (workspace-readFile, workspace-writeFile) to:

  1. Resolve entities against the provided index (avoiding duplicate creation)
  2. Create or update markdown files in knowledge/People/, knowledge/Projects/, etc.
  3. Merge information when the same entity appears across multiple sources

The agent operates on the principle that the knowledge graph build and update process is idempotent: running the pipeline multiple times on the same file produces the same result because the agent uses the live index to deduplicate entities.

Extending the Pipeline to New Data Sources

Adding a new source requires only two steps:

  1. Create a sync script that writes markdown files to a new folder (e.g., slack_sync/) in the workspace
  2. Register the folder in the SOURCE_FOLDERS array in apps/x/packages/core/src/knowledge/build_graph.ts

The existing knowledge graph build and update infrastructure automatically begins tracking the new folder, detecting changes, and processing files through the same agent pipeline without additional configuration.

Summary

  • Rowboat maintains an Obsidian-style knowledge graph as plain markdown files in the knowledge/ directory
  • Raw data ingestion normalizes emails, transcripts, and memos to markdown via source-specific sync scripts like sync_gmail.ts
  • Incremental processing uses knowledge_graph_state.json to track file mtime and SHA-256 hashes, avoiding redundant work
  • Live indexing rebuilds the entity index on every run, providing the agent with current context for deduplication
  • Agent-driven updates use the note_creation agent to extract entities and write updates, ensuring the knowledge graph build and update process remains consistent across all data sources

Frequently Asked Questions

How does Rowboat detect new emails or meeting notes without reprocessing everything?

Rowboat stores the last modified time (mtime) and SHA-256 hash of every processed file in knowledge_graph_state.json. The getFilesToProcess function in apps/x/packages/core/src/knowledge/graph_state.ts compares the current filesystem state against this record, returning only files that have changed or are new.

What happens if the same person is mentioned in both an email and a meeting transcript?

The note_creation agent receives a knowledge index built by buildKnowledgeIndex() that contains all existing entities (people, organizations, projects). When processing new files, the agent uses this index to resolve entities by name, email, or alias, merging new information into the existing note rather than creating duplicates.

Can I add support for Slack or other proprietary data sources?

Yes. To extend the knowledge graph build and update pipeline to new sources, create a sync script that writes markdown files to a dedicated folder (e.g., slack_sync/), then add that folder name to the SOURCE_FOLDERS array in apps/x/packages/core/src/knowledge/build_graph.ts. The existing state tracking and agent processing will handle the new source automatically.

How does the system handle voice memos differently from emails?

Voice memos are written directly to knowledge/Voice Memos/<date>/ as markdown files by the Electron UI, while emails are staged in gmail_sync/ before processing. However, both follow the identical knowledge graph build and update cycle: the getFilesToProcess function detects them, the index is rebuilt, and the note_creation agent extracts entities and updates the vault.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →