# How Rowboat Builds and Updates Its Knowledge Graph from Emails and Meeting Notes

> Discover how Rowboat builds and updates its knowledge graph from emails and meeting notes. Learn about its efficient entity extraction and live index rebuilding process.

- Repository: [RowBoat Labs/rowboat](https://github.com/rowboatlabs/rowboat)
- Tags: how-to-guide
- Published: 2026-02-16

---

**Rowboat incrementally builds and updates its knowledge graph by detecting new or changed markdown files across source folders, rebuilding a live index of existing entities, and invoking an LLM agent to extract entities and update notes in the `knowledge/` directory.**

The `rowboatlabs/rowboat` repository implements a file-centric knowledge graph build and update pipeline that transforms raw data sources into an Obsidian-style vault. By treating every email thread, meeting transcript, and voice memo as a markdown file, Rowboat enables incremental updates without full reprocessing.

## Ingesting Raw Data into the Knowledge Graph Pipeline

Rowboat consumes heterogeneous data by first normalizing it to markdown. Each source writes to a dedicated folder under the workspace, monitored by the graph builder.

### Gmail Synchronization via sync_gmail.ts

The Gmail integration polls the API and persists threads as markdown. In [`apps/x/packages/core/src/knowledge/sync_gmail.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/sync_gmail.ts), the `processThread` function fetches thread data and writes it to `gmail_sync/<threadId>.md`:

```typescript
// apps/x/packages/core/src/knowledge/sync_gmail.ts
const SYNC_DIR = path.join(WorkDir, 'gmail_sync');

async function processThread(auth, threadId, syncDir, attachmentsDir) {
  const gmail = google.gmail({ version: 'v1', auth });
  const res = await gmail.users.threads.get({ userId: 'me', id: threadId });
  const thread = res.data;
  // …build markdown with subject, messages, attachments…
  fs.writeFileSync(path.join(syncDir, `${threadId}.md`), mdContent);
}

```

Each execution writes or overwrites the file, preserving a stable identifier (the thread ID) so subsequent runs can detect changes via filesystem metadata.

### Meeting Transcripts and Voice Memos

Other sources follow the same pattern:

- **Fireflies transcripts** land in `fireflies_transcripts/` as individual markdown files
- **Granola notes** are written directly to `granola_notes/` via UI components
- **Voice memos** are saved to `knowledge/Voice Memos/<date>/voice-memo-*.md` by the Electron UI

All folders are registered in the `SOURCE_FOLDERS` array within [`apps/x/packages/core/src/knowledge/build_graph.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/build_graph.ts), enabling the pipeline to monitor them collectively.

## Tracking Changes for Incremental Knowledge Graph Updates

To avoid reprocessing unchanged files, Rowboat maintains a lightweight state database.

### The State File: knowledge_graph_state.json

The system records the **mtime** and **SHA-256 hash** of every processed file in [`knowledge_graph_state.json`](https://github.com/rowboatlabs/rowboat/blob/main/knowledge_graph_state.json). The state management logic resides in [`apps/x/packages/core/src/knowledge/graph_state.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/graph_state.ts):

- **`loadState()`** and **`saveState()`** handle persistence
- **`hasFileChanged()`** performs a fast mtime check followed by hash verification
- **`markFileAsProcessed()`** commits the new metadata after successful ingestion

### Detecting Modified Files with graph_state.ts

The `getFilesToProcess` function recursively scans source directories and filters for markdown files that have changed since the last run:

```typescript
// apps/x/packages/core/src/knowledge/graph_state.ts
export function getFilesToProcess(sourceDir: string, state: GraphState): string[] {
  // Recursively walk the directory …
  if (stat.isFile() && entry.endsWith('.md') && hasFileChanged(fullPath, state)) {
    filesToProcess.push(fullPath);
  }
}

```

This function is invoked for each entry in `SOURCE_FOLDERS`, producing a deduplicated list of files requiring processing.

## Indexing Existing Entities

Before processing new data, Rowboat constructs a comprehensive index of the current knowledge vault to provide the agent with context for entity resolution.

### Building the Knowledge Index with knowledge_index.ts

The `buildKnowledgeIndex` function in [`apps/x/packages/core/src/knowledge/knowledge_index.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/knowledge_index.ts) scans the `knowledge/` directory recursively and parses markdown frontmatter and content to extract:

- **People** (names, emails, aliases)
- **Organizations**
- **Projects**
- **Topics**
- **Other** entities

The parser uses helper functions like `extractField()` and `extractList()` to pull structured data from the markdown.

The index is then formatted for LLM consumption via `formatIndexForPrompt()`, which renders the entities as a markdown table. This index is injected into the agent's system prompt, ensuring the agent can resolve entities (e.g., determining that "John Doe" in a new email is the same person as "J. Doe" in an existing note) without performing expensive filesystem searches during inference.

## The Graph Builder Loop

The orchestration logic runs continuously, coordinating detection, indexing, and agent invocation.

### Orchestrating the Knowledge Graph Build with build_graph.ts

The main loop in [`apps/x/packages/core/src/knowledge/build_graph.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/build_graph.ts) executes every 30 seconds:

```typescript
// apps/x/packages/core/src/knowledge/build_graph.ts
async function runBuilder() {
  const state = loadState();
  const files = SOURCE_FOLDERS.flatMap(dir => getFilesToProcess(path.join(WorkDir, dir), state));
  if (!files.length) return;            // nothing new

  const index = buildKnowledgeIndex();  // current view of all notes
  const fileContents = await readFileContents(files);
  const run = await createRun({ agentId: NOTE_CREATION_AGENT });

  // Build the system prompt – index + raw files
  const message = `
    Process the following ${files.length} source files and create/update obsidian notes.
    ---\n${formatIndexForPrompt(index)}\n---\n${fileContents.map(...).join('\n')}`;
  await createMessage({ runId: run.id, role: 'user', content: message });

  await waitForRunCompletion(run.id);
  files.forEach(fp => markFileAsProcessed(fp, state));
  saveState(state);
}

```

This function implements the complete **knowledge graph build and update** cycle: detecting changes, indexing the current vault, prompting the agent with context, and committing state updates.

### The Note Creation Agent

The agent responsible for entity extraction is defined in [`apps/x/packages/core/src/pre_built/note_creation_high.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/pre_built/note_creation_high.ts). It receives the knowledge index and raw source content, then uses workspace tools (`workspace-readFile`, `workspace-writeFile`) to:

1. Resolve entities against the provided index (avoiding duplicate creation)
2. Create or update markdown files in `knowledge/People/`, `knowledge/Projects/`, etc.
3. Merge information when the same entity appears across multiple sources

The agent operates on the principle that the **knowledge graph build and update** process is idempotent: running the pipeline multiple times on the same file produces the same result because the agent uses the live index to deduplicate entities.

## Extending the Pipeline to New Data Sources

Adding a new source requires only two steps:

1. **Create a sync script** that writes markdown files to a new folder (e.g., `slack_sync/`) in the workspace
2. **Register the folder** in the `SOURCE_FOLDERS` array in [`apps/x/packages/core/src/knowledge/build_graph.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/build_graph.ts)

The existing **knowledge graph build and update** infrastructure automatically begins tracking the new folder, detecting changes, and processing files through the same agent pipeline without additional configuration.

## Summary

- Rowboat maintains an **Obsidian-style knowledge graph** as plain markdown files in the `knowledge/` directory
- **Raw data ingestion** normalizes emails, transcripts, and memos to markdown via source-specific sync scripts like [`sync_gmail.ts`](https://github.com/rowboatlabs/rowboat/blob/main/sync_gmail.ts)
- **Incremental processing** uses [`knowledge_graph_state.json`](https://github.com/rowboatlabs/rowboat/blob/main/knowledge_graph_state.json) to track file mtime and SHA-256 hashes, avoiding redundant work
- **Live indexing** rebuilds the entity index on every run, providing the agent with current context for deduplication
- **Agent-driven updates** use the `note_creation` agent to extract entities and write updates, ensuring the knowledge graph build and update process remains consistent across all data sources

## Frequently Asked Questions

### How does Rowboat detect new emails or meeting notes without reprocessing everything?

Rowboat stores the last modified time (mtime) and SHA-256 hash of every processed file in [`knowledge_graph_state.json`](https://github.com/rowboatlabs/rowboat/blob/main/knowledge_graph_state.json). The `getFilesToProcess` function in [`apps/x/packages/core/src/knowledge/graph_state.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/graph_state.ts) compares the current filesystem state against this record, returning only files that have changed or are new.

### What happens if the same person is mentioned in both an email and a meeting transcript?

The `note_creation` agent receives a **knowledge index** built by `buildKnowledgeIndex()` that contains all existing entities (people, organizations, projects). When processing new files, the agent uses this index to resolve entities by name, email, or alias, merging new information into the existing note rather than creating duplicates.

### Can I add support for Slack or other proprietary data sources?

Yes. To extend the knowledge graph build and update pipeline to new sources, create a sync script that writes markdown files to a dedicated folder (e.g., `slack_sync/`), then add that folder name to the `SOURCE_FOLDERS` array in [`apps/x/packages/core/src/knowledge/build_graph.ts`](https://github.com/rowboatlabs/rowboat/blob/main/apps/x/packages/core/src/knowledge/build_graph.ts). The existing state tracking and agent processing will handle the new source automatically.

### How does the system handle voice memos differently from emails?

Voice memos are written directly to `knowledge/Voice Memos/<date>/` as markdown files by the Electron UI, while emails are staged in `gmail_sync/` before processing. However, both follow the identical knowledge graph build and update cycle: the `getFilesToProcess` function detects them, the index is rebuilt, and the `note_creation` agent extracts entities and updates the vault.