How the Hyperresearch Research Vault Is Structured: Markdown Layout and SQLite Schema

Hyperresearch stores a research vault as a dual-layer architecture with markdown files in a visible research/ directory and SQLite metadata in .hyperresearch/hyperresearch.db, coordinated through src/hyperresearch/core/vault.py and src/hyperresearch/core/db.py.

The research vault in the jordan-gibbs/hyperresearch open-source project separates human-readable content from machine-queryable indexes. This design allows researchers to edit markdown files directly while the system maintains high-performance queries via a relational database.

Directory Layout and Markdown Structure

Hyperresearch organizes vault contents into two distinct zones: a hidden control directory for system files and a visible research directory for user content. This layout is implemented in src/hyperresearch/core/vault.py through property methods that resolve paths relative to the vault root.

Hidden Control Directory (.hyperresearch/)

The .hyperresearch/ directory contains system-level assets that remain outside the user workspace:

  • .hyperresearch/config.toml – Human-readable TOML configuration defining the vault name, visible research directory, and operational settings.
  • .hyperresearch/hyperresearch.db – The single SQLite database file indexing all metadata.
  • .hyperresearch/templates/ – Default markdown templates, including note.md used when creating new notes.
  • .hyperresearch/exports/ – Destination for CLI-generated artifacts such as CSV or JSON exports.

Visible Research Directory (research/)

The research/ directory (configurable via config.toml) serves as the primary workspace. By default, this directory contains several specialized subdirectories:

  • research/notes/ – Stores individual markdown files for each note, synchronized to the SQLite notes and note_content tables.
  • research/index/ – Holds generated markdown indexes such as statistics pages and tables of contents.
  • research/temp/ – Staging area for auto-generated stubs and agent artifacts. Files here are synced to the database but excluded from ordinary directory listings.
  • research/runs/<run-tag>/ – Per-run workspaces for pipeline artifacts, decompositions, and logs. Run tags are validated by validate_run_tag in the CLI to ensure safe filesystem slugs.

SQLite Database Schema

All vault metadata lives in a single SQLite file located at .hyperresearch/hyperresearch.db. The schema is defined by the SCHEMA_SQL multiline string in src/hyperresearch/core/db.py and includes tables for content, relationships, vectors, and auxiliary caches.

Core Entity Tables

The foundation of the schema stores note identities and content separately:

  • notes – Contains core metadata including id, title, status, type, timestamps, and content hashes.
  • note_content – Stores the raw markdown body alongside a plain-text version for processing.

Relationship and Metadata Tables

Hyperresearch tracks connections and classifications through normalized tables:

  • tags and aliases – Many-to-one mappings linking tag strings and alternative titles to note records.
  • links – Source-target relationships with line numbers and contextual snippets, indexed for fast bi-directional lookups.

Vector and Claim Storage

Advanced features require specialized storage:

  • embeddings – Vector embeddings generated by LLMs for semantic search.
  • claims – Structured claim objects extracted from note content during analysis.

Caching and Auxiliary Tables

The schema includes utility tables for operational resilience:

  • api_cache – Caches HTTP responses to avoid redundant external calls.
  • escalations – Tracks escalated fetch operations requiring manual review.
  • sources – Indexes external source references.
  • assets – Manages attached file metadata.

Full-Text Search Implementation

Beyond standard tables, the database supports full-text search via FTS5 virtual tables. The FTS_SQL string in src/hyperresearch/core/db.py defines notes_fts and claims_fts virtual tables, enabling fast natural-language queries across note bodies and extracted claims.

Vault Initialization Workflow

When initializing a new vault via Vault.init(), the system executes a deterministic setup sequence. The implementation in src/hyperresearch/core/vault.py performs the following actions:

  1. Creates the hidden .hyperresearch/ directory with templates/ and exports/ subdirectories.
  2. Creates the visible research/ directory tree, including notes/, index/, and temp/.
  3. Writes a default config.toml and the note.md template file.
  4. Opens a SQLite connection and executes init_schema() from src/hyperresearch/core/db.py to materialize all tables and indexes.

Synchronizing Markdown and SQLite

The vault maintains consistency between filesystem state and database indexes through the sync mechanism in src/hyperresearch/core/sync.py. When markdown files are added, modified, or removed from research/notes/ or research/temp/, the auto_sync() method incrementally updates the corresponding database records without requiring a full rebuild. This ensures that links, tags, and full-text search indexes remain current while preserving the markdown files as the source of truth.

Practical Code Examples

Creating a New Vault

from pathlib import Path
from hyperresearch.core.vault import Vault

# Initialize a vault at ./my-vault

vault_path = Path("./my-vault")
vault = Vault.init(vault_path, name="My Research Vault")
print(vault.research_dir)            # => my-vault/research

print(vault.db_path)                 # => my-vault/.hyperresearch/hyperresearch.db

Adding a Markdown Note

import uuid, datetime
from hyperresearch.core.vault import Vault

vault = Vault.discover()                     # Finds the nearest vault from cwd

note_id = str(uuid.uuid4())
note_path = vault.notes_dir / f"{note_id}.md"

# Write a simple markdown note

note_path.write_text(
    f"""---
title: "Sample Note"
id: "{note_id}"
tags: []
status: draft
type: note
created: {datetime.datetime.utcnow().isoformat()}
---

# Sample Note

This is a sample note inside the Hyperresearch vault.
"""
)

# Sync to database

vault.auto_sync()

Querying the SQLite Database

import sqlite3
from hyperresearch.core.vault import Vault

vault = Vault.discover()
conn: sqlite3.Connection = vault.db

# List all draft notes

cur = conn.execute(
    "SELECT id, title, created FROM notes WHERE status = 'draft' ORDER BY created DESC"
)
for row in cur:
    print(row["id"], row["title"], row["created"])

# Register a link from note A to note B

conn.execute(
    """
    INSERT INTO links (source_id, target_ref, target_id, line_number, context)
    VALUES (?, ?, ?, ?, ?)
    """,
    ("noteA-id", "B", "noteB-id", 12, "See discussion in B")
)
conn.commit()

Summary

  • Hybrid Storage: Hyperresearch uses markdown files in research/ for content and SQLite in .hyperresearch/ for metadata.
  • Schema Completeness: The database schema in src/hyperresearch/core/db.py includes tables for notes, content, tags, links, embeddings, claims, and full-text search.
  • Initialization: The Vault.init() method creates both the directory structure and database schema atomically.
  • Synchronization: The sync.py module incrementally updates database indexes to match filesystem changes.
  • Extensibility: Developers can query the SQLite database directly via vault.db while maintaining the vault's integrity.

Frequently Asked Questions

Where is the Hyperresearch SQLite database file located?

The database is stored at .hyperresearch/hyperresearch.db relative to the vault root. This path is provided by the db_path property in src/hyperresearch/core/vault.py, which ensures all metadata remains colocated with the vault but hidden from normal directory views.

What tables are created in the Hyperresearch vault database?

The schema includes core tables (notes, note_content), relationship tables (links, tags, aliases), vector storage (embeddings), structured data (claims), and auxiliary tables (api_cache, escalations, sources, assets). Additionally, FTS5 virtual tables notes_fts and claims_fts enable full-text search capabilities.

How does Hyperresearch keep markdown files synchronized with the database?

The system uses incremental synchronization logic in src/hyperresearch/core/sync.py. When vault.auto_sync() is invoked, the system scans markdown files in research/notes/ and research/temp/, computing content hashes to detect changes and updating the SQLite records accordingly without rebuilding the entire index.

Can I manually move or rename markdown files within the vault?

While markdown files serve as the source of truth, manual moves require caution. After relocating files, you must run vault.auto_sync() to update the database paths and preserve link integrity. The links table stores source_id and target_id references that remain valid as long as the note IDs within the file frontmatter remain unchanged.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →