# How the Hyperresearch Research Vault Is Structured: Markdown Layout and SQLite Schema

> Explore the Hyperresearch vault structure, detailing its markdown layout and SQLite schema. Understand how research data is organized for efficient retrieval and management.

- Repository: [Jordan Gibbs/hyperresearch](https://github.com/jordan-gibbs/hyperresearch)
- Tags: internals
- Published: 2026-09-13

---

**Hyperresearch stores a research vault as a dual-layer architecture with markdown files in a visible `research/` directory and SQLite metadata in `.hyperresearch/hyperresearch.db`, coordinated through [`src/hyperresearch/core/vault.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/vault.py) and [`src/hyperresearch/core/db.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/db.py).**

The **research vault** in the jordan-gibbs/hyperresearch open-source project separates human-readable content from machine-queryable indexes. This design allows researchers to edit markdown files directly while the system maintains high-performance queries via a relational database.

## Directory Layout and Markdown Structure

Hyperresearch organizes vault contents into two distinct zones: a hidden control directory for system files and a visible research directory for user content. This layout is implemented in [`src/hyperresearch/core/vault.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/vault.py) through property methods that resolve paths relative to the vault root.

### Hidden Control Directory (.hyperresearch/)

The `.hyperresearch/` directory contains system-level assets that remain outside the user workspace:

- **[`.hyperresearch/config.toml`](https://github.com/jordan-gibbs/hyperresearch/blob/main/.hyperresearch/config.toml)** – Human-readable TOML configuration defining the vault name, visible research directory, and operational settings.
- **`.hyperresearch/hyperresearch.db`** – The single SQLite database file indexing all metadata.
- **`.hyperresearch/templates/`** – Default markdown templates, including [`note.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/note.md) used when creating new notes.
- **`.hyperresearch/exports/`** – Destination for CLI-generated artifacts such as CSV or JSON exports.

### Visible Research Directory (research/)

The `research/` directory (configurable via [`config.toml`](https://github.com/jordan-gibbs/hyperresearch/blob/main/config.toml)) serves as the primary workspace. By default, this directory contains several specialized subdirectories:

- **`research/notes/`** – Stores individual markdown files for each note, synchronized to the SQLite `notes` and `note_content` tables.
- **`research/index/`** – Holds generated markdown indexes such as statistics pages and tables of contents.
- **`research/temp/`** – Staging area for auto-generated stubs and agent artifacts. Files here are synced to the database but excluded from ordinary directory listings.
- **`research/runs/<run-tag>/`** – Per-run workspaces for pipeline artifacts, decompositions, and logs. Run tags are validated by `validate_run_tag` in the CLI to ensure safe filesystem slugs.

## SQLite Database Schema

All vault metadata lives in a single SQLite file located at `.hyperresearch/hyperresearch.db`. The schema is defined by the `SCHEMA_SQL` multiline string in [`src/hyperresearch/core/db.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/db.py) and includes tables for content, relationships, vectors, and auxiliary caches.

### Core Entity Tables

The foundation of the schema stores note identities and content separately:

- **`notes`** – Contains core metadata including `id`, `title`, `status`, `type`, timestamps, and content hashes.
- **`note_content`** – Stores the raw markdown body alongside a plain-text version for processing.

### Relationship and Metadata Tables

Hyperresearch tracks connections and classifications through normalized tables:

- **`tags`** and **`aliases`** – Many-to-one mappings linking tag strings and alternative titles to note records.
- **`links`** – Source-target relationships with line numbers and contextual snippets, indexed for fast bi-directional lookups.

### Vector and Claim Storage

Advanced features require specialized storage:

- **`embeddings`** – Vector embeddings generated by LLMs for semantic search.
- **`claims`** – Structured claim objects extracted from note content during analysis.

### Caching and Auxiliary Tables

The schema includes utility tables for operational resilience:

- **`api_cache`** – Caches HTTP responses to avoid redundant external calls.
- **`escalations`** – Tracks escalated fetch operations requiring manual review.
- **`sources`** – Indexes external source references.
- **`assets`** – Manages attached file metadata.

### Full-Text Search Implementation

Beyond standard tables, the database supports full-text search via FTS5 virtual tables. The `FTS_SQL` string in [`src/hyperresearch/core/db.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/db.py) defines `notes_fts` and `claims_fts` virtual tables, enabling fast natural-language queries across note bodies and extracted claims.

## Vault Initialization Workflow

When initializing a new vault via `Vault.init()`, the system executes a deterministic setup sequence. The implementation in [`src/hyperresearch/core/vault.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/vault.py) performs the following actions:

1. Creates the hidden `.hyperresearch/` directory with `templates/` and `exports/` subdirectories.
2. Creates the visible `research/` directory tree, including `notes/`, `index/`, and `temp/`.
3. Writes a default [`config.toml`](https://github.com/jordan-gibbs/hyperresearch/blob/main/config.toml) and the [`note.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/note.md) template file.
4. Opens a SQLite connection and executes `init_schema()` from [`src/hyperresearch/core/db.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/db.py) to materialize all tables and indexes.

## Synchronizing Markdown and SQLite

The vault maintains consistency between filesystem state and database indexes through the sync mechanism in [`src/hyperresearch/core/sync.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/sync.py). When markdown files are added, modified, or removed from `research/notes/` or `research/temp/`, the `auto_sync()` method incrementally updates the corresponding database records without requiring a full rebuild. This ensures that links, tags, and full-text search indexes remain current while preserving the markdown files as the source of truth.

## Practical Code Examples

### Creating a New Vault

```python
from pathlib import Path
from hyperresearch.core.vault import Vault

# Initialize a vault at ./my-vault

vault_path = Path("./my-vault")
vault = Vault.init(vault_path, name="My Research Vault")
print(vault.research_dir)            # => my-vault/research

print(vault.db_path)                 # => my-vault/.hyperresearch/hyperresearch.db

```

### Adding a Markdown Note

```python
import uuid, datetime
from hyperresearch.core.vault import Vault

vault = Vault.discover()                     # Finds the nearest vault from cwd

note_id = str(uuid.uuid4())
note_path = vault.notes_dir / f"{note_id}.md"

# Write a simple markdown note

note_path.write_text(
    f"""---
title: "Sample Note"
id: "{note_id}"
tags: []
status: draft
type: note
created: {datetime.datetime.utcnow().isoformat()}
---

# Sample Note

This is a sample note inside the Hyperresearch vault.
"""
)

# Sync to database

vault.auto_sync()

```

### Querying the SQLite Database

```python
import sqlite3
from hyperresearch.core.vault import Vault

vault = Vault.discover()
conn: sqlite3.Connection = vault.db

# List all draft notes

cur = conn.execute(
    "SELECT id, title, created FROM notes WHERE status = 'draft' ORDER BY created DESC"
)
for row in cur:
    print(row["id"], row["title"], row["created"])

```

### Inserting Link Records

```python

# Register a link from note A to note B

conn.execute(
    """
    INSERT INTO links (source_id, target_ref, target_id, line_number, context)
    VALUES (?, ?, ?, ?, ?)
    """,
    ("noteA-id", "B", "noteB-id", 12, "See discussion in B")
)
conn.commit()

```

## Summary

- **Hybrid Storage**: Hyperresearch uses markdown files in `research/` for content and SQLite in `.hyperresearch/` for metadata.
- **Schema Completeness**: The database schema in [`src/hyperresearch/core/db.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/db.py) includes tables for notes, content, tags, links, embeddings, claims, and full-text search.
- **Initialization**: The `Vault.init()` method creates both the directory structure and database schema atomically.
- **Synchronization**: The [`sync.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/sync.py) module incrementally updates database indexes to match filesystem changes.
- **Extensibility**: Developers can query the SQLite database directly via `vault.db` while maintaining the vault's integrity.

## Frequently Asked Questions

### Where is the Hyperresearch SQLite database file located?

The database is stored at `.hyperresearch/hyperresearch.db` relative to the vault root. This path is provided by the `db_path` property in [`src/hyperresearch/core/vault.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/vault.py), which ensures all metadata remains colocated with the vault but hidden from normal directory views.

### What tables are created in the Hyperresearch vault database?

The schema includes core tables (`notes`, `note_content`), relationship tables (`links`, `tags`, `aliases`), vector storage (`embeddings`), structured data (`claims`), and auxiliary tables (`api_cache`, `escalations`, `sources`, `assets`). Additionally, FTS5 virtual tables `notes_fts` and `claims_fts` enable full-text search capabilities.

### How does Hyperresearch keep markdown files synchronized with the database?

The system uses incremental synchronization logic in [`src/hyperresearch/core/sync.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/sync.py). When `vault.auto_sync()` is invoked, the system scans markdown files in `research/notes/` and `research/temp/`, computing content hashes to detect changes and updating the SQLite records accordingly without rebuilding the entire index.

### Can I manually move or rename markdown files within the vault?

While markdown files serve as the source of truth, manual moves require caution. After relocating files, you must run `vault.auto_sync()` to update the database paths and preserve link integrity. The `links` table stores `source_id` and `target_id` references that remain valid as long as the note IDs within the file frontmatter remain unchanged.