How to Perform Full‑Text Search on a Hyperresearch Vault

Hyperresearch enables fast, ranked full‑text search across your entire note collection using an embedded SQLite FTS5 virtual table exposed through both a CLI and a Python API.

Every note in a Hyperresearch vault is indexed in an SQLite database that includes an FTS5 virtual table named notes_fts. According to the jordan-gibbs/hyperresearch source code, this architecture allows you to query your knowledge base using Boolean logic, prefix wildcards, and phrase matching while applying sophisticated filters for tags, status, and graph relationships.

Understanding the Hyperresearch Search Architecture

The search system consists of three coordinated components that transform raw queries into ranked results.

The FTS5 Virtual Table (notes_fts)

At the storage layer, Hyperresearch maintains the notes_fts virtual table inside your vault’s SQLite database. This table tokenizes note content to support fast substring and phrase matching. When you execute a search, the system joins this virtual table against the main notes table to retrieve metadata such as titles, paths, tags, and status flags.

Core Search Components

The implementation spans three critical files:

When you invoke a search, the query string first passes through preprocess_query in fts.py (lines 31–63), which automatically adds prefix wildcards, splits alphanumeric tokens, and preserves quoted phrases.

Searching via the Command Line Interface

The CLI offers the fastest way to query your vault without writing code. The command builds a Vault object, constructs SearchFilters from your flags, and executes the search against notes_fts.

Basic Syntax and Examples


# Basic keyword search with automatic prefix matching

hyperresearch search "machine learning"

# Restrict results to specific tags and note status

hyperresearch search "deep learning" --tag AI --status evergreen

# Export results as JSON with full note bodies included

hyperresearch search "climate change" --limit 25 --json --include-body

The CLI automatically handles query preprocessing, so entering gpt4o will match gpt4o-mini and gpt4o-generated due to the prefix wildcard logic implemented in the underlying Python layer.

Searching Programmatically with Python

For automation scripts or integrations, import the search functions directly from the Hyperresearch package. The Python API gives you granular control over BM25 ranking weights and graph‑based constraints.

from hyperresearch.core.vault import Vault
from hyperresearch.search.fts import search_fts, SearchQueryError
from hyperresearch.search.filters import SearchFilters

# Auto-discover the .hyperresearch directory and sync the database

vault = Vault.discover()
vault.auto_sync()

# Configure filters using AND logic across multiple criteria

filters = SearchFilters(
    tags=["AI", "research"],
    status="evergreen",
    after="2023-01-01",
)

# Execute the search with custom ranking parameters

try:
    results = search_fts(
        vault.db,
        query="gpt4o",
        filters=filters,
        limit=20,
        ranking={
            "title_weight": 12.0,
            "body_weight": 1.0,
            "tags_weight": 5.0,
            "aliases_weight": 3.0,
        },
        quality_ranked=True,  # Adjust scores by source quality

    )
except SearchQueryError as exc:
    raise SystemExit(f"Invalid query: {exc}")

# Process results containing id, title, snippet, and score

for r in results:
    print(f"{r['id']}: {r['title']} (score {r['score']:.2f})")
    print(f"Snippet: {r['snippet']}\n")

The search_fts function constructs a SQL statement that joins notes_fts with the notes table, applies filters via SearchFilters.to_sql (lines 27–105 in filters.py), and returns a list of dictionaries containing id, title, snippet, score, and optionally the full note body.

Advanced Query Techniques

Beyond simple keyword matching, the API supports sophisticated filtering and ranking strategies.

Custom BM25 Ranking Weights

The ranking parameter accepts a dictionary to adjust BM25 scoring weights. Increase title_weight to prioritize matches in note titles, or boost tags_weight to surface results with specific tag relevance.

Graph‑Based Filters

Filter results based on link relationships within your knowledge graph:


# Find notes mentioning "data provenance" that link to a specific note ID

filters = SearchFilters(linked_to="note-1234")
results = search_fts(vault.db, query="data provenance", filters=filters)

This executes a WHERE clause that constrains the search to notes having an outbound link to note-1234, effectively combining full‑text search with graph traversal logic.

Date and Metadata Filtering

Use SearchFilters to constrain results by creation date, modification date, note type, or custom YAML frontmatter fields. The after and before parameters accept ISO‑8601 date strings that generate SQL date comparisons against the vault’s metadata schema.

Summary

  • Architecture: Hyperresearch uses SQLite FTS5 via the notes_fts virtual table joined to the main notes table for metadata retrieval.
  • Entry Points: Access search via hyperresearch search CLI commands or the search_fts function in src/hyperresearch/search/fts.py.
  • Query Processing: Raw queries are automatically preprocessed to add prefix wildcards and handle quoted phrases before execution.
  • Filtering: Apply constraints using SearchFilters in src/hyperresearch/search/filters.py for tags, status, date ranges, and graph links.
  • Ranking: Customize BM25 weights or enable quality‑aware ranking to prioritize high‑credibility sources.

Frequently Asked Questions

The search relies on SQLite’s built‑in FTS5 tokenizer, which splits text into terms and applies case‑insensitive matching. The preprocess_query function in src/hyperresearch/search/fts.py further processes input by splitting alphanumeric boundaries and appending prefix wildcards (*) to tokens, enabling partial word matching without requiring explicit stemming rules.

Can I search for exact phrases with Boolean operators?

Yes. The query preprocessor preserves double‑quoted strings as exact phrases. You can combine these with standard FTS5 Boolean operators: use AND (implicit between terms), OR, and NOT (or - prefix) inside your query string. For example, "machine learning" OR "deep learning" -tutorial returns notes containing either phrase but excluding the word "tutorial".

What is the performance impact of adding multiple filters?

Filters translate directly to SQL WHERE clauses that execute within the SQLite query planner. Since tags, status, and path columns are typically indexed, combining filters with full‑text queries remains efficient even on vaults containing tens of thousands of notes. The SearchFilters.to_sql method optimizes clause ordering to leverage SQLite’s query optimization.

The current implementation focuses on lexical matching via FTS5 BM25 ranking. While quality_ranked=True adjusts scores based on note metadata quality, true semantic or vector search is not implemented in search_fts. For fuzzy matching, the automatic prefix wildcards provide approximate matching for truncated words.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →