# How to Perform Full‑Text Search on a Hyperresearch Vault

> Learn to perform full-text search on your Hyperresearch vault. Discover how this powerful tool uses SQLite FTS5 for fast, ranked search via CLI and Python API.

- Repository: [Jordan Gibbs/hyperresearch](https://github.com/jordan-gibbs/hyperresearch)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Hyperresearch enables fast, ranked full‑text search across your entire note collection using an embedded SQLite FTS5 virtual table exposed through both a CLI and a Python API.**

Every note in a Hyperresearch vault is indexed in an SQLite database that includes an FTS5 virtual table named `notes_fts`. According to the jordan-gibbs/hyperresearch source code, this architecture allows you to query your knowledge base using Boolean logic, prefix wildcards, and phrase matching while applying sophisticated filters for tags, status, and graph relationships.

## Understanding the Hyperresearch Search Architecture

The search system consists of three coordinated components that transform raw queries into ranked results.

### The FTS5 Virtual Table (`notes_fts`)

At the storage layer, Hyperresearch maintains the `notes_fts` virtual table inside your vault’s SQLite database. This table tokenizes note content to support fast substring and phrase matching. When you execute a search, the system joins this virtual table against the main `notes` table to retrieve metadata such as titles, paths, tags, and status flags.

### Core Search Components

The implementation spans three critical files:

- **[`src/hyperresearch/search/fts.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/search/fts.py)** – Contains the `search_fts` function and `preprocess_query` logic that converts free‑text input into FTS5‑compatible syntax.
- **[`src/hyperresearch/search/filters.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/search/filters.py)** – Implements the `SearchFilters` class, which generates SQL `WHERE` clauses for tag, status, and date constraints.
- **[`src/hyperresearch/cli/search.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/search.py)** – Provides the CLI wrapper that parses arguments, instantiates `SearchFilters`, and calls `search_fts`.

When you invoke a search, the query string first passes through `preprocess_query` in [`fts.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/fts.py) (lines 31–63), which automatically adds prefix wildcards, splits alphanumeric tokens, and preserves quoted phrases.

## Searching via the Command Line Interface

The CLI offers the fastest way to query your vault without writing code. The command builds a `Vault` object, constructs `SearchFilters` from your flags, and executes the search against `notes_fts`.

### Basic Syntax and Examples

```bash

# Basic keyword search with automatic prefix matching

hyperresearch search "machine learning"

# Restrict results to specific tags and note status

hyperresearch search "deep learning" --tag AI --status evergreen

# Export results as JSON with full note bodies included

hyperresearch search "climate change" --limit 25 --json --include-body

```

The CLI automatically handles query preprocessing, so entering `gpt4o` will match `gpt4o-mini` and `gpt4o-generated` due to the prefix wildcard logic implemented in the underlying Python layer.

## Searching Programmatically with Python

For automation scripts or integrations, import the search functions directly from the Hyperresearch package. The Python API gives you granular control over BM25 ranking weights and graph‑based constraints.

### Initializing the Vault and Running a Search

```python
from hyperresearch.core.vault import Vault
from hyperresearch.search.fts import search_fts, SearchQueryError
from hyperresearch.search.filters import SearchFilters

# Auto-discover the .hyperresearch directory and sync the database

vault = Vault.discover()
vault.auto_sync()

# Configure filters using AND logic across multiple criteria

filters = SearchFilters(
    tags=["AI", "research"],
    status="evergreen",
    after="2023-01-01",
)

# Execute the search with custom ranking parameters

try:
    results = search_fts(
        vault.db,
        query="gpt4o",
        filters=filters,
        limit=20,
        ranking={
            "title_weight": 12.0,
            "body_weight": 1.0,
            "tags_weight": 5.0,
            "aliases_weight": 3.0,
        },
        quality_ranked=True,  # Adjust scores by source quality

    )
except SearchQueryError as exc:
    raise SystemExit(f"Invalid query: {exc}")

# Process results containing id, title, snippet, and score

for r in results:
    print(f"{r['id']}: {r['title']} (score {r['score']:.2f})")
    print(f"Snippet: {r['snippet']}\n")

```

The `search_fts` function constructs a SQL statement that joins `notes_fts` with the `notes` table, applies filters via `SearchFilters.to_sql` (lines 27–105 in [`filters.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/filters.py)), and returns a list of dictionaries containing `id`, `title`, `snippet`, `score`, and optionally the full note body.

## Advanced Query Techniques

Beyond simple keyword matching, the API supports sophisticated filtering and ranking strategies.

### Custom BM25 Ranking Weights

The `ranking` parameter accepts a dictionary to adjust BM25 scoring weights. Increase `title_weight` to prioritize matches in note titles, or boost `tags_weight` to surface results with specific tag relevance.

### Graph‑Based Filters

Filter results based on link relationships within your knowledge graph:

```python

# Find notes mentioning "data provenance" that link to a specific note ID

filters = SearchFilters(linked_to="note-1234")
results = search_fts(vault.db, query="data provenance", filters=filters)

```

This executes a `WHERE` clause that constrains the search to notes having an outbound link to `note-1234`, effectively combining full‑text search with graph traversal logic.

### Date and Metadata Filtering

Use `SearchFilters` to constrain results by creation date, modification date, note type, or custom YAML frontmatter fields. The `after` and `before` parameters accept ISO‑8601 date strings that generate SQL date comparisons against the vault’s metadata schema.

## Summary

- **Architecture**: Hyperresearch uses SQLite FTS5 via the `notes_fts` virtual table joined to the main `notes` table for metadata retrieval.
- **Entry Points**: Access search via `hyperresearch search` CLI commands or the `search_fts` function in [`src/hyperresearch/search/fts.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/search/fts.py).
- **Query Processing**: Raw queries are automatically preprocessed to add prefix wildcards and handle quoted phrases before execution.
- **Filtering**: Apply constraints using `SearchFilters` in [`src/hyperresearch/search/filters.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/search/filters.py) for tags, status, date ranges, and graph links.
- **Ranking**: Customize BM25 weights or enable quality‑aware ranking to prioritize high‑credibility sources.

## Frequently Asked Questions

### How does Hyperresearch handle stemming and tokenization during full‑text search?

The search relies on SQLite’s built‑in FTS5 tokenizer, which splits text into terms and applies case‑insensitive matching. The `preprocess_query` function in [`src/hyperresearch/search/fts.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/search/fts.py) further processes input by splitting alphanumeric boundaries and appending prefix wildcards (`*`) to tokens, enabling partial word matching without requiring explicit stemming rules.

### Can I search for exact phrases with Boolean operators?

Yes. The query preprocessor preserves double‑quoted strings as exact phrases. You can combine these with standard FTS5 Boolean operators: use `AND` (implicit between terms), `OR`, and `NOT` (or `-` prefix) inside your query string. For example, `"machine learning" OR "deep learning" -tutorial` returns notes containing either phrase but excluding the word "tutorial".

### What is the performance impact of adding multiple filters?

Filters translate directly to SQL `WHERE` clauses that execute within the SQLite query planner. Since `tags`, `status`, and `path` columns are typically indexed, combining filters with full‑text queries remains efficient even on vaults containing tens of thousands of notes. The `SearchFilters.to_sql` method optimizes clause ordering to leverage SQLite’s query optimization.

### Does the API support fuzzy or semantic search?

The current implementation focuses on lexical matching via FTS5 BM25 ranking. While `quality_ranked=True` adjusts scores based on note metadata quality, true semantic or vector search is not implemented in `search_fts`. For fuzzy matching, the automatic prefix wildcards provide approximate matching for truncated words.