Hister Data Storage Mechanisms: SQLite, PostgreSQL, and Vector Embeddings Explained

Hister stores structured application data in SQLite by default (with optional PostgreSQL support), maintains dense embedding vectors in a separate SQLite database using the sqlite-vec extension, and implements full-text search via SQLite FTS5 virtual tables.

Hister is an open-source web archiving and semantic search platform written in Go. Understanding the Hister data storage mechanisms is critical for administrators planning backup strategies and developers extending the platform. The architecture deliberately separates relational metadata from vector embeddings while leveraging SQLite's advanced extensions for both semantic similarity and lexical search capabilities.

Primary Relational Database: SQLite or PostgreSQL

Hister persists core application data—users, history entries, crawl jobs, document versions, and embedding jobs—through a relational database layer abstracted by GORM (the Go Object-Relational Mapper).

GORM Integration and Configuration

The database connection is initialized in server/model/model.go at lines 35‑44. During startup, the application reads config.Config.DatabaseConnection to determine the driver type (sqlite or postgres) and corresponding Data Source Name (DSN). The default configuration uses SQLite, storing data in db.sqlite3 (with the path configurable via config.Config.Database as defined in config/config.go at line 516).

When the user selects PostgreSQL, the same GORM models and migration logic apply, simply switching the underlying driver to gorm.io/driver/postgres. This design ensures schema consistency regardless of the backend.

Schema Migrations

Database schema creation and versioning are handled through the automigrate() function within server/model/model.go. This utility automatically creates tables, indexes, and virtual tables based on the current model definitions. Custom migration helpers supplement the auto-migration process for complex schema transformations across releases.

Hister implements a dedicated vector store to power its "search-by-embedding" semantic search functionality. This store maintains dense vector representations of document content for similarity queries.

The sqlite-vec Extension

The vector store relies on sqlite-vec, a specialized C extension for SQLite that provides optimized vector operations. The extension auto-loads via sqlitevec.Auto() in server/vectorstore/sqlitevec/vec.go. Because this extension is SQLite-specific, the vector store always uses SQLite regardless of whether the primary database runs on PostgreSQL.

Vector Storage Implementation

The vector database resides in a separate file named vectors.sqlite3, created adjacent to the main relational database. The Go wrapper implementing the storage interface lives in server/vectorstore/sqlite.go (lines 28‑33), exposing methods for inserting chunks and performing similarity searches.

import (
    "github.com/asciimoo/hister/server/vectorstore"
)

func storeEmbedding(docID string, vec []float32) error {
    // vectors.sqlite3 lives next to the main DB file
    store, err := vectorstore.NewSQLiteVectorStore("vectors.sqlite3")
    if err != nil {
        return err
    }
    defer store.Close()

    // A slice of chunks (each chunk gets its own vector)
    chunks := []vectorstore.Chunk{
        {DocID: docID, Content: "…", Vector: vec},
    }
    return store.PutChunks(docID, 0, chunks)
}

To query the vector store for similar documents, use the Search method, which accepts a query vector, the number of results to return (topK), a minimum similarity threshold, and a user ID for access control (passing 0 disables per-user filtering to search across all vectors):

func similarDocs(queryVec []float32, topK int) ([]vectorstore.Result, error) {
    store, err := vectorstore.NewSQLiteVectorStore("vectors.sqlite3")
    if err != nil {
        return nil, err
    }
    defer store.Close()

    // userID = 0 disables per-user filtering (all vectors are visible)
    return store.Search(queryVec, topK, 0.0, 0)
}

Full-Text Search with SQLite FTS5

Hister builds a full-text index on document textual content using SQLite's FTS5 (Full-Text Search version 5) virtual tables. The relevant schema is generated automatically during the automigrate() process (see server/model/model.go lines 75‑82).

The DocumentVersion model includes a Content column that GORM indexes via FTS5, enabling fast keyword searches without external search engines. GORM automatically routes MATCH queries against the underlying virtual table:

import "github.com/asciimoo/hister/server/model"

func searchText(term string) ([]model.DocumentVersion, error) {
    var docs []model.DocumentVersion
    // GORM will use the underlying FTS5 virtual table automatically
    if err := model.DB.
        Where("content MATCH ?", term).
        Find(&docs).Error; err != nil {
        return nil, err
    }
    return docs, nil
}

Working with Hister's Data Layer

Initializing the relational database requires loading the configuration and invoking the model package's initialization routine:

import (
    "github.com/asciimoo/hister/config"
    "github.com/asciimoo/hister/server/model"
)

func main() {
    cfg, _ := config.Load()           // reads config.yaml / defaults
    // Initialise GORM + run migrations
    if err := model.Init(cfg); err != nil {
        log.Fatalf("DB init failed: %v", err)
    }
    // DB is now available as model.DB
}

Summary

  • Primary storage uses SQLite by default (db.sqlite3) with optional PostgreSQL support, managed through GORM in server/model/model.go.
  • Vector embeddings reside in a separate vectors.sqlite3 file utilizing the sqlite-vec C extension for high-performance similarity search.
  • Full-text indexing leverages SQLite FTS5 virtual tables created automatically during schema migrations, enabling native keyword search on document content.
  • Configuration is centralized in config/config.go, controlling database paths, connection strings, and driver selection.

Frequently Asked Questions

Does Hister require PostgreSQL to run?

No. Hister runs entirely on SQLite by default, including both the relational database (db.sqlite3) and the vector store (vectors.sqlite3). PostgreSQL is optional and can be enabled by changing the DatabaseConnection configuration value to use the postgres driver with GORM.

Why does Hister use a separate database file for vectors?

The vector store requires the sqlite-vec C extension, which is only available for SQLite. Even when the primary database runs on PostgreSQL, Hister maintains the vector store in a separate SQLite file (vectors.sqlite3) to leverage this specialized extension for efficient vector similarity operations.

How does Hister handle full-text searching?

Hister implements full-text search using SQLite FTS5 virtual tables. During the automigrate() process in server/model/model.go, GORM creates a FTS5 virtual table indexing the Content column of the DocumentVersion model. This allows the application to execute high-performance keyword queries using standard SQL MATCH clauses without external search infrastructure.

Where are the database files located?

By default, both db.sqlite3 (relational data) and vectors.sqlite3 (embeddings) are created in the working directory or at the path specified by config.Config.Database (defined in config/config.go line 516). These are standard SQLite files that can be backed up, moved, or inspected using any SQLite client or command-line tool.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →