# Data Flow for Adding a Document to Hister: Complete Pipeline Guide

> Explore the eight-stage data flow pipeline for adding documents to Hister. Transform raw input into searchable entries with embeddings and update the SQLite database for instant retrieval.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: architecture
- Published: 2026-09-01

---

**When you add a document to Hister, the system executes an eight-stage pipeline that transforms raw input into a searchable database entry with semantic embeddings, persisting data in SQLite and updating the search schema for immediate retrieval.**

The open-source search engine **asciimoo/hister** processes document ingestion through a carefully orchestrated flow spanning the CLI client, HTTP API, and backend indexing services. Understanding this data flow helps developers debug import issues, optimize embedding performance, and extend the ingestion pipeline for custom document formats.

## Step 1: CLI Argument Parsing and HTTP Request Construction

The ingestion process begins in [`cmd/documents.go`](https://github.com/asciimoo/hister/blob/main/cmd/documents.go), where the CLI parses the `hister add` or `hister import` subcommands. The client reads local files or fetches remote URLs, populating a `client.Document` struct with the URL and content fields.

The `client.AddDocument` method in [`client/document.go`](https://github.com/asciimoo/hister/blob/main/client/document.go) marshals this struct into JSON and issues an HTTP **POST** request to the server’s `/api/documents` endpoint:

```go
// cmd/documents.go – part of the import sub-command
func runImport(cmd *cobra.Command, args []string) {
    path := args[0]
    content, _ := os.ReadFile(path)
    doc := client.Document{
        URL:     "file://" + path,
        Content: string(content),
    }
    c := newClient()
    if err := c.AddDocument(&doc); err != nil {
        exit(1, "add failed: "+err.Error())
    }
    fmt.Println("Document added:", doc.URL)
}

```

This client-side abstraction ensures that both direct API consumers and CLI users trigger identical server-side processing logic.

## Step 2: HTTP Handler and Request Validation

The server receives the request in [`server/api.go`](https://github.com/asciimoo/hister/blob/main/server/api.go) through the `addDocumentHandler` function. This handler unmarshals the JSON body into a `document.Document` model and validates the incoming payload before triggering the parsing pipeline:

```go
// server/api.go – addDocumentHandler
func (s *Server) addDocumentHandler(w http.ResponseWriter, r *http.Request) {
    var d document.Document
    if err := json.NewDecoder(r.Body).Decode(&d); err != nil {
        http.Error(w, "bad request", http.StatusBadRequest)
        return
    }
    // parse / clean HTML, detect language
    if err := d.LoadFromHTML(); err != nil {
        http.Error(w, err.Error(), http.StatusUnprocessableEntity)
        return
    }
    // store in DB and enqueue embedding
    if err := s.Indexer.Add(&d); err != nil {
        http.Error(w, err.Error(), http.StatusInternalServerError)
        return
    }
    json.NewEncoder(w).Encode(map[string]any{
        "id":    d.ID,
        "url":   d.URL,
        "added": d.Added,
    })
}

```

The handler acts as the gatekeeper, rejecting malformed requests before they reach the indexing subsystem.

## Step 3: Document Parsing and HTML Normalization

Once validated, the raw content flows into the document processing layer defined in [`server/document/fromhtml.go`](https://github.com/asciimoo/hister/blob/main/server/document/fromhtml.go). The `LoadFromHTML()` method extracts the title, strips HTML tags, and normalizes the text content into a clean, indexable format.

This transformation converts unstructured web pages or rich text into a plain-text representation suitable for full-text search indexing while preserving metadata like the original URL and extracted title.

## Step 4: Language Detection and Localization

After parsing, [`server/document/language.go`](https://github.com/asciimoo/hister/blob/main/server/document/language.go) executes the `DetectLanguage` function to populate the `Language` field of the document struct. This lightweight detection mechanism identifies the document’s primary language, which drives language-specific indexing strategies and enables filtered search queries by language code later in the retrieval phase.

## Step 5: SQLite Database Indexing and Persistence

The fully populated `Document` struct moves to [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go), where the `Indexer.Add()` method executes the core persistence logic. This component writes the document record into the SQLite database across multiple tables including `documents` and `versions`, registering searchable fields such as URL, title, text content, and tags.

The indexer maintains transactional integrity during this write operation, ensuring that document metadata and content remain consistent before embedding generation begins.

## Step 6: Asynchronous Embedding Generation

Following the database insert, the indexer enqueues an embedding job handled by [`server/vectorstore/vectorstore.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/vectorstore.go). The `VectorStore` fetches the document text and calls the configured embedder— implemented via OpenAI or Ollama APIs— to generate a high-dimensional vector representation of the content.

The `AddEmbedding` method stores the resulting vector in the `vectors` table, enabling semantic similarity search:

```go
// server/vectorstore/vectorstore.go – AddEmbedding
func (vs *VectorStore) AddEmbedding(docID int64, text string) error {
    vec, err := vs.embedder.Embed(text) // calls external LLM provider
    if err != nil {
        return err
    }
    _, err = vs.db.Exec(`INSERT INTO vectors (doc_id, vec) VALUES (?, ?)`,
        docID, vec)
    return err
}

```

This step decouples vector generation from the HTTP response, allowing the API to return quickly while background workers process the computationally expensive embedding operations.

## Step 7: Search Schema Registration

Finally, the indexer updates the `searchschema.Schema` defined in [`server/indexer/searchschema/schema.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/searchschema/schema.go). This registration makes the new document discoverable through structured search queries by exposing fields like `added`, `language`, and `domain` to the query engine.

The schema update ensures that immediately after ingestion, the document appears in filtered searches and full-text queries without requiring a manual index rebuild.

## Summary

- **The data flow for adding a document to Hister** begins with CLI parsing in [`cmd/documents.go`](https://github.com/asciimoo/hister/blob/main/cmd/documents.go) and ends with search schema registration, traversing eight distinct pipeline stages.
- **HTTP transport** bridges the client and server via the `/api/documents` endpoint handled by `addDocumentHandler` in [`server/api.go`](https://github.com/asciimoo/hister/blob/main/server/api.go).
- **Content normalization** occurs in [`server/document/fromhtml.go`](https://github.com/asciimoo/hister/blob/main/server/document/fromhtml.go) through the `LoadFromHTML()` method, which extracts titles and cleans HTML.
- **Language detection** populates the `Language` field using logic in [`server/document/language.go`](https://github.com/asciimoo/hister/blob/main/server/document/language.go) to enable locale-aware indexing.
- **Persistence** happens in SQLite through [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go), which manages the `documents`, `versions`, and searchable field tables.
- **Vector embeddings** are generated asynchronously by [`server/vectorstore/vectorstore.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/vectorstore.go) using external LLM providers and stored in the `vectors` table.
- **Searchability** is activated when [`server/indexer/searchschema/schema.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/searchschema/schema.go) registers the document’s metadata fields for filtered queries.

## Frequently Asked Questions

### What database does Hister use for document storage?

Hister uses **SQLite** as its primary persistence layer. The `Indexer.Add()` method in [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go) writes document records into tables including `documents`, `versions`, and `vectors`, providing ACID compliance and zero-configuration deployment.

### How does Hister generate embeddings for documents?

The system generates embeddings through the `VectorStore.AddEmbedding()` function in [`server/vectorstore/vectorstore.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/vectorstore.go). This component calls an external embedder interface— configured for OpenAI or Ollama APIs— to convert document text into high-dimensional vectors, which are then stored in the `vectors` table for semantic similarity search.

### What file formats does Hister support when adding documents?

According to the source code in [`server/document/fromhtml.go`](https://github.com/asciimoo/hister/blob/main/server/document/fromhtml.go), Hister processes raw HTML and plain text content. The `LoadFromHTML()` method specifically handles HTML parsing to extract titles and clean markup, while the CLI in [`cmd/documents.go`](https://github.com/asciimoo/hister/blob/main/cmd/documents.go) can read arbitrary files as byte content for text-based ingestion.

### Is the embedding generation synchronous or asynchronous?

Embedding generation is **asynchronous**. While [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go) initiates the process immediately after the database write, the HTTP handler returns the document ID to the client before vector generation completes. This decoupling prevents API timeouts during slow LLM provider calls while ensuring vectors are eventually stored via the `VectorStore` background workflow.