Data Flow for Adding a Document to Hister: Complete Pipeline Guide

When you add a document to Hister, the system executes an eight-stage pipeline that transforms raw input into a searchable database entry with semantic embeddings, persisting data in SQLite and updating the search schema for immediate retrieval.

The open-source search engine asciimoo/hister processes document ingestion through a carefully orchestrated flow spanning the CLI client, HTTP API, and backend indexing services. Understanding this data flow helps developers debug import issues, optimize embedding performance, and extend the ingestion pipeline for custom document formats.

Step 1: CLI Argument Parsing and HTTP Request Construction

The ingestion process begins in cmd/documents.go, where the CLI parses the hister add or hister import subcommands. The client reads local files or fetches remote URLs, populating a client.Document struct with the URL and content fields.

The client.AddDocument method in client/document.go marshals this struct into JSON and issues an HTTP POST request to the server’s /api/documents endpoint:

// cmd/documents.go – part of the import sub-command
func runImport(cmd *cobra.Command, args []string) {
    path := args[0]
    content, _ := os.ReadFile(path)
    doc := client.Document{
        URL:     "file://" + path,
        Content: string(content),
    }
    c := newClient()
    if err := c.AddDocument(&doc); err != nil {
        exit(1, "add failed: "+err.Error())
    }
    fmt.Println("Document added:", doc.URL)
}

This client-side abstraction ensures that both direct API consumers and CLI users trigger identical server-side processing logic.

Step 2: HTTP Handler and Request Validation

The server receives the request in server/api.go through the addDocumentHandler function. This handler unmarshals the JSON body into a document.Document model and validates the incoming payload before triggering the parsing pipeline:

// server/api.go – addDocumentHandler
func (s *Server) addDocumentHandler(w http.ResponseWriter, r *http.Request) {
    var d document.Document
    if err := json.NewDecoder(r.Body).Decode(&d); err != nil {
        http.Error(w, "bad request", http.StatusBadRequest)
        return
    }
    // parse / clean HTML, detect language
    if err := d.LoadFromHTML(); err != nil {
        http.Error(w, err.Error(), http.StatusUnprocessableEntity)
        return
    }
    // store in DB and enqueue embedding
    if err := s.Indexer.Add(&d); err != nil {
        http.Error(w, err.Error(), http.StatusInternalServerError)
        return
    }
    json.NewEncoder(w).Encode(map[string]any{
        "id":    d.ID,
        "url":   d.URL,
        "added": d.Added,
    })
}

The handler acts as the gatekeeper, rejecting malformed requests before they reach the indexing subsystem.

Step 3: Document Parsing and HTML Normalization

Once validated, the raw content flows into the document processing layer defined in server/document/fromhtml.go. The LoadFromHTML() method extracts the title, strips HTML tags, and normalizes the text content into a clean, indexable format.

This transformation converts unstructured web pages or rich text into a plain-text representation suitable for full-text search indexing while preserving metadata like the original URL and extracted title.

Step 4: Language Detection and Localization

After parsing, server/document/language.go executes the DetectLanguage function to populate the Language field of the document struct. This lightweight detection mechanism identifies the document’s primary language, which drives language-specific indexing strategies and enables filtered search queries by language code later in the retrieval phase.

Step 5: SQLite Database Indexing and Persistence

The fully populated Document struct moves to server/indexer/indexer.go, where the Indexer.Add() method executes the core persistence logic. This component writes the document record into the SQLite database across multiple tables including documents and versions, registering searchable fields such as URL, title, text content, and tags.

The indexer maintains transactional integrity during this write operation, ensuring that document metadata and content remain consistent before embedding generation begins.

Step 6: Asynchronous Embedding Generation

Following the database insert, the indexer enqueues an embedding job handled by server/vectorstore/vectorstore.go. The VectorStore fetches the document text and calls the configured embedder— implemented via OpenAI or Ollama APIs— to generate a high-dimensional vector representation of the content.

The AddEmbedding method stores the resulting vector in the vectors table, enabling semantic similarity search:

// server/vectorstore/vectorstore.go – AddEmbedding
func (vs *VectorStore) AddEmbedding(docID int64, text string) error {
    vec, err := vs.embedder.Embed(text) // calls external LLM provider
    if err != nil {
        return err
    }
    _, err = vs.db.Exec(`INSERT INTO vectors (doc_id, vec) VALUES (?, ?)`,
        docID, vec)
    return err
}

This step decouples vector generation from the HTTP response, allowing the API to return quickly while background workers process the computationally expensive embedding operations.

Step 7: Search Schema Registration

Finally, the indexer updates the searchschema.Schema defined in server/indexer/searchschema/schema.go. This registration makes the new document discoverable through structured search queries by exposing fields like added, language, and domain to the query engine.

The schema update ensures that immediately after ingestion, the document appears in filtered searches and full-text queries without requiring a manual index rebuild.

Summary

  • The data flow for adding a document to Hister begins with CLI parsing in cmd/documents.go and ends with search schema registration, traversing eight distinct pipeline stages.
  • HTTP transport bridges the client and server via the /api/documents endpoint handled by addDocumentHandler in server/api.go.
  • Content normalization occurs in server/document/fromhtml.go through the LoadFromHTML() method, which extracts titles and cleans HTML.
  • Language detection populates the Language field using logic in server/document/language.go to enable locale-aware indexing.
  • Persistence happens in SQLite through server/indexer/indexer.go, which manages the documents, versions, and searchable field tables.
  • Vector embeddings are generated asynchronously by server/vectorstore/vectorstore.go using external LLM providers and stored in the vectors table.
  • Searchability is activated when server/indexer/searchschema/schema.go registers the document’s metadata fields for filtered queries.

Frequently Asked Questions

What database does Hister use for document storage?

Hister uses SQLite as its primary persistence layer. The Indexer.Add() method in server/indexer/indexer.go writes document records into tables including documents, versions, and vectors, providing ACID compliance and zero-configuration deployment.

How does Hister generate embeddings for documents?

The system generates embeddings through the VectorStore.AddEmbedding() function in server/vectorstore/vectorstore.go. This component calls an external embedder interface— configured for OpenAI or Ollama APIs— to convert document text into high-dimensional vectors, which are then stored in the vectors table for semantic similarity search.

What file formats does Hister support when adding documents?

According to the source code in server/document/fromhtml.go, Hister processes raw HTML and plain text content. The LoadFromHTML() method specifically handles HTML parsing to extract titles and clean markup, while the CLI in cmd/documents.go can read arbitrary files as byte content for text-based ingestion.

Is the embedding generation synchronous or asynchronous?

Embedding generation is asynchronous. While server/indexer/indexer.go initiates the process immediately after the database write, the HTTP handler returns the document ID to the client before vector generation completes. This decoupling prevents API timeouts during slow LLM provider calls while ensuring vectors are eventually stored via the VectorStore background workflow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →