# How Search Queries Are Processed in Hister: From CLI to Vector Retrieval

> Discover how Hister processes search queries. Learn about the three-layer pipeline from CLI arguments to vector retrieval for efficient document searching.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: internals
- Published: 2026-09-01

---

**Hister processes search queries through a three-layer pipeline where CLI arguments become an `indexer.Query` object, travel via HTTP GET to the server, and execute vector-based document retrieval against the embedded document index.**

Hister is an open-source search engine that routes user queries from a command-line interface through a structured pipeline to perform semantic document retrieval. Understanding how search queries are processed in Hister requires tracing the journey from argument parsing in [`cmd/search.go`](https://github.com/asciimoo/hister/blob/main/cmd/search.go) through HTTP transport to vector-based matching in the server-side `indexer` package.

## Architecture Overview: The Query Processing Pipeline

The search flow follows a deterministic path across three architectural layers. First, the CLI gathers arguments and flags into a structured query object. Next, the client layer marshals this data for HTTP transport. Finally, the server unmarshals the request, executes the search against vector embeddings, and returns paginated results.

The complete data flow traverses these components in sequence:

```

CLI args → indexer.Query → client.Search (GET /search) → server handler → indexer.Search → Results → CLI output

```

This architecture cleanly separates concerns: the CLI handles argument parsing, the client manages network transport, and the server performs the computational heavy lifting of vector similarity search.

## Building the Query on the CLI Side

The process begins in [`cmd/search.go`](https://github.com/asciimoo/hister/blob/main/cmd/search.go) where user input transforms into a structured search request through argument parsing and flag extraction.

### Parsing Arguments and Flags

Lines **60‑71** of [`cmd/search.go`](https://github.com/asciimoo/hister/blob/main/cmd/search.go) handle argument concatenation and flag processing. The command joins positional arguments into a single query string using `strings.Join(args, " ")`. Flags such as `--format`, `--limit`, `--sort`, and `--fields` configure request behavior and output formatting.

The `searchFilterMap` helper (lines **36‑45**) processes field selections, setting `includeHTML` to `true` when the user explicitly requests the `html` field in the `--fields` flag.

```go
// Inside cmd/search.go – building the Query
qs := strings.Join(args, " ")
includeHTML := false
if fieldsRaw, _ := cmd.Flags().GetString("fields"); strings.Contains(fieldsRaw, "html") {
    includeHTML = true
}
c := newClient()
res, err := c.Search(&indexer.Query{
    Text:        qs,
    IncludeHTML: includeHTML,
    PageKey:     "",        // first page
    Sort:        "date",    // from --sort flag
})

```

### Constructing the indexer.Query Object

Once parsed, the CLI creates an **`indexer.Query`** struct populated with the search text, pagination key, sort mode, and HTML inclusion flag. This object serves as the canonical representation of the user's request throughout the system, ensuring type safety across the client-server boundary.

## Transporting the Query to the Server

The client layer handles serialization and network transfer in [`client/search.go`](https://github.com/asciimoo/hister/blob/main/client/search.go), abstracting the HTTP communication details from the CLI.

### JSON Marshaling and HTTP GET Request

The `client.Search` function (lines **13‑23**) marshals the `indexer.Query` to JSON using `json.Marshal(q)`, URL-encodes the payload, and issues a **GET** request to the server's `/search` endpoint. This design embeds the query as a query parameter rather than a request body, making the requests cacheable and bookmarkable.

```go
// Inside client/search.go – sending the request
func (c *Client) Search(q *indexer.Query) (_ *indexer.Results, err error) {
    qJSON, err := json.Marshal(q)
    if err != nil { return nil, err }
    u := "/search?query=" + url.QueryEscape(string(qJSON))
    req, err := c.newRequest("GET", u, nil)
    // … perform request and decode response …
}

```

## Server-Side Query Processing

Upon arrival, the server decodes the request and executes the core search logic through the `indexer` package.

### Request Handling and Routing

The HTTP router registered in [`server/api.go`](https://github.com/asciimoo/hister/blob/main/server/api.go) maps the `/search` path to a handler that unmarshals the query JSON back into an `indexer.Query` struct. This handler immediately delegates to the core search routine, passing the structured query to `indexer.Search`.

### Core Search Execution and Vector Retrieval

The **`indexer.Search`** function performs the computational work. It first tokenizes the query text using **`indexer.Analyzer`**, then queries the **vector store** (`vectorstore`) for nearest-neighbor embeddings matching the query vector. This vector-based approach, implemented in the server-side indexing engine, enables semantic search capabilities beyond exact keyword matching.

The function retrieves matching documents from [`server/document/document.go`](https://github.com/asciimoo/hister/blob/main/server/document/document.go), which defines the `Document` struct containing fields such as `Title`, `URL`, `Score`, and `HTML`.

### Sorting and Pagination

Results are ordered according to the requested **sort mode**: `relevance` (default vector similarity), `date`, `domain`, or `visits`. The **`PageKey`** field enables cursor-based pagination, allowing the client to request subsequent result pages without offset-based performance degradation. When `IncludeHTML` is true, the server includes raw HTML content from the document store in the response payload.

## Formatting and Returning Results

The server serializes the **`indexer.Results`** struct—containing a slice of `document.Document` objects and a new `PageKey` for the next page—to JSON for the return journey.

Back in the CLI, lines **102‑174** of [`cmd/search.go`](https://github.com/asciimoo/hister/blob/main/cmd/search.go) handle output formatting. A loop iterates through paginated results, respects the `--limit` flag, and streams formatted output to stdout via `printDoc`, supporting JSON, CSV, or plain text formats based on the `--format` flag.

```go
// Example: Running a search from the CLI
// $ hister search --format json --limit 5 --sort date "golang concurrency"

```

## Summary

- **CLI Parsing**: Hister converts CLI arguments into an `indexer.Query` struct in [`cmd/search.go`](https://github.com/asciimoo/hister/blob/main/cmd/search.go), processing flags like `--sort` and `--fields` before transmission.
- **HTTP Transport**: The client marshals queries to JSON and transmits them via HTTP GET to `/search` as implemented in [`client/search.go`](https://github.com/asciimoo/hister/blob/main/client/search.go).
- **Vector Retrieval**: Server-side processing in the `indexer` package tokenizes queries using `indexer.Analyzer` and performs vector-based retrieval against [`server/vectorstore/vectorstore.go`](https://github.com/asciimoo/hister/blob/main/server/vectorstore/vectorstore.go).
- **Sorting and Pagination**: Results support four sort modes (relevance, date, domain, visits) and cursor-based pagination via the `PageKey` field.
- **Output Formatting**: The CLI layer handles final rendering in JSON, CSV, or plain text, iterating through paginated results while respecting the `--limit` constraint.

## Frequently Asked Questions

### What format does Hister use to transmit search queries between client and server?

Hister transmits queries as URL-encoded JSON payloads via HTTP GET requests to the `/search` endpoint. The `client.Search` function in [`client/search.go`](https://github.com/asciimoo/hister/blob/main/client/search.go) marshals the `indexer.Query` struct and embeds it as a query parameter, enabling stateless, cacheable search requests that can be bookmarked or cached by intermediate proxies.

### How does Hister handle pagination in search results?

The server implements **cursor-based pagination** using the `PageKey` field in the `indexer.Query` struct. When the server returns results, it includes a new `PageKey` value that the client passes in subsequent requests to retrieve the next page. This approach avoids the performance degradation associated with offset-based queries when retrieving deep result sets from the vector store.

### What sorting options are available when processing search queries in Hister?

According to the `asciimoo/hister` source code, Hister supports four sort modes: **`relevance`** (default vector similarity), **`date`** (chronological ordering), **`domain`** (grouping by source domain), and **`visits`** (popularity-based ranking). These are specified via the `--sort` CLI flag and processed in the core `indexer.Search` routine after vector retrieval completes.

### How does Hister determine whether to include raw HTML in search results?

The CLI checks the `--fields` flag for the presence of "html" in [`cmd/search.go`](https://github.com/asciimoo/hister/blob/main/cmd/search.go) (lines **36‑45**). If detected, it sets `IncludeHTML` to `true` on the `indexer.Query` object. The server includes raw document HTML in the response only when this boolean flag is set, reducing payload size and processing overhead for standard searches that do not require full document content.