How Search Queries Are Processed in Hister: From CLI to Vector Retrieval
Hister processes search queries through a three-layer pipeline where CLI arguments become an indexer.Query object, travel via HTTP GET to the server, and execute vector-based document retrieval against the embedded document index.
Hister is an open-source search engine that routes user queries from a command-line interface through a structured pipeline to perform semantic document retrieval. Understanding how search queries are processed in Hister requires tracing the journey from argument parsing in cmd/search.go through HTTP transport to vector-based matching in the server-side indexer package.
Architecture Overview: The Query Processing Pipeline
The search flow follows a deterministic path across three architectural layers. First, the CLI gathers arguments and flags into a structured query object. Next, the client layer marshals this data for HTTP transport. Finally, the server unmarshals the request, executes the search against vector embeddings, and returns paginated results.
The complete data flow traverses these components in sequence:
CLI args → indexer.Query → client.Search (GET /search) → server handler → indexer.Search → Results → CLI output
This architecture cleanly separates concerns: the CLI handles argument parsing, the client manages network transport, and the server performs the computational heavy lifting of vector similarity search.
Building the Query on the CLI Side
The process begins in cmd/search.go where user input transforms into a structured search request through argument parsing and flag extraction.
Parsing Arguments and Flags
Lines 60‑71 of cmd/search.go handle argument concatenation and flag processing. The command joins positional arguments into a single query string using strings.Join(args, " "). Flags such as --format, --limit, --sort, and --fields configure request behavior and output formatting.
The searchFilterMap helper (lines 36‑45) processes field selections, setting includeHTML to true when the user explicitly requests the html field in the --fields flag.
// Inside cmd/search.go – building the Query
qs := strings.Join(args, " ")
includeHTML := false
if fieldsRaw, _ := cmd.Flags().GetString("fields"); strings.Contains(fieldsRaw, "html") {
includeHTML = true
}
c := newClient()
res, err := c.Search(&indexer.Query{
Text: qs,
IncludeHTML: includeHTML,
PageKey: "", // first page
Sort: "date", // from --sort flag
})
Constructing the indexer.Query Object
Once parsed, the CLI creates an indexer.Query struct populated with the search text, pagination key, sort mode, and HTML inclusion flag. This object serves as the canonical representation of the user's request throughout the system, ensuring type safety across the client-server boundary.
Transporting the Query to the Server
The client layer handles serialization and network transfer in client/search.go, abstracting the HTTP communication details from the CLI.
JSON Marshaling and HTTP GET Request
The client.Search function (lines 13‑23) marshals the indexer.Query to JSON using json.Marshal(q), URL-encodes the payload, and issues a GET request to the server's /search endpoint. This design embeds the query as a query parameter rather than a request body, making the requests cacheable and bookmarkable.
// Inside client/search.go – sending the request
func (c *Client) Search(q *indexer.Query) (_ *indexer.Results, err error) {
qJSON, err := json.Marshal(q)
if err != nil { return nil, err }
u := "/search?query=" + url.QueryEscape(string(qJSON))
req, err := c.newRequest("GET", u, nil)
// … perform request and decode response …
}
Server-Side Query Processing
Upon arrival, the server decodes the request and executes the core search logic through the indexer package.
Request Handling and Routing
The HTTP router registered in server/api.go maps the /search path to a handler that unmarshals the query JSON back into an indexer.Query struct. This handler immediately delegates to the core search routine, passing the structured query to indexer.Search.
Core Search Execution and Vector Retrieval
The indexer.Search function performs the computational work. It first tokenizes the query text using indexer.Analyzer, then queries the vector store (vectorstore) for nearest-neighbor embeddings matching the query vector. This vector-based approach, implemented in the server-side indexing engine, enables semantic search capabilities beyond exact keyword matching.
The function retrieves matching documents from server/document/document.go, which defines the Document struct containing fields such as Title, URL, Score, and HTML.
Sorting and Pagination
Results are ordered according to the requested sort mode: relevance (default vector similarity), date, domain, or visits. The PageKey field enables cursor-based pagination, allowing the client to request subsequent result pages without offset-based performance degradation. When IncludeHTML is true, the server includes raw HTML content from the document store in the response payload.
Formatting and Returning Results
The server serializes the indexer.Results struct—containing a slice of document.Document objects and a new PageKey for the next page—to JSON for the return journey.
Back in the CLI, lines 102‑174 of cmd/search.go handle output formatting. A loop iterates through paginated results, respects the --limit flag, and streams formatted output to stdout via printDoc, supporting JSON, CSV, or plain text formats based on the --format flag.
// Example: Running a search from the CLI
// $ hister search --format json --limit 5 --sort date "golang concurrency"
Summary
- CLI Parsing: Hister converts CLI arguments into an
indexer.Querystruct incmd/search.go, processing flags like--sortand--fieldsbefore transmission. - HTTP Transport: The client marshals queries to JSON and transmits them via HTTP GET to
/searchas implemented inclient/search.go. - Vector Retrieval: Server-side processing in the
indexerpackage tokenizes queries usingindexer.Analyzerand performs vector-based retrieval againstserver/vectorstore/vectorstore.go. - Sorting and Pagination: Results support four sort modes (relevance, date, domain, visits) and cursor-based pagination via the
PageKeyfield. - Output Formatting: The CLI layer handles final rendering in JSON, CSV, or plain text, iterating through paginated results while respecting the
--limitconstraint.
Frequently Asked Questions
What format does Hister use to transmit search queries between client and server?
Hister transmits queries as URL-encoded JSON payloads via HTTP GET requests to the /search endpoint. The client.Search function in client/search.go marshals the indexer.Query struct and embeds it as a query parameter, enabling stateless, cacheable search requests that can be bookmarked or cached by intermediate proxies.
How does Hister handle pagination in search results?
The server implements cursor-based pagination using the PageKey field in the indexer.Query struct. When the server returns results, it includes a new PageKey value that the client passes in subsequent requests to retrieve the next page. This approach avoids the performance degradation associated with offset-based queries when retrieving deep result sets from the vector store.
What sorting options are available when processing search queries in Hister?
According to the asciimoo/hister source code, Hister supports four sort modes: relevance (default vector similarity), date (chronological ordering), domain (grouping by source domain), and visits (popularity-based ranking). These are specified via the --sort CLI flag and processed in the core indexer.Search routine after vector retrieval completes.
How does Hister determine whether to include raw HTML in search results?
The CLI checks the --fields flag for the presence of "html" in cmd/search.go (lines 36‑45). If detected, it sets IncludeHTML to true on the indexer.Query object. The server includes raw document HTML in the response only when this boolean flag is set, reducing payload size and processing overhead for standard searches that do not require full document content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →