# How Are Local Files Indexed by Hister? A Deep Dive into the Indexing Pipeline

> Discover how Hister indexes local files. Learn about its three-stage pipeline: file discovery, content extraction with parsers, and storage in Bleve search indexes.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: deep-dive
- Published: 2026-09-01

---

**Hister indexes local files by converting them into `RemoteFile` documents through a three-stage pipeline: discovering files via directory walking, extracting content using MIME-type parsers, and storing results in Bleve search indexes under a custom `remote-file://` URL scheme.**

Hister unifies local filesystem content with web pages in a single searchable index. Understanding how local files are indexed by Hister requires examining the tight integration between the command-line import utilities, the filesystem watcher, and the core server indexing engine in the `asciimoo/hister` repository.

## Stage 1: Discovering Files in Watched Directories

The indexing process begins in **[`cmd/import_file.go`](https://github.com/asciimoo/hister/blob/main/cmd/import_file.go)**, where the `expandImportInputs` function (lines 72-106) orchestrates file discovery. This stage translates user input—either explicit paths or configured watched directories—into a validated list of import candidates.

### Directory Walking and Filtering

When you run `hister import file`, the system uses `filepath.WalkDir` to traverse directories recursively. The walker applies exclusion logic defined in **[`files/files.go`](https://github.com/asciimoo/hister/blob/main/files/files.go)**:

- **`ShouldSkipDir`** implements default exclusions for directories like `node_modules`, `__pycache__`, and `.git`, while respecting user-defined patterns.
- **`ExpandHome`** expands `~` into absolute paths.
- **`DirectoryMatchesPath`** validates that files belong to configured directories.

The function returns a slice of `importFileInput` structs containing the file path and label, ready for content extraction.

## Stage 2: Content Extraction and URL Generation

Once discovered, each file passes through `importRemoteFile` in **[`cmd/import_file.go`](https://github.com/asciimoo/hister/blob/main/cmd/import_file.go)** (lines 85-130). This stage transforms raw filesystem data into a structured `document.Document` ready for indexing.

### The Remote-File URL Scheme

Hister does not store raw filesystem paths in the index. Instead, it generates a pseudo-URL using the `remote-file://` scheme via the `remoteFileURL` function. This normalizes the source hostname and converts absolute paths to POSIX-style strings, producing identifiers like:

```

remote-file://my-laptop/home/user/documents/report.pdf

```

This URL becomes the document's unique identifier (`Document.URL`) and ensures consistency between local and web content.

### MIME-Type Extraction Pipeline

After URL generation, the system reads the file with `os.ReadFile` and passes the bytes to **`indexer.PrepareFileContent`** in [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go). This function executes `extractor.ExtractContext`, which:

- Detects MIME types to determine the appropriate parser.
- Populates `Document.Text` with extracted plain text.
- Populates `Document.HTML` with HTML representations when available.
- Optionally extracts preview images for supported formats.

## Stage 3: Storing Documents in the Bleve Index

With content extracted, the document moves from client to server for persistent storage.

### Client-Server API Communication

The `importRemoteFile` function calls `c.AddDocumentJSON(d)`, where `c` is a `client.Client` instance. This sends a `POST /documents` request to the Hister server API, serializing the `Document` struct as JSON.

### Persistent Storage Operations

On the server side, **[`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go)** handles the request through `AddDocumentContext` (lines 979-1002). This critical function:

1. Writes the document to the **Bleve** index via `bleve.Index.Index`.
2. Updates the **vector store** with embeddings for semantic search capabilities.
3. Maintains referential integrity between full-text and vector indexes.

The result is a fully searchable document accessible through Hister’s unified query interface alongside web pages.

## Real-Time Indexing with File System Watching

Beyond one-off imports, Hister supports continuous indexing via filesystem monitoring. The **[`files/files.go`](https://github.com/asciimoo/hister/blob/main/files/files.go)** module implements `WatchDirectories`, which instantiates an `fsnotify.Watcher` to monitor configured directories for write events.

### Debouncing and Automatic Re-Indexing

To prevent index thrashing during rapid file saves, the watcher implements a **200ms debounce timer** (`debounceTime = 200 ms`). When a file modification is detected, the system waits for the debounce period to elapse before triggering the same three-stage import pipeline, ensuring the index reflects the latest content without excessive CPU or I/O overhead.

## Summary

- **Discovery** uses `filepath.WalkDir` in [`cmd/import_file.go`](https://github.com/asciimoo/hister/blob/main/cmd/import_file.go) with configurable include/exclude rules via `files.ShouldSkipDir`.
- **Normalization** converts local paths to `remote-file://` URLs to create unique, host-specific document identifiers.
- **Extraction** relies on `indexer.PrepareFileContent` to parse MIME types and generate searchable text and HTML content.
- **Storage** persists documents in Bleve indexes through `AddDocumentContext`, with optional vector embeddings for semantic search.
- **Watching** uses `fsnotify` with 200ms debouncing to maintain index freshness automatically.

## Frequently Asked Questions

### What URL scheme does Hister use to identify local files?

Hister uses the `remote-file://` scheme to identify local documents. As implemented in [`cmd/import_file.go`](https://github.com/asciimoo/hister/blob/main/cmd/import_file.go), the `remoteFileURL` function normalizes the hostname and filepath to create unique identifiers like `remote-file://my-laptop/path/to/file.txt`, ensuring local files coexist seamlessly with web URLs in the search index.

### How does Hister handle file modifications after the initial import?

Hister monitors configured directories using `fsnotify.Watcher` implemented in [`files/files.go`](https://github.com/asciimoo/hister/blob/main/files/files.go). When files change, the system applies a 200ms debounce timer to batch rapid successive writes, then automatically re-runs the three-stage import pipeline to update the Bleve index with the latest content.

### What file types can Hister index from the local filesystem?

Hister can index any file type supported by the extractor pipeline defined in [`server/indexer/indexer.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/indexer.go). The `PrepareFileContent` function delegates to `extractor.ExtractContext`, which selects parsers based on MIME type detection, enabling text extraction from PDFs, Office documents, Markdown, code files, and other formats.

### Where is the local file indexing configuration defined?

Directorywatching configuration is defined in **[`config/config.go`](https://github.com/asciimoo/hister/blob/main/config/config.go)** through `Directory` structs that specify paths, labels, and exclusion patterns. The CLI import logic in **[`cmd/import_file.go`](https://github.com/asciimoo/hister/blob/main/cmd/import_file.go)** reads these configurations when no explicit paths are provided to the `hister import file` command.