How Are Local Files Indexed by Hister? A Deep Dive into the Indexing Pipeline
Hister indexes local files by converting them into RemoteFile documents through a three-stage pipeline: discovering files via directory walking, extracting content using MIME-type parsers, and storing results in Bleve search indexes under a custom remote-file:// URL scheme.
Hister unifies local filesystem content with web pages in a single searchable index. Understanding how local files are indexed by Hister requires examining the tight integration between the command-line import utilities, the filesystem watcher, and the core server indexing engine in the asciimoo/hister repository.
Stage 1: Discovering Files in Watched Directories
The indexing process begins in cmd/import_file.go, where the expandImportInputs function (lines 72-106) orchestrates file discovery. This stage translates user input—either explicit paths or configured watched directories—into a validated list of import candidates.
Directory Walking and Filtering
When you run hister import file, the system uses filepath.WalkDir to traverse directories recursively. The walker applies exclusion logic defined in files/files.go:
ShouldSkipDirimplements default exclusions for directories likenode_modules,__pycache__, and.git, while respecting user-defined patterns.ExpandHomeexpands~into absolute paths.DirectoryMatchesPathvalidates that files belong to configured directories.
The function returns a slice of importFileInput structs containing the file path and label, ready for content extraction.
Stage 2: Content Extraction and URL Generation
Once discovered, each file passes through importRemoteFile in cmd/import_file.go (lines 85-130). This stage transforms raw filesystem data into a structured document.Document ready for indexing.
The Remote-File URL Scheme
Hister does not store raw filesystem paths in the index. Instead, it generates a pseudo-URL using the remote-file:// scheme via the remoteFileURL function. This normalizes the source hostname and converts absolute paths to POSIX-style strings, producing identifiers like:
remote-file://my-laptop/home/user/documents/report.pdf
This URL becomes the document's unique identifier (Document.URL) and ensures consistency between local and web content.
MIME-Type Extraction Pipeline
After URL generation, the system reads the file with os.ReadFile and passes the bytes to indexer.PrepareFileContent in server/indexer/indexer.go. This function executes extractor.ExtractContext, which:
- Detects MIME types to determine the appropriate parser.
- Populates
Document.Textwith extracted plain text. - Populates
Document.HTMLwith HTML representations when available. - Optionally extracts preview images for supported formats.
Stage 3: Storing Documents in the Bleve Index
With content extracted, the document moves from client to server for persistent storage.
Client-Server API Communication
The importRemoteFile function calls c.AddDocumentJSON(d), where c is a client.Client instance. This sends a POST /documents request to the Hister server API, serializing the Document struct as JSON.
Persistent Storage Operations
On the server side, server/indexer/indexer.go handles the request through AddDocumentContext (lines 979-1002). This critical function:
- Writes the document to the Bleve index via
bleve.Index.Index. - Updates the vector store with embeddings for semantic search capabilities.
- Maintains referential integrity between full-text and vector indexes.
The result is a fully searchable document accessible through Hister’s unified query interface alongside web pages.
Real-Time Indexing with File System Watching
Beyond one-off imports, Hister supports continuous indexing via filesystem monitoring. The files/files.go module implements WatchDirectories, which instantiates an fsnotify.Watcher to monitor configured directories for write events.
Debouncing and Automatic Re-Indexing
To prevent index thrashing during rapid file saves, the watcher implements a 200ms debounce timer (debounceTime = 200 ms). When a file modification is detected, the system waits for the debounce period to elapse before triggering the same three-stage import pipeline, ensuring the index reflects the latest content without excessive CPU or I/O overhead.
Summary
- Discovery uses
filepath.WalkDirincmd/import_file.gowith configurable include/exclude rules viafiles.ShouldSkipDir. - Normalization converts local paths to
remote-file://URLs to create unique, host-specific document identifiers. - Extraction relies on
indexer.PrepareFileContentto parse MIME types and generate searchable text and HTML content. - Storage persists documents in Bleve indexes through
AddDocumentContext, with optional vector embeddings for semantic search. - Watching uses
fsnotifywith 200ms debouncing to maintain index freshness automatically.
Frequently Asked Questions
What URL scheme does Hister use to identify local files?
Hister uses the remote-file:// scheme to identify local documents. As implemented in cmd/import_file.go, the remoteFileURL function normalizes the hostname and filepath to create unique identifiers like remote-file://my-laptop/path/to/file.txt, ensuring local files coexist seamlessly with web URLs in the search index.
How does Hister handle file modifications after the initial import?
Hister monitors configured directories using fsnotify.Watcher implemented in files/files.go. When files change, the system applies a 200ms debounce timer to batch rapid successive writes, then automatically re-runs the three-stage import pipeline to update the Bleve index with the latest content.
What file types can Hister index from the local filesystem?
Hister can index any file type supported by the extractor pipeline defined in server/indexer/indexer.go. The PrepareFileContent function delegates to extractor.ExtractContext, which selects parsers based on MIME type detection, enabling text extraction from PDFs, Office documents, Markdown, code files, and other formats.
Where is the local file indexing configuration defined?
Directorywatching configuration is defined in config/config.go through Directory structs that specify paths, labels, and exclusion patterns. The CLI import logic in cmd/import_file.go reads these configurations when no explicit paths are provided to the hister import file command.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →