How Are Local Files Indexed by Hister? A Deep Dive into the Indexing Pipeline

Hister indexes local files by converting them into RemoteFile documents through a three-stage pipeline: discovering files via directory walking, extracting content using MIME-type parsers, and storing results in Bleve search indexes under a custom remote-file:// URL scheme.

Hister unifies local filesystem content with web pages in a single searchable index. Understanding how local files are indexed by Hister requires examining the tight integration between the command-line import utilities, the filesystem watcher, and the core server indexing engine in the asciimoo/hister repository.

Stage 1: Discovering Files in Watched Directories

The indexing process begins in cmd/import_file.go, where the expandImportInputs function (lines 72-106) orchestrates file discovery. This stage translates user input—either explicit paths or configured watched directories—into a validated list of import candidates.

Directory Walking and Filtering

When you run hister import file, the system uses filepath.WalkDir to traverse directories recursively. The walker applies exclusion logic defined in files/files.go:

  • ShouldSkipDir implements default exclusions for directories like node_modules, __pycache__, and .git, while respecting user-defined patterns.
  • ExpandHome expands ~ into absolute paths.
  • DirectoryMatchesPath validates that files belong to configured directories.

The function returns a slice of importFileInput structs containing the file path and label, ready for content extraction.

Stage 2: Content Extraction and URL Generation

Once discovered, each file passes through importRemoteFile in cmd/import_file.go (lines 85-130). This stage transforms raw filesystem data into a structured document.Document ready for indexing.

The Remote-File URL Scheme

Hister does not store raw filesystem paths in the index. Instead, it generates a pseudo-URL using the remote-file:// scheme via the remoteFileURL function. This normalizes the source hostname and converts absolute paths to POSIX-style strings, producing identifiers like:


remote-file://my-laptop/home/user/documents/report.pdf

This URL becomes the document's unique identifier (Document.URL) and ensures consistency between local and web content.

MIME-Type Extraction Pipeline

After URL generation, the system reads the file with os.ReadFile and passes the bytes to indexer.PrepareFileContent in server/indexer/indexer.go. This function executes extractor.ExtractContext, which:

  • Detects MIME types to determine the appropriate parser.
  • Populates Document.Text with extracted plain text.
  • Populates Document.HTML with HTML representations when available.
  • Optionally extracts preview images for supported formats.

Stage 3: Storing Documents in the Bleve Index

With content extracted, the document moves from client to server for persistent storage.

Client-Server API Communication

The importRemoteFile function calls c.AddDocumentJSON(d), where c is a client.Client instance. This sends a POST /documents request to the Hister server API, serializing the Document struct as JSON.

Persistent Storage Operations

On the server side, server/indexer/indexer.go handles the request through AddDocumentContext (lines 979-1002). This critical function:

  1. Writes the document to the Bleve index via bleve.Index.Index.
  2. Updates the vector store with embeddings for semantic search capabilities.
  3. Maintains referential integrity between full-text and vector indexes.

The result is a fully searchable document accessible through Hister’s unified query interface alongside web pages.

Real-Time Indexing with File System Watching

Beyond one-off imports, Hister supports continuous indexing via filesystem monitoring. The files/files.go module implements WatchDirectories, which instantiates an fsnotify.Watcher to monitor configured directories for write events.

Debouncing and Automatic Re-Indexing

To prevent index thrashing during rapid file saves, the watcher implements a 200ms debounce timer (debounceTime = 200 ms). When a file modification is detected, the system waits for the debounce period to elapse before triggering the same three-stage import pipeline, ensuring the index reflects the latest content without excessive CPU or I/O overhead.

Summary

  • Discovery uses filepath.WalkDir in cmd/import_file.go with configurable include/exclude rules via files.ShouldSkipDir.
  • Normalization converts local paths to remote-file:// URLs to create unique, host-specific document identifiers.
  • Extraction relies on indexer.PrepareFileContent to parse MIME types and generate searchable text and HTML content.
  • Storage persists documents in Bleve indexes through AddDocumentContext, with optional vector embeddings for semantic search.
  • Watching uses fsnotify with 200ms debouncing to maintain index freshness automatically.

Frequently Asked Questions

What URL scheme does Hister use to identify local files?

Hister uses the remote-file:// scheme to identify local documents. As implemented in cmd/import_file.go, the remoteFileURL function normalizes the hostname and filepath to create unique identifiers like remote-file://my-laptop/path/to/file.txt, ensuring local files coexist seamlessly with web URLs in the search index.

How does Hister handle file modifications after the initial import?

Hister monitors configured directories using fsnotify.Watcher implemented in files/files.go. When files change, the system applies a 200ms debounce timer to batch rapid successive writes, then automatically re-runs the three-stage import pipeline to update the Bleve index with the latest content.

What file types can Hister index from the local filesystem?

Hister can index any file type supported by the extractor pipeline defined in server/indexer/indexer.go. The PrepareFileContent function delegates to extractor.ExtractContext, which selects parsers based on MIME type detection, enabling text extraction from PDFs, Office documents, Markdown, code files, and other formats.

Where is the local file indexing configuration defined?

Directorywatching configuration is defined in config/config.go through Directory structs that specify paths, labels, and exclusion patterns. The CLI import logic in cmd/import_file.go reads these configurations when no explicit paths are provided to the hister import file command.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →