What File Types Does Hister Support for Local File Indexing?
Hister natively indexes PDF, DOCX, Markdown, Org-mode, and any UTF-8 encoded text file by routing each extension to a dedicated handler that extracts content for the Bleve search index.
Hister is an open-source local document indexer from the asciimoo/hister repository that transforms static files into searchable content. According to the source code, it uses a handler-based architecture to recognize specific file extensions and parse them accordingly, while gracefully falling back to plain-text extraction for any valid UTF-8 encoded file.
Core Supported File Types
Hister implements specialized handlers for common document formats. Each handler lives in server/indexer/filetypes.go and is selected based on file extension matching.
PDF Documents (.pdf)
The pdfFileType handler in server/indexer/pdf.go processes Adobe PDF files using the github.com/asciimoo/pdf library. It extracts plain text from the binary PDF structure and automatically injects a type=pdf metadata field into the indexed document.
Microsoft Word Documents (.docx)
For modern Word files, the docxFileType handler in server/indexer/docx.go parses the OOXML archive structure. It extracts the document's textual content while ignoring formatting metadata, making Word documents fully searchable alongside other file types.
Markdown Files (.md, .markdown)
The markdownFileType handler in server/indexer/markdown.go recognizes both .md and .markdown extensions. Rather than rendering the Markdown to HTML, Hister treats these as plain-text documents, preserving the original markup syntax in the search index so you can query for specific formatting characters if needed.
Org-Mode Files (.org)
Emacs Org-mode files are handled by orgFileType in server/indexer/org.go. This extractor focuses on the body text of Org files, stripping structural metadata while preserving the content for full-text search within the Bleve index.
Plain Text and Source Code (Fallback)
Any file that does not match the specific handlers above is passed to plainTextFileType in server/indexer/filetypes.go. This fallback handler validates that the file is valid UTF-8; if so, it stores the raw bytes as the document's text. This provides out-of-the-box support for:
- Configuration files (
.yaml,.json,.toml) - Log files (
.log) - Programming source code (
.go,.js,.py,.rs, etc.) - Generic text files (
.txt)
Binary files that fail UTF-8 validation are rejected and excluded from the index.
How File Type Detection Works
The selection logic resides in server/indexer/filetypes.go within the PrepareFileContent function. The indexer iterates over handlers in a specific priority order: PDF, DOCX, Markdown, Org-mode, and finally plain text. The first matching handler processes the file; if none match, the system defaults to the UTF-8 plain-text validator.
This architecture means you can drop an entire directory of mixed content into Hister, and it will automatically apply the correct parser for each file while still indexing your source code and configuration files as searchable text.
Indexing Files by Type
You can index entire directories or restrict Hister to specific file patterns using command-line flags:
# Recursively index all supported file types in a directory
hister index --recursive /path/to/my/documents
# Index only Markdown and Org-mode files
hister index --recursive --allowed-pattern="\.md$|\.org$" /path/to/notes
When executed, the index command invokes PrepareFileContent for each file, which delegates to the appropriate handler (pdfFileType, docxFileType, markdownFileType, orgFileType, or plainTextFileType). The extracted text is then wrapped in a document.Document struct and added to the Bleve search index.
Summary
- PDF (
.pdf): Extracted viapdfFileTypeusing theasciimoo/pdflibrary withtype=pdfmetadata - Word (
.docx): Parsed bydocxFileTypefrom the OOXML archive inserver/indexer/docx.go - Markdown (
.md,.markdown): Handled bymarkdownFileTypeinserver/indexer/markdown.go - Org-mode (
.org): Processed byorgFileTypeinserver/indexer/org.go - Plain text: Fallback
plainTextFileTypeinserver/indexer/filetypes.goindexes any UTF-8 file, including source code
Frequently Asked Questions
Does Hister support older Microsoft Word .doc files?
No. According to the source code in server/indexer/filetypes.go, Hister only recognizes the modern .docx format via the docxFileType handler. Legacy binary .doc files are not supported and will be treated as binary data, causing them to be rejected during the UTF-8 validation step.
Can I index programming source code files?
Yes. Because Hister treats any UTF-8 encoded file as plain text, you can index source code files with extensions like .go, .js, .py, .rs, or .java. The plainTextFileType handler in server/indexer/filetypes.go validates the encoding and stores the raw file content, making your entire codebase searchable alongside documentation.
What happens if I try to index a binary file like an image or video?
Binary files that do not match specific handlers (PDF or DOCX) and fail UTF-8 validation are silently rejected. The plainTextFileType handler checks for valid UTF-8 encoding before indexing; if the check fails, the file is skipped and not added to the Bleve search index.
How do I index only specific file types?
Use the --allowed-pattern flag with a regular expression to filter files during indexing. For example, hister index --recursive --allowed-pattern="\.pdf$|\.md$" will only process PDF and Markdown files, ignoring all other extensions even if they are normally supported.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →