# What File Types Does Hister Support for Local File Indexing?

> Discover the file types Hister supports for local file indexing including PDF DOCX Markdown Org-mode and UTF-8 text files for efficient content search.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: how-to-guide
- Published: 2026-08-27

---

**Hister natively indexes PDF, DOCX, Markdown, Org-mode, and any UTF-8 encoded text file by routing each extension to a dedicated handler that extracts content for the Bleve search index.**

Hister is an open-source local document indexer from the `asciimoo/hister` repository that transforms static files into searchable content. According to the source code, it uses a handler-based architecture to recognize specific file extensions and parse them accordingly, while gracefully falling back to plain-text extraction for any valid UTF-8 encoded file.

## Core Supported File Types

Hister implements specialized handlers for common document formats. Each handler lives in [`server/indexer/filetypes.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/filetypes.go) and is selected based on file extension matching.

### PDF Documents (`.pdf`)

The `pdfFileType` handler in [`server/indexer/pdf.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/pdf.go) processes Adobe PDF files using the `github.com/asciimoo/pdf` library. It extracts plain text from the binary PDF structure and automatically injects a `type=pdf` metadata field into the indexed document.

### Microsoft Word Documents (`.docx`)

For modern Word files, the `docxFileType` handler in [`server/indexer/docx.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/docx.go) parses the OOXML archive structure. It extracts the document's textual content while ignoring formatting metadata, making Word documents fully searchable alongside other file types.

### Markdown Files (`.md`, `.markdown`)

The `markdownFileType` handler in [`server/indexer/markdown.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/markdown.go) recognizes both `.md` and `.markdown` extensions. Rather than rendering the Markdown to HTML, Hister treats these as plain-text documents, preserving the original markup syntax in the search index so you can query for specific formatting characters if needed.

### Org-Mode Files (`.org`)

Emacs Org-mode files are handled by `orgFileType` in [`server/indexer/org.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/org.go). This extractor focuses on the body text of Org files, stripping structural metadata while preserving the content for full-text search within the Bleve index.

### Plain Text and Source Code (Fallback)

Any file that does not match the specific handlers above is passed to `plainTextFileType` in [`server/indexer/filetypes.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/filetypes.go). This fallback handler validates that the file is valid UTF-8; if so, it stores the raw bytes as the document's text. This provides out-of-the-box support for:

- Configuration files (`.yaml`, `.json`, `.toml`)
- Log files (`.log`)
- Programming source code (`.go`, `.js`, `.py`, `.rs`, etc.)
- Generic text files (`.txt`)

Binary files that fail UTF-8 validation are rejected and excluded from the index.

## How File Type Detection Works

The selection logic resides in [`server/indexer/filetypes.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/filetypes.go) within the `PrepareFileContent` function. The indexer iterates over handlers in a specific priority order: PDF, DOCX, Markdown, Org-mode, and finally plain text. The first matching handler processes the file; if none match, the system defaults to the UTF-8 plain-text validator.

This architecture means you can drop an entire directory of mixed content into Hister, and it will automatically apply the correct parser for each file while still indexing your source code and configuration files as searchable text.

## Indexing Files by Type

You can index entire directories or restrict Hister to specific file patterns using command-line flags:

```bash

# Recursively index all supported file types in a directory

hister index --recursive /path/to/my/documents

# Index only Markdown and Org-mode files

hister index --recursive --allowed-pattern="\.md$|\.org$" /path/to/notes

```

When executed, the `index` command invokes `PrepareFileContent` for each file, which delegates to the appropriate handler (`pdfFileType`, `docxFileType`, `markdownFileType`, `orgFileType`, or `plainTextFileType`). The extracted text is then wrapped in a `document.Document` struct and added to the Bleve search index.

## Summary

- **PDF** (`.pdf`): Extracted via `pdfFileType` using the `asciimoo/pdf` library with `type=pdf` metadata
- **Word** (`.docx`): Parsed by `docxFileType` from the OOXML archive in [`server/indexer/docx.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/docx.go)
- **Markdown** (`.md`, `.markdown`): Handled by `markdownFileType` in [`server/indexer/markdown.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/markdown.go)
- **Org-mode** (`.org`): Processed by `orgFileType` in [`server/indexer/org.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/org.go)
- **Plain text**: Fallback `plainTextFileType` in [`server/indexer/filetypes.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/filetypes.go) indexes any UTF-8 file, including source code

## Frequently Asked Questions

### Does Hister support older Microsoft Word .doc files?

No. According to the source code in [`server/indexer/filetypes.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/filetypes.go), Hister only recognizes the modern `.docx` format via the `docxFileType` handler. Legacy binary `.doc` files are not supported and will be treated as binary data, causing them to be rejected during the UTF-8 validation step.

### Can I index programming source code files?

Yes. Because Hister treats any UTF-8 encoded file as plain text, you can index source code files with extensions like `.go`, `.js`, `.py`, `.rs`, or `.java`. The `plainTextFileType` handler in [`server/indexer/filetypes.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/filetypes.go) validates the encoding and stores the raw file content, making your entire codebase searchable alongside documentation.

### What happens if I try to index a binary file like an image or video?

Binary files that do not match specific handlers (PDF or DOCX) and fail UTF-8 validation are silently rejected. The `plainTextFileType` handler checks for valid UTF-8 encoding before indexing; if the check fails, the file is skipped and not added to the Bleve search index.

### How do I index only specific file types?

Use the `--allowed-pattern` flag with a regular expression to filter files during indexing. For example, `hister index --recursive --allowed-pattern="\.pdf$|\.md$"` will only process PDF and Markdown files, ignoring all other extensions even if they are normally supported.