Complete Guide to Hister CLI Commands for Server and Data Management
The hister binary exposes sub-commands for server lifecycle (listen), data ingestion (index, import), index maintenance (reindex, cleanup), and administrative tasks, all registered via the Cobra framework in cmd/root.go.
The asciimoo/hister project provides a self-hosted document indexing and search platform controlled entirely through a command-line interface. Mastering the Hister CLI commands allows administrators to launch services, recursively crawl websites, import bookmark collections, and perform essential housekeeping without writing custom scripts.
Server Lifecycle Management
Starting the HTTP Server with listen
The listen sub-command boots the HTTP API and optional file system watchers. Defined in cmd/root.go within the command run block (lines 59-63), it delegates to server.Listen in server/server.go to initialize the indexer, configure routes for document CRUD and search endpoints, and bind to the address specified in cfg.Server.Address.
To start the server on a custom port with public access:
hister listen --address 0.0.0.0:8080 --public
Indexing and Crawling Commands
One-Off URL Indexing
The index command, implemented in cmd/index.go, adds individual URLs to the search index. It checks for duplicates using the crawlerSkipOptions logic (lines 97-111) unless the --force flag is provided, and submits documents via client.AddDocumentJSON.
# Index a single URL (skips if exists)
hister index https://example.com
# Force re-indexing to overwrite existing entries
hister index --force https://example.com
Recursive and Persistent Crawling
Supplying the --recursive flag transforms the operation into a persistent crawl job stored in model.CrawlJob. The crawl loop resides in crawlAndIndex (lines 58-81), which respects crawlerSkipOptions for duplicate detection. Administrators can resume interrupted jobs by specifying an existing --job-id without the --recursive flag.
# Start a recursive crawl with a specific job ID
hister index --recursive --job-id myjob https://myblog.com
# Resume the crawl later from the saved queue state
hister index --job-id myjob
Managing Crawl Jobs and Reindexing
The crawl command inspects and manages persistent crawl jobs (list, show, delete), while reindex rebuilds the entire index—essential after modifying analyzer settings or language detection configurations.
# Rebuild the index after configuration changes
hister reindex
Data Import and Export Workflows
Importing from External Services
The import command aggregates service-specific importers for Linkding, Wallabag, Shaarli, and browsers. Registered in cmd/root.go (lines 82-86), each sub-command (e.g., import linkding) accepts flags defined by addServiceImportFlags (lines 29-34). The import flow constructs a crawler backend via applyCrawlerBackendFlags (supporting HTTP, ChromeDP, or Bidi), retrieves bookmarks from the external API, and pipes them through the standard indexing pipeline.
# Import from Linkding using an environment variable for the API token
hister import linkding --api-token "$LINKDING_TOKEN"
Exporting Document Archives
The export command streams documents from the server to stdout or a file, supporting text, json, and csv formats. It accepts --start-date and --end-date filters and invokes the client's ExportDocuments endpoint.
# Export the entire index to a JSON backup file
hister export --format json > hister-backup.json
Maintenance and Administrative Tasks
Document Cleanup and Deletion
The delete command removes documents matching query parameters, requiring explicit confirmation or the --yes flag for safety. The --dry flag enables simulation mode to preview deletions. The cleanup command performs deeper housekeeping, purging orphaned files, expired crawl jobs, and stale metadata from the database.
# Simulate deletion of documents older than 2022-01-01
hister delete --end-date 2022-01-01 --dry
# Prune orphaned data and expired jobs
hister cleanup
User Management and System Utilities
Administrators can manage local users through create-user, show-user, update-user, and delete-user commands. The search command allows direct CLI queries against the index, and check-update queries the GitHub releases feed to verify version currency.
Global Configuration and Client Architecture
All sub-commands share global flags defined in cmd/root.go, including --config for configuration files, --log-level for verbosity, and --token for authentication. Each command constructs an RPC client using the newClient helper (lines 78-89) defined in the same file, ensuring consistent communication with the server via client/client.go.
Summary
- The
listencommand incmd/root.goinitializes the HTTP server and indexer, binding tocfg.Server.Address. - Use
indexfor one-off URL addition or persistent recursive crawling, with duplicate detection handled bycrawlerSkipOptionsincmd/index.go. - The
importcommand supports multiple bookmark services via flags defined inaddServiceImportFlags, using crawler backends configured throughapplyCrawlerBackendFlags. - Maintenance tasks like
reindex,cleanup, anddeleteutilize the sharednewClientRPC helper to interact with the server. - Global flags
--config,--log-level, and--tokenapply to all sub-commands in the Cobra-based CLI structure.
Frequently Asked Questions
How do I resume an interrupted crawl job in Hister?
Supply the existing --job-id to the index command without the --recursive flag. Hister loads the persisted queue state from the database and continues processing URLs from where the previous execution stopped.
What is the difference between cleanup and delete commands?
The delete command removes specific documents matching query criteria (e.g., date ranges), while cleanup performs system-wide housekeeping by purging orphaned files, expired crawl jobs, and stale metadata that no longer reference valid documents.
Can I import bookmarks from services not listed in the Hister documentation?
Yes, if the service provides a standard export format (like HTML bookmarks or API access). The import command structure in cmd/root.go is extensible, allowing developers to add new service importers using the existing addServiceImportFlags pattern and applyCrawlerBackendFlags for content retrieval.
Where are the HTTP server routes defined when running hister listen?
The listen command delegates to server.Listen in server/server.go, which configures routes for search, document CRUD operations, and user management, while cmd/root.go (lines 59-63) handles the command initialization and indexer setup.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →