How Hister Manages Crawl Job State and Resumability

Hister persists crawl progress in a SQLite or PostgreSQL database using a JSON-encoded State field within the Crawl struct, enabling automatic resumption of interrupted crawls from the exact point of failure.

The asciimoo/hister search engine implements robust crawl job state management by treating each crawl as a first-class database entity. By serializing internal progress markers to a relational store via server/model/crawl.go, Hister ensures that long-running indexing operations survive process restarts, crashes, and graceful shutdowns without data loss.

Core Architecture Components

The Crawl Struct and State Persistence

In server/model/crawl.go, the Crawl struct serves as the central abstraction for job state tracking. It contains metadata fields including ID, UserID, Status, StartedAt, FinishedAt, and Error, alongside a critical State field that stores JSON-encoded progress data.

The State field tracks dynamic crawl progress such as the last processed URL, pagination tokens, or batch offsets. During execution, the crawler periodically invokes SaveState to flush the current progress to the database, creating a durable checkpoint that survives unexpected termination.

Database Schema and ACID Guarantees

The underlying schema is defined in server/model/migration.go, which creates the crawls table with a TEXT column for the JSON state snapshot. Because Hister uses standard relational backends—either SQLite or PostgreSQL—the persistence layer guarantees ACID consistency, ensuring that state updates are atomic and durable even under power loss or container restarts.

How Resumability Works

State Serialization Flow

When a crawl launches, the system inserts a new row with Status set to "running" and an empty State buffer. After processing each document batch, the crawler marshals its internal progress map into JSON and updates the State column via GORM. This incremental persistence strategy minimizes work duplication while maintaining low database write overhead.

Recovery and Resume Logic

The resumability mechanism activates when the hister crawl command is invoked with the --resume flag, implemented in cmd/crawl.go. The command queries the crawls table for incomplete jobs marked as running or paused, retrieves the stored JSON state, and deserializes it back into the crawler's memory.

On the server side, server/indexer/update.go handles the Update RPC by loading the persisted Crawl record, invoking DecodeState to reconstruct the progress object, and injecting it into the indexing pipeline. This allows the worker to continue processing from the exact URL or offset recorded in the last checkpoint.

Implementation Examples

Starting and Resuming Crawls via CLI

The command-line interface provides straightforward controls for managing crawl lifecycle states.

To initiate a new crawl:

// Launch a fresh crawl from the command line
cmd := exec.Command("hister", "crawl", "--url", "https://example.com", "--depth", "3")
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
cmd.Run()

To resume an interrupted session:

// Resume the most recent unfinished crawl
cmd := exec.Command("hister", "crawl", "--resume")
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
cmd.Run()

Programmatic State Management

Developers can interact directly with the state persistence layer through the server/model package.

import (
    "context"
    "github.com/asciimoo/hister/server/model"
)

// Retrieve a crawl and resume it from Go code
func resumeCrawl(ctx context.Context, db *gorm.DB, crawlID uint) error {
    var c model.Crawl
    if err := db.First(&c, crawlID).Error; err != nil {
        return err
    }
    // Unmarshal the saved state
    state, err := c.DecodeState()
    if err != nil {
        return err
    }
    // Pass the state to the crawler implementation
    return model.ResumeCrawler(ctx, db, &c, state)
}

Summary

  • State Persistence: Hister stores crawl progress in a JSON State field within the Crawl struct defined in server/model/crawl.go.
  • Database Backing: The crawls table schema in server/model/migration.go provides ACID-compliant storage via SQLite or PostgreSQL.
  • Automatic Recovery: The CLI --resume flag and server/indexer/update.go handler reconstruct crawler state from database snapshots.
  • Granular Checkpoints: Periodic calls to SaveState minimize reprocessing by recording last-processed URLs and pagination indices.

Frequently Asked Questions

What data format does Hister use to store crawl state?

Hister serializes crawl progress as JSON in the State column of the crawls table. This text-based format stores dynamic data such as last processed URLs, page indices, and crawler-specific pagination tokens.

Can Hister resume crawls after a system crash?

Yes. Because state updates are written to the database through ACID-compliant transactions, the persisted State remains consistent even if the process terminates unexpectedly. Upon restart, the --resume flag or server-side recovery logic reloads the last checkpoint from server/model/crawl.go.

How does Hister distinguish between running and completed crawls?

The Status field in the Crawl struct tracks lifecycle states including "running", "finished", and error conditions. When resuming, queries filter for rows with resumable statuses, ensuring completed jobs are not accidentally restarted.

Is it possible to resume specific crawl jobs by ID?

Yes. While the CLI --resume flag typically targets the most recent incomplete crawl, programmatic access via server/model allows direct retrieval by crawlID. The resumeCrawl function demonstrates loading a specific record and invoking DecodeState to restore its execution context.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →