# How Hister Manages Crawl Job State and Resumability

> Learn how Hister manages crawl job state and resumability. Discover how Hister's state persistence with SQLite or PostgreSQL enables automatic crawl resumption from failure points.

- Repository: [Adam Tauber/hister](https://github.com/asciimoo/hister)
- Tags: internals
- Published: 2026-09-01

---

**Hister persists crawl progress in a SQLite or PostgreSQL database using a JSON-encoded State field within the Crawl struct, enabling automatic resumption of interrupted crawls from the exact point of failure.**

The `asciimoo/hister` search engine implements robust crawl job state management by treating each crawl as a first-class database entity. By serializing internal progress markers to a relational store via [`server/model/crawl.go`](https://github.com/asciimoo/hister/blob/main/server/model/crawl.go), Hister ensures that long-running indexing operations survive process restarts, crashes, and graceful shutdowns without data loss.

## Core Architecture Components

### The Crawl Struct and State Persistence

In [`server/model/crawl.go`](https://github.com/asciimoo/hister/blob/main/server/model/crawl.go), the **Crawl struct** serves as the central abstraction for job state tracking. It contains metadata fields including `ID`, `UserID`, `Status`, `StartedAt`, `FinishedAt`, and `Error`, alongside a critical `State` field that stores JSON-encoded progress data.

The `State` field tracks dynamic crawl progress such as the last processed URL, pagination tokens, or batch offsets. During execution, the crawler periodically invokes `SaveState` to flush the current progress to the database, creating a durable checkpoint that survives unexpected termination.

### Database Schema and ACID Guarantees

The underlying schema is defined in [`server/model/migration.go`](https://github.com/asciimoo/hister/blob/main/server/model/migration.go), which creates the `crawls` table with a `TEXT` column for the JSON state snapshot. Because Hister uses standard relational backends—either SQLite or PostgreSQL—the persistence layer guarantees **ACID consistency**, ensuring that state updates are atomic and durable even under power loss or container restarts.

## How Resumability Works

### State Serialization Flow

When a crawl launches, the system inserts a new row with `Status` set to `"running"` and an empty `State` buffer. After processing each document batch, the crawler marshals its internal progress map into JSON and updates the `State` column via GORM. This incremental persistence strategy minimizes work duplication while maintaining low database write overhead.

### Recovery and Resume Logic

The resumability mechanism activates when the `hister crawl` command is invoked with the `--resume` flag, implemented in [`cmd/crawl.go`](https://github.com/asciimoo/hister/blob/main/cmd/crawl.go). The command queries the `crawls` table for incomplete jobs marked as `running` or `paused`, retrieves the stored JSON state, and deserializes it back into the crawler's memory.

On the server side, [`server/indexer/update.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/update.go) handles the **Update RPC** by loading the persisted `Crawl` record, invoking `DecodeState` to reconstruct the progress object, and injecting it into the indexing pipeline. This allows the worker to continue processing from the exact URL or offset recorded in the last checkpoint.

## Implementation Examples

### Starting and Resuming Crawls via CLI

The command-line interface provides straightforward controls for managing crawl lifecycle states.

To initiate a new crawl:

```go
// Launch a fresh crawl from the command line
cmd := exec.Command("hister", "crawl", "--url", "https://example.com", "--depth", "3")
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
cmd.Run()

```

To resume an interrupted session:

```go
// Resume the most recent unfinished crawl
cmd := exec.Command("hister", "crawl", "--resume")
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
cmd.Run()

```

### Programmatic State Management

Developers can interact directly with the state persistence layer through the `server/model` package.

```go
import (
    "context"
    "github.com/asciimoo/hister/server/model"
)

// Retrieve a crawl and resume it from Go code
func resumeCrawl(ctx context.Context, db *gorm.DB, crawlID uint) error {
    var c model.Crawl
    if err := db.First(&c, crawlID).Error; err != nil {
        return err
    }
    // Unmarshal the saved state
    state, err := c.DecodeState()
    if err != nil {
        return err
    }
    // Pass the state to the crawler implementation
    return model.ResumeCrawler(ctx, db, &c, state)
}

```

## Summary

- **State Persistence**: Hister stores crawl progress in a JSON `State` field within the `Crawl` struct defined in [`server/model/crawl.go`](https://github.com/asciimoo/hister/blob/main/server/model/crawl.go).
- **Database Backing**: The `crawls` table schema in [`server/model/migration.go`](https://github.com/asciimoo/hister/blob/main/server/model/migration.go) provides ACID-compliant storage via SQLite or PostgreSQL.
- **Automatic Recovery**: The CLI `--resume` flag and [`server/indexer/update.go`](https://github.com/asciimoo/hister/blob/main/server/indexer/update.go) handler reconstruct crawler state from database snapshots.
- **Granular Checkpoints**: Periodic calls to `SaveState` minimize reprocessing by recording last-processed URLs and pagination indices.

## Frequently Asked Questions

### What data format does Hister use to store crawl state?

Hister serializes crawl progress as JSON in the `State` column of the `crawls` table. This text-based format stores dynamic data such as last processed URLs, page indices, and crawler-specific pagination tokens.

### Can Hister resume crawls after a system crash?

Yes. Because state updates are written to the database through ACID-compliant transactions, the persisted `State` remains consistent even if the process terminates unexpectedly. Upon restart, the `--resume` flag or server-side recovery logic reloads the last checkpoint from [`server/model/crawl.go`](https://github.com/asciimoo/hister/blob/main/server/model/crawl.go).

### How does Hister distinguish between running and completed crawls?

The `Status` field in the `Crawl` struct tracks lifecycle states including `"running"`, `"finished"`, and error conditions. When resuming, queries filter for rows with resumable statuses, ensuring completed jobs are not accidentally restarted.

### Is it possible to resume specific crawl jobs by ID?

Yes. While the CLI `--resume` flag typically targets the most recent incomplete crawl, programmatic access via `server/model` allows direct retrieval by `crawlID`. The `resumeCrawl` function demonstrates loading a specific record and invoking `DecodeState` to restore its execution context.