# How Heretic Resumes Interrupted Optimization Runs: A Deep Dive into Checkpoint Recovery

> Learn how Heretic resumes interrupted optimization runs. Discover its journal-storage backend, checkpoint recovery, and seamless study state reloading for continuous optimization.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: deep-dive
- Published: 2026-02-19

---

**Heretic uses Optuna's journal-storage backend to persist the full study state to a JSON-Lines checkpoint file, enabling seamless resumption by reloading saved Pydantic settings and continuing optimization from the last completed trial.**

Heretic, an open-source neural architecture search tool by `p-e-w/heretic`, builds its optimization loop on top of Optuna. When running expensive hyperparameter searches that can take hours or days, the ability to resume interrupted optimization runs is critical. Heretic solves this through a robust checkpointing system that preserves every trial's state on disk.

## Understanding Heretic's Checkpoint Architecture

Heretic's resume capability rests on a persistent storage layer that records every study mutation atomically. This design ensures that even if the process crashes during a trial, the study remains consistent and recoverable.

### Journal Storage Backend

At the core of the checkpoint system is Optuna's `JournalFileBackend`. When Heretic initializes a study, it creates a storage backend that appends every operation to a JSON-Lines file:

```python
backend = JournalFileBackend(study_checkpoint_file, lock_obj=lock_obj)
storage = JournalStorage(backend)                      # src/heretic/main.py#L245-L248

```

The `JournalStorage` wrapper translates high-level Optuna operations into append-only log entries. This append-only design minimizes write amplification and ensures that partial writes cannot corrupt the entire study history.

### Checkpoint File Location

By default, Heretic stores checkpoint files in the directory defined by `settings.study_checkpoint_dir`. Each study receives a unique JSON-Lines file that contains the complete trial history, user attributes, and study metadata. This file persists across process restarts, enabling the resume functionality.

## The Resume Detection and Recovery Process

When you restart Heretic with the same configuration, the application detects the existing checkpoint and initiates a recovery workflow. This process involves user interaction, configuration restoration, and trial accounting.

### Detecting Existing Studies

Upon startup, Heretic queries the `JournalStorage` to check for existing studies. If a checkpoint file is present, the application presents three options (implemented around lines 750-795 in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py)):

- Show previous results (if the run already finished)
- **Continue the previous run** (if it was interrupted)
- Restart from scratch (delete the checkpoint)

This interactive prompt ensures that users consciously choose whether to resume interrupted optimization runs or begin fresh.

### Loading Saved Configuration

When you select "Continue the previous run," Heretic retrieves the original configuration from the study's user attributes. The application stores a JSON-serialized version of the initial Pydantic `Settings` object within the Optuna study.

The configuration system in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) uses a custom settings source called `init_settings` that overrides all other configuration sources during a resume operation:

```python
return (init_settings,  # Used during resume – should override *all* other sources.

        CliSettingsSource(...),
        EnvSettingsSource(...), ...)

# src/heretic/config.py#L326-L337

```

This design ensures that the original hyperparameters and search space definitions take precedence over environment variables or command-line arguments that might differ from the initial run.

### Calculating Remaining Trials

Before resuming optimization, Heretic counts the number of already completed trials using `count_completed_trials()`. The application then calculates the remaining trials by subtracting this count from `settings.n_trials`:

```python
study.optimize(objective_wrapper,
               n_trials=settings.n_trials - count_completed_trials())

# src/heretic/main.py#L594-L597

```

This calculation ensures that the total number of trials across the original and resumed runs equals the user's original request. Heretic prints a "Resuming existing study." message (lines 888-894 in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py)) to confirm the resume operation.

## Practical Implementation: Code Examples

To resume an interrupted optimization run, simply execute the same Heretic command you used initially:

```bash

# First run – interrupted by system crash or Ctrl-C

$ heretic --model meta-llama/Meta-Llama-3-8B --n-trials 100

# Restart the same command

$ heretic --model meta-llama/Meta-Llama-3-8B --n-trials 100

```

Upon restart, Heretic detects the checkpoint and displays:

```

Resuming existing study.
Found 47 completed trials. Continuing with 53 remaining trials.

```

If you prefer to discard the interrupted run and start fresh:

```bash

# Remove the checkpoint directory

$ rm -rf ~/.cache/heretic/studies

# Or select "Ignore and start from scratch" in the interactive prompt

$ heretic --model meta-llama/Meta-Llama-3-8B --n-trials 100

```

## Summary

Heretic implements robust resumption of interrupted optimization runs through:

- **Journal-based checkpointing**: Uses Optuna's `JournalFileBackend` to append study state to a JSON-Lines file at `src/heretic/main.py#L245-L248`
- **Automatic detection**: Identifies existing studies on startup and offers interactive choices around `src/heretic/main.py#L750-L795`
- **Configuration preservation**: Reloads original Pydantic settings via the `init_settings` source that overrides all other inputs at `src/heretic/config.py#L326-L337`
- **Trial accounting**: Calculates remaining trials with `count_completed_trials()` and resumes optimization at `src/heretic/main.py#L594-L597`

## Frequently Asked Questions

### How does Heretic store optimization progress between runs?

Heretic stores optimization progress in a JSON-Lines checkpoint file managed by Optuna's `JournalFileBackend`. This file contains every trial's parameters, values, and study metadata. The backend appends each operation atomically to the file located in `settings.study_checkpoint_dir`, ensuring that even if the process crashes mid-trial, the study remains consistent and recoverable.

### Can I change hyperparameters when resuming an interrupted run?

No, you cannot change hyperparameters when resuming an interrupted optimization run. Heretic deliberately prevents this by loading the original `Settings` object from the study's user attributes and using the `init_settings` source to override any new command-line arguments or environment variables. This design ensures experimental consistency—resuming continues the exact same search space and configuration that was interrupted.

### What happens if I reach the trial limit but want to add more trials later?

If you reach the trial limit but want to add more trials later, Heretic allows you to extend the study. You would update `settings.n_trials` to the new desired total and restart the command. Heretic detects the existing study, counts the already completed trials, and calls `study.optimize` with only the remaining number of trials needed to reach the new limit. The study's user attribute `finished` is updated accordingly to reflect the extended state.