# How to Use the Reasonix Recovery System When It Crashes or Gets Stuck

> Encountered a Reasonix recovery system crash or hang? Learn how to automatically recover your session, persist state, and choose to continue, revise, or discard your work.

- Repository: [YHH/DeepSeek-Reasonix](https://github.com/esengine/DeepSeek-Reasonix)
- Tags: how-to-guide
- Published: 2026-08-08

---

**The Reasonix recovery system automatically creates a Recovery Episode when the engine crashes or hangs, persisting the current state to disk and presenting a prompt to continue, revise, or discard the session.**

The `esengine/DeepSeek-Reasonix` repository implements a robust **Reasonix recovery system** designed to protect long-running reasoning sessions from unexpected failures. When the core agent detects a crash, missing reasoning error, or indefinite hang, the system captures the current execution state and pauses gracefully, allowing users to recover progress rather than starting from scratch.

## How the Reasonix Recovery System Detects Failures

The recovery workflow begins with the **Recovery Gate** in [`internal/recovery/gate.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/gate.go). The core agent monitors turn outcomes through the `ObserveResult` method, checking for `TurnOutcomeRecoveryPaused` or missing-reasoning errors. When `EpisodeStopped` flags a failure, the system immediately initiates a Recovery Episode.

### Persisting Checkpoint State

Upon detection, the system writes a checkpoint file via `SessionRecoveryState` in [`internal/store/session.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/store/session.go). This file stores the Recovery Episode ID and any partial task results under the session directory, ensuring that even if the process terminates abruptly, the recovery state remains intact.

### Presenting the Recovery Prompt

Once persisted, the web UI or CLI displays a recovery prompt tracked by the `recovery_human_prompt` metric in [`workers/crash-report/src/stats.ts`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/workers/crash-report/src/stats.ts). The endpoint handler `sessionLeaseRecoveryHandler` in [`internal/serve/serve.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/serve/serve.go) routes user decisions to the appropriate recovery action.

## Recovery Actions: Continue, Revise, or Discard

The **Recovery Gate** exposes three distinct actions defined in [`internal/recovery/types.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/types.go) and implemented in [`gate.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/gate.go):

- **Continue**: Resumes execution of the original episode using the persisted checkpoint.
- **Revise**: Rejects the pending action and starts a fresh Recovery Episode via `Gate.Revise`.
- **Discard**: Clears the session state entirely, allowing the user to start a new task without historical baggage.

Each action is logged through the telemetry sink in [`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go) via `RecordProtocolRecovery`, creating an audit trail for failure analysis.

## Manual Recovery Triggers

While the system auto-detects failures, you can manually invoke the **Reasonix recovery system** through both CLI flags and web interface controls.

### CLI Recovery Flags

Add the `--recover` flag to force the controller to load existing checkpoints:

```bash

# Resume the latest checkpoint

reasonix run --recover ~/.reasonix/sessions/2024-07-31-recovery-abc123

# Discard and start fresh

reasonix run --reset-session

```

### Web UI Recovery Flow

When the UI displays a "Recovery paused" banner, clicking **Continue**, **Revise**, or **Discard** triggers a POST to `/session/recover`, which internally invokes `sessionLeaseRecoveryHandler` to process your selection.

## Auditing and Cleanup

Every recovery event is recorded via `RecordProtocolRecovery` in [`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go), enabling developers to analyze patterns like `ProtocolRecoveryMissingReasoningRetryRecovered`. Once an episode completes successfully or aborts, `taskRuntime.clearTaskRecoveryState` in [`internal/recovery/state.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/state.go) removes the temporary checkpoint files to prevent disk bloat.

## Practical Code Examples

The following examples demonstrate programmatic interaction with the recovery system:

```bash

# Run normally

reasonix run my_task.yaml

# Check for recovery checkpoints after a crash

ls ~/.reasonix/sessions/*-recovery-*

# Resume specific checkpoint

reasonix run --recover ~/.reasonix/sessions/2024-07-31-recovery-abc123

```

For web UI integrations, resume recovery programmatically:

```go
func resumeRecovery(sessPath string) error {
    // Load persisted state
    recState := recovery.Load(sessPath) // internal/recovery/persist.go
    
    // Initialize controller with recovery handler
    ctrl := control.NewController()
    ctrl.SetOnSessionRecovered(sessionLeaseRecoveryHandler(leases))
    
    // Continue execution
    return ctrl.ContinueRecovery(recState)
}

```

## Summary

- The **Reasonix recovery system** in `esengine/DeepSeek-Reasonix` automatically detects crashes via the Recovery Gate ([`internal/recovery/gate.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/gate.go)).
- Checkpoints are persisted to disk using `SessionRecoveryState` ([`internal/store/session.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/store/session.go)) before presenting user prompts.
- Users can **Continue**, **Revise**, or **Discard** stuck sessions through CLI flags (`--recover`, `--reset-session`) or the web UI (`/session/recover`).
- All recovery actions are audited via `RecordProtocolRecovery` ([`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go)) and cleaned up via `clearTaskRecoveryState` ([`internal/recovery/state.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/state.go)).

## Frequently Asked Questions

### What happens when Reasonix detects a crash?

The system invokes `ObserveResult` in [`internal/recovery/gate.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/gate.go) to identify `TurnOutcomeRecoveryPaused` or missing-reasoning errors. It then writes a checkpoint via `SessionRecoveryState` and displays a recovery prompt tracked by `recovery_human_prompt` metrics.

### How do I manually trigger recovery from the command line?

Use the `--recover` flag followed by the session path to resume a specific checkpoint, or `--reset-session` to discard recovery state and start fresh. The CLI forces the controller to load existing checkpoints and present the recovery menu.

### What is the difference between Revise and Discard actions?

**Revise** (implemented via `Gate.Revise` in [`gate.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/gate.go)) rejects the pending corrupted action and spawns a new Recovery Episode while preserving context, whereas **Discard** calls `clearTaskRecoveryState` to wipe all recovery data and returns the system to a clean slate.

### Where are recovery checkpoints stored?

Checkpoints are written to the session directory under paths like `~/.reasonix/sessions/*-recovery-*`, defined in [`internal/store/session.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/store/session.go). The `SessionRecoveryState` structure maintains the Recovery Episode ID and partial results until explicitly cleared.