How Automatic Error Recovery Works in DeepSeek-Reasonix: Auto-Mode Resilience Explained

DeepSeek-Reasonix implements an Auto-mode recovery system that persists session state to side-car files, detects crashes via transaction locks, and automatically resumes execution from the last safe checkpoint without user intervention.

DeepSeek-Reasonix is engineered to withstand unexpected interruptions during long-running reasoning workflows. The repository implements a robust automatic error recovery mechanism that combines atomic checkpointing, crash-resistant inbox queues, and catalog-based lineage tracking. This architecture ensures that even if a process terminates abruptly, the engine can detect the failure, preserve all progress, and seamlessly continue execution when resources become available again.

Recovery Checkpoint Persistence via Side-Car Files

The foundation of automatic error recovery in DeepSeek-Reasonix lies in atomic checkpoint persistence. Each session maintains a side-car file named <session>.recovery.json, derived by appending .recovery.json to the normal session file path as implemented in internal/store/session.go (lines 31-38).

This checkpoint contains the current run ID, pending items, and a recovery ledger that logs partial work. The side-car approach ensures that state snapshots remain collocated with session data while maintaining independence from the main workflow file, enabling atomic updates without corrupting the primary session state.

The following Go implementation demonstrates how the system persists recovery state:

func saveRecovery(sessionPath string, state RecoveryState) error {
    // Derive side‑car name
    recoveryPath := sessionPath + ".recovery.json"
    // Marshal the state and write atomically
    data, _ := json.Marshal(state)
    return os.WriteFile(recoveryPath, data, 0o600)
}

Cross-Process Crash Recovery with Session Inbox

To handle process-level failures, DeepSeek-Reasonix utilizes a sessioninbox layer that implements transaction locks surviving process crashes. According to internal/sessioninbox/store.go (lines 43-45 and 212-215), when a crash is detected, the inbox inspects the recovery ledger and re-queues any orphaned items for the next run.

This mechanism guarantees exactly-once semantics for work items even when the process terminates unexpectedly. The inbox maintains durability through filesystem-level locks that persist beyond process boundaries, allowing subsequent runs to identify incomplete transactions and reclaim orphaned work without manual intervention.

Recovery Groups and Branch Lineage Tracking

Beyond simple state restoration, DeepSeek-Reasonix tracks recovery semantics through catalog metadata to maintain provenance. The system maintains recovery groups representing proven recovery lineages in internal/sessioncatalog/recovery_groups.go (lines 64-70 and 122-130).

Sessions marked as recovery branches carry a recovery_group_id along with metadata fields including recovery_state, recovery_reason, and recovery_digest. These fields enable the system to distinguish between authoritative branches and temporary recovery copies, preventing contamination of proven results while allowing speculative continuation of interrupted workflows.

Telemetry and Observability for Recovery Events

Every recovery event is instrumented through the telemetry sink defined in internal/telemetry/sink.go (lines 75-87 and 131-132). The system emits specific metrics including recovery_failure, recovery_rule_continue, and recovery_human_prompt to provide visibility into automatic error recovery patterns.

This telemetry allows operators to analyze recovery frequency, identify systemic failure modes, and optimize recovery rules. The metrics distinguish between automated continuations and those requiring human intervention, providing quantitative insight into the self-healing capabilities of the system.

The telemetry integration is implemented as follows:

func (r *Reporter) RecordRecovery(m recovery.Metrics) {
    // Increment counters for each metric type
    addMetric(counts, "recovery_failure", "count", m.FailureEvents)
    // …
}

Automatic Continuation Without User Intervention

When DeepSeek-Reasonix initializes a session, the loader automatically checks for existing recovery side-cars via the loadRecovery function. If a valid checkpoint exists, the system restores the saved state, replays recorded events, and resumes the workflow from the last recorded ProtocolRecovery point.

This auto-mode operates without explicit user commands, enabling unattended operation for batch processing and long-running reasoning tasks. The loader validates recovery digests to ensure state integrity before continuation, preventing execution from corrupted partial writes.

The restoration logic follows this pattern:

func loadRecovery(sessionPath string) (*RecoveryState, error) {
    recoveryPath := sessionPath + ".recovery.json"
    b, err := os.ReadFile(recoveryPath)
    if err != nil {
        return nil, err
    }
    var state RecoveryState
    if err := json.Unmarshal(b, &state); err != nil {
        return nil, err
    }
    return &state, nil
}

Summary

  • Side-car persistence: DeepSeek-Reasonix stores recovery state in <session>.recovery.json files alongside main session data, implemented in internal/store/session.go.
  • Crash-resistant inbox: The sessioninbox layer uses transaction locks that survive process crashes to re-queue orphaned work, ensuring no data loss during unexpected termination.
  • Lineage tracking: Recovery groups and branch metadata in internal/sessioncatalog/recovery_groups.go maintain provenance and distinguish between authoritative and temporary recovery states.
  • Telemetry integration: Recovery events emit metrics like recovery_failure and recovery_rule_continue through internal/telemetry/sink.go for operational visibility.
  • Seamless resumption: The auto-mode loader automatically detects and restores from checkpoints without user intervention, replaying events from the last ProtocolRecovery point.

Frequently Asked Questions

How does DeepSeek-Reasonix detect process crashes?

DeepSeek-Reasonix detects process crashes through filesystem-based transaction locks in the sessioninbox layer that persist beyond process boundaries. When a new run starts, the inbox inspects the recovery ledger for incomplete transactions and orphaned items, as implemented in internal/sessioninbox/store.go (lines 43-45). This allows the system to identify work that was in-flight when the previous process terminated.

What data is stored in the recovery checkpoint?

The recovery checkpoint stored in <session>.recovery.json contains the current run ID, pending work items, and a recovery ledger logging partial work. According to internal/store/session.go (lines 31-38), this side-car file captures sufficient state to reconstruct the session context and resume from the last safe point without losing progress on completed or in-progress tasks.

Can automatic recovery corrupt existing session data?

No, the automatic error recovery system prevents corruption through atomic side-car files and digest validation. The recovery metadata includes recovery_digest fields that verify state integrity before resumption, as defined in internal/sessioncatalog/recovery_groups.go (lines 122-130). Additionally, recovery branches are tracked separately from authoritative sessions to isolate speculative recovery attempts from proven results.

How can I monitor recovery events in production?

Recovery events are automatically reported through the telemetry sink in internal/telemetry/sink.go (lines 75-87), which emits metrics including recovery_failure, recovery_rule_continue, and recovery_human_prompt. These metrics can be consumed by monitoring systems to track recovery patterns, identify recurring failure modes, and alert when manual intervention is required versus when the system successfully auto-recovers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →