Reasonix Recovery Mechanisms: How the Framework Handles Stuck Agents and Token Limits

Reasonix uses a dedicated recovery subsystem with structured episodes, session checkpointing, automatic retry logic, and human-in-the-loop approvals to recover agents from token limits, panics, and stall conditions.

When building autonomous AI agents, recovery from failure is as critical as the core execution logic. The Reasonix framework (from esengine/DeepSeek-Reasonix) isolates all resilience concerns into a purpose-built recovery package that handles everything from transient token-budget overflows to complete process crashes. This article examines the concrete mechanisms that bring stuck agents back to productive operation.

The Recovery Episode Lifecycle

Every recovery operation in Reasonix follows a strict four-phase state machine defined in internal/recovery/types.go. Understanding these phases is essential to troubleshooting agent behavior.

Phase 1: Idle

In PhaseIdle, the agent runs normally with no active failure detection. This is the default state for all sessions.

Phase 2: Diagnosing

When the runtime detects a failure—whether a panic in a tool handler, token-limit error from the LLM provider, or an explicit "stuck" signal from the controller—it transitions to PhaseDiagnosing. The system:

  • Records a FailureEvent with classification metadata
  • Creates a PendingProposal describing possible remediation actions

This phase lives in internal/recovery/state.go, where the RecordFailure() method captures failure context.

Phase 3: Awaiting Decision

The session pauses and surfaces a recovery card via PhaseAwaitingDecision. Each card contains:

  • Full failure context (stack traces, tool outputs, token usage)
  • A suggested remediation strategy
  • A review verdict requiring user or automated confirmation

Card construction is centralized in ToEventApproval (internal/recovery/types.go, lines 80-95), which builds the event.RecoveryApproval payload dispatched to the UI layer.

Phase 4: Resume or Revise

Based on the verdict, internal/recovery/state.go executes one of three actions:

Action Handler Function Behavior
Continue ActionContinue Retry the original plan unchanged
Revise ActionRevise Switch strategy, scope, or tools
Abort clearTaskRecoveryState Terminate episode, surface final error

Core Recovery Mechanisms

Session-Level Checkpointing

Every session persists a recovery checkpoint to <session>.recovery.json. When a process crashes or token budget exhausts, the next startup loads this checkpoint via SessionRecoveryState in internal/store/session.go and restores the exact episode state—including pending proposals and retry counters.

This enables crash-only recovery: no special shutdown logic required, just resume from the last consistent checkpoint.

Recovery-Copy Branching

When a recovery episode creates a divergent plan, the original content is marked as a recovery copy so future merges can distinguish between original and recovery-generated artifacts. This logic spans two files:

Automatic Retry with Token-Budget Management

Transient failures (classified as FailureClassTransient) trigger automatic retry without human intervention. The task runtime in internal/agent/taskruntime.go maintains a SafeRetryLeft counter; each retry attempt decrements this value until exhausted or success occurs.

Token-limit errors from LLM providers are automatically classified as transient, allowing seamless budget replenishment and continuation.

Human-in-the-Loop Recovery Cards

Not all recoveries proceed automatically. For uncertain failures or exhausted retry budgets, the framework constructs detailed recovery approval requests that surface in desktop, web, or CLI interfaces. Users choose to:

  • Continue the current approach
  • Confirm a revised plan with adjusted parameters
  • Abort and investigate manually

Telemetry and Observability

All recovery actions emitstructured metrics via internal/telemetry/sink.go that enable operational visibility:

Metric When Incremented
recovery_failure Any failure event recorded
recovery_rule_continue Automatic retry triggered
recovery_human_prompt Recovery card surfaced to user
recovery_success Episode resolved successfully

These counters enable SLI/SLO tracking for agent reliability.

Session-Lease Recovery Hooks

The serve layer (internal/serve/serve.go, lines 115-119) registers a callback via ctrl.SetOnSessionRecovered that re-instantiates the controller after any session lease recovery—whether from crash, token exhaustion, or deliberate pause. This ensures stateful recovery handlers persist across process restarts.

Implementation Example: Detecting Token-Budget Overflow

The following pattern demonstrates how tool implementations surface token limits to the recovery subsystem:

// In a tool implementation (e.g., an LLM call wrapper)
func callLLM(ctx context.Context, prompt string) (string, error) {
    resp, err := provider.Generate(ctx, prompt)
    if err != nil {
        // Detect token-budget overflow (provider returns a specific error type)
        if errors.Is(err, provider.ErrTokenLimitExceeded) {
            // Convert to a FailureEvent for the recovery subsystem
            fe := &recovery.FailureEvent{
                Class:      recovery.FailureClassTransient,
                Tool:       "llm",
                ErrSummary: err.Error(),
                SafeRetryLeft: 2, // allow a couple of automatic retries
            }
            // Record the failure – this will cause a Recovery Episode to start
            taskRuntime.RecordFailure(fe)
            return "", fmt.Errorf("recovery: token limit exceeded")
        }
        return "", err
    }
    return resp, nil
}

The FailureClassTransient classification enables automatic retry, while SafeRetryLeft: 2 caps the retry attempts before escalation to human review.

Implementation Example: Automatic Session Resumption

For infrastructure-level recovery, the serve layer implements checkpoint restoration:

// In the serve layer – automatically resume after a paused recovery
func sessionLeaseRecoveryHandler(k *control.SessionLeaseKeeper) func(control.SessionRecoveryInfo) error {
    return func(info control.SessionRecoveryInfo) error {
        // Load the persisted recovery checkpoint
        ctrl, err := control.RestoreFromCheckpoint(info.Path)
        if err != nil {
            return err
        }
        // Re-attach the same recovery hook so further failures are handled
        ctrl.SetOnSessionRecovered(sessionLeaseRecoveryHandler(k))
        // Continue the session – the task runtime will pick up the saved episode
        return nil
    }
}

This handler chains indefinitely: every recovered session re-registers its own recovery handler, ensuring recursive resilience.

Key Source Files Reference

File Purpose
internal/recovery/types.go Phase definitions, FailureEvent, PendingProposal, ToEventApproval helper
internal/recovery/state.go Episode lifecycle management, action handlers
internal/store/session.go Checkpoint persistence and loading
internal/sessioncatalog/reconcile.go Recovery-copy merge logic
internal/serve/serve.go SetOnSessionRecovered callback registration
internal/telemetry/sink.go Recovery metrics emission

Summary

  • Recovery episodes follow a strict four-phase lifecycle (Idle → Diagnosing → Awaiting Decision → Resume/Revise) defined in internal/recovery/types.go.
  • Session checkpointing to *.recovery.json enables crash-only recovery without graceful shutdown requirements.
  • Automatic retry with SafeRetryLeft counters handles transient failures including token-limit overflows.
  • Recovery-copy branching preserves original content when recovery generates divergent plans.
  • Human-in-the-loop approval surfaces detailed recovery cards when automatic resolution is uncertain or exhausted.
  • Telemetry integration provides operational visibility into recovery frequency and outcomes.

Frequently Asked Questions

How does Reasonix distinguish between recoverable and non-recoverable failures?

The framework uses failure classification via the FailureClass enum in internal/recovery/types.go. FailureClassTransient (token limits, temporary rate limits) triggers automatic retry. FailureClassPermanent (invalid API keys, irreparable tool errors) skips retry and surfaces immediately for human decision. The classification logic resides in tool wrappers and controller error handlers.

Can recovery episodes nest—that is, can a recovery itself fail and enter another recovery?

Yes. The state machine in internal/recovery/state.go supports recursive recovery episodes. If a recovery action (Continue or Revise) itself fails, the runtime records a new FailureEvent and enters a fresh Diagnosing phase. To prevent infinite recursion, the SafeRetryLeft counter decrements across nested episodes until human intervention is required.

What happens to in-flight tool calls when a token limit triggers recovery?

Active tool calls receive context cancellation through the standard Go context tree. The task runtime in internal/agent/taskruntime.go cancels the operation's context before entering PhaseDiagnosing, ensuring tools release resources promptly. Partial results may be preserved in the session checkpoint depending on the tool's checkpointing implementation.

How can operators monitor recovery behavior across a fleet of Reasonix agents?

The telemetry sink in internal/telemetry/sink.go emits structured metrics including recovery_failure, recovery_rule_continue, recovery_human_prompt, and recovery_success. These integrate with standard observability stacks (Prometheus, OpenTelemetry). Additionally, session checkpoints contain recovery episode history for post-hoc analysis of agent reliability patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →