# Reasonix Recovery Mechanisms: How the Framework Handles Stuck Agents and Token Limits

> Discover Reasonix recovery mechanisms for stuck agents and token limits. Learn how structured episodes, checkpointing, and retries ensure agent resilience. Explore the esengine DeepSeek Reasonix framework.

- Repository: [YHH/DeepSeek-Reasonix](https://github.com/esengine/DeepSeek-Reasonix)
- Tags: internals
- Published: 2026-08-11

---

**Reasonix uses a dedicated recovery subsystem with structured episodes, session checkpointing, automatic retry logic, and human-in-the-loop approvals to recover agents from token limits, panics, and stall conditions.**

When building autonomous AI agents, **recovery from failure** is as critical as the core execution logic. The Reasonix framework (from `esengine/DeepSeek-Reasonix`) isolates all resilience concerns into a purpose-built `recovery` package that handles everything from transient token-budget overflows to complete process crashes. This article examines the concrete mechanisms that bring stuck agents back to productive operation.

## The Recovery Episode Lifecycle

Every recovery operation in Reasonix follows a strict **four-phase state machine** defined in [`internal/recovery/types.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/types.go). Understanding these phases is essential to troubleshooting agent behavior.

### Phase 1: Idle

In `PhaseIdle`, the agent runs normally with no active failure detection. This is the default state for all sessions.

### Phase 2: Diagnosing

When the runtime detects a failure—whether a **panic in a tool handler**, **token-limit error from the LLM provider**, or an explicit "stuck" signal from the controller—it transitions to `PhaseDiagnosing`. The system:

- Records a `FailureEvent` with classification metadata
- Creates a `PendingProposal` describing possible remediation actions

This phase lives in [`internal/recovery/state.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/state.go), where the `RecordFailure()` method captures failure context.

### Phase 3: Awaiting Decision

The session pauses and surfaces a **recovery card** via `PhaseAwaitingDecision`. Each card contains:

- Full failure context (stack traces, tool outputs, token usage)
- A suggested remediation strategy
- A **review verdict** requiring user or automated confirmation

Card construction is centralized in `ToEventApproval` ([`internal/recovery/types.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/types.go), lines 80-95), which builds the `event.RecoveryApproval` payload dispatched to the UI layer.

### Phase 4: Resume or Revise

Based on the verdict, [`internal/recovery/state.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/state.go) executes one of three actions:

| Action | Handler Function | Behavior |
|--------|------------------|----------|
| **Continue** | `ActionContinue` | Retry the original plan unchanged |
| **Revise** | `ActionRevise` | Switch strategy, scope, or tools |
| **Abort** | `clearTaskRecoveryState` | Terminate episode, surface final error |

## Core Recovery Mechanisms

### Session-Level Checkpointing

Every session persists a **recovery checkpoint** to `<session>.recovery.json`. When a process crashes or token budget exhausts, the next startup loads this checkpoint via `SessionRecoveryState` in [`internal/store/session.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/store/session.go) and restores the exact episode state—including pending proposals and retry counters.

This enables **crash-only recovery**: no special shutdown logic required, just resume from the last consistent checkpoint.

### Recovery-Copy Branching

When a recovery episode creates a divergent plan, the original content is marked as a **recovery copy** so future merges can distinguish between original and recovery-generated artifacts. This logic spans two files:

- [`internal/sessioncatalog/types.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/sessioncatalog/types.go) — defines `RecoveryCopy` metadata
- [`internal/sessioncatalog/reconcile.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/sessioncatalog/reconcile.go) — implements three-way merge reconciliation

### Automatic Retry with Token-Budget Management

Transient failures (classified as `FailureClassTransient`) trigger **automatic retry** without human intervention. The task runtime in [`internal/agent/taskruntime.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/agent/taskruntime.go) maintains a `SafeRetryLeft` counter; each retry attempt decrements this value until exhausted or success occurs.

Token-limit errors from LLM providers are automatically classified as transient, allowing seamless budget replenishment and continuation.

### Human-in-the-Loop Recovery Cards

Not all recoveries proceed automatically. For uncertain failures or exhausted retry budgets, the framework constructs detailed **recovery approval requests** that surface in desktop, web, or CLI interfaces. Users choose to:

- Continue the current approach
- Confirm a revised plan with adjusted parameters
- Abort and investigate manually

### Telemetry and Observability

All recovery actions emitstructured metrics via [`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go) that enable operational visibility:

| Metric | When Incremented |
|--------|----------------|
| `recovery_failure` | Any failure event recorded |
| `recovery_rule_continue` | Automatic retry triggered |
| `recovery_human_prompt` | Recovery card surfaced to user |
| `recovery_success` | Episode resolved successfully |

These counters enable SLI/SLO tracking for agent reliability.

### Session-Lease Recovery Hooks

The serve layer ([`internal/serve/serve.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/serve/serve.go), lines 115-119) registers a callback via `ctrl.SetOnSessionRecovered` that re-instantiates the controller after any session lease recovery—whether from crash, token exhaustion, or deliberate pause. This ensures **stateful recovery handlers** persist across process restarts.

## Implementation Example: Detecting Token-Budget Overflow

The following pattern demonstrates how tool implementations surface token limits to the recovery subsystem:

```go
// In a tool implementation (e.g., an LLM call wrapper)
func callLLM(ctx context.Context, prompt string) (string, error) {
    resp, err := provider.Generate(ctx, prompt)
    if err != nil {
        // Detect token-budget overflow (provider returns a specific error type)
        if errors.Is(err, provider.ErrTokenLimitExceeded) {
            // Convert to a FailureEvent for the recovery subsystem
            fe := &recovery.FailureEvent{
                Class:      recovery.FailureClassTransient,
                Tool:       "llm",
                ErrSummary: err.Error(),
                SafeRetryLeft: 2, // allow a couple of automatic retries
            }
            // Record the failure – this will cause a Recovery Episode to start
            taskRuntime.RecordFailure(fe)
            return "", fmt.Errorf("recovery: token limit exceeded")
        }
        return "", err
    }
    return resp, nil
}

```

The `FailureClassTransient` classification enables automatic retry, while `SafeRetryLeft: 2` caps the retry attempts before escalation to human review.

## Implementation Example: Automatic Session Resumption

For infrastructure-level recovery, the serve layer implements checkpoint restoration:

```go
// In the serve layer – automatically resume after a paused recovery
func sessionLeaseRecoveryHandler(k *control.SessionLeaseKeeper) func(control.SessionRecoveryInfo) error {
    return func(info control.SessionRecoveryInfo) error {
        // Load the persisted recovery checkpoint
        ctrl, err := control.RestoreFromCheckpoint(info.Path)
        if err != nil {
            return err
        }
        // Re-attach the same recovery hook so further failures are handled
        ctrl.SetOnSessionRecovered(sessionLeaseRecoveryHandler(k))
        // Continue the session – the task runtime will pick up the saved episode
        return nil
    }
}

```

This handler chains indefinitely: every recovered session re-registers its own recovery handler, ensuring **recursive resilience**.

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`internal/recovery/types.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/types.go) | Phase definitions, `FailureEvent`, `PendingProposal`, `ToEventApproval` helper |
| [`internal/recovery/state.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/state.go) | Episode lifecycle management, action handlers |
| [`internal/store/session.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/store/session.go) | Checkpoint persistence and loading |
| [`internal/sessioncatalog/reconcile.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/sessioncatalog/reconcile.go) | Recovery-copy merge logic |
| [`internal/serve/serve.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/serve/serve.go) | `SetOnSessionRecovered` callback registration |
| [`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go) | Recovery metrics emission |

## Summary

- **Recovery episodes** follow a strict four-phase lifecycle (Idle → Diagnosing → Awaiting Decision → Resume/Revise) defined in [`internal/recovery/types.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/types.go).
- **Session checkpointing** to `*.recovery.json` enables crash-only recovery without graceful shutdown requirements.
- **Automatic retry** with `SafeRetryLeft` counters handles transient failures including token-limit overflows.
- **Recovery-copy branching** preserves original content when recovery generates divergent plans.
- **Human-in-the-loop approval** surfaces detailed recovery cards when automatic resolution is uncertain or exhausted.
- **Telemetry integration** provides operational visibility into recovery frequency and outcomes.

## Frequently Asked Questions

### How does Reasonix distinguish between recoverable and non-recoverable failures?

The framework uses **failure classification** via the `FailureClass` enum in [`internal/recovery/types.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/types.go). `FailureClassTransient` (token limits, temporary rate limits) triggers automatic retry. `FailureClassPermanent` (invalid API keys, irreparable tool errors) skips retry and surfaces immediately for human decision. The classification logic resides in tool wrappers and controller error handlers.

### Can recovery episodes nest—that is, can a recovery itself fail and enter another recovery?

Yes. The state machine in [`internal/recovery/state.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/recovery/state.go) supports **recursive recovery episodes**. If a recovery action (Continue or Revise) itself fails, the runtime records a new `FailureEvent` and enters a fresh Diagnosing phase. To prevent infinite recursion, the `SafeRetryLeft` counter decrements across nested episodes until human intervention is required.

### What happens to in-flight tool calls when a token limit triggers recovery?

Active tool calls receive **context cancellation** through the standard Go context tree. The task runtime in [`internal/agent/taskruntime.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/agent/taskruntime.go) cancels the operation's context before entering `PhaseDiagnosing`, ensuring tools release resources promptly. Partial results may be preserved in the session checkpoint depending on the tool's checkpointing implementation.

### How can operators monitor recovery behavior across a fleet of Reasonix agents?

The telemetry sink in [`internal/telemetry/sink.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/telemetry/sink.go) emits structured metrics including `recovery_failure`, `recovery_rule_continue`, `recovery_human_prompt`, and `recovery_success`. These integrate with standard observability stacks (Prometheus, OpenTelemetry). Additionally, session checkpoints contain recovery episode history for post-hoc analysis of agent reliability patterns.