# AIOX Recovery System: How Automatic Failure Recovery Works in SynkraAI/aiox-core

> Discover the AIOX recovery system in SynkraAI/aiox-core. Learn how this self-healing layer automatically detects, classifies, and remediates failures like retries, rollbacks, and escalations without human intervention.

- Repository: [SynkraAI/aiox-core](https://github.com/synkraai/aiox-core)
- Tags: how-to-guide
- Published: 2026-03-15

---

**The AIOX Recovery System is a self-healing orchestration layer that detects epic failures, classifies error types, and executes automated remediation strategies—including retries, rollbacks, and escalation—without human intervention.**

The **AIOX recovery system** provides the resilience infrastructure for the AIOX Autonomous Development Engine (ADE), ensuring that pipeline epics (Spec, Execution, QA) recover automatically from transient errors, state corruption, or dependency conflicts. Implemented in the `SynkraAI/aiox-core` repository, this **automatic failure recovery** mechanism operates through a centralized `RecoveryHandler` class that coordinates with detection and rollback subsystems to maintain continuous autonomous development workflows.

## Core Architecture of the AIOX Recovery System

The recovery architecture consists of a central handler supported by three lazy-loaded infrastructure modules and integrated by the master orchestrator.

### RecoveryHandler (Central Coordinator)

The `RecoveryHandler` class in [`.aiox-core/core/orchestration/recovery-handler.js`](https://github.com/SynkraAI/aiox-core/blob/main/.aiox-core/core/orchestration/recovery-handler.js) serves as the brain of the operation. It maintains an internal `attempts` map to track failure history, evaluates whether the pipeline is stuck using circular detection logic, and executes remediation strategies. When an epic fails, the handler records the timestamp, error message, and approach, then consults its strategy selection algorithm to determine the next action.

### Infrastructure Components

Three specialized modules provide the underlying capabilities:

- **StuckDetector** (`../../infrastructure/scripts/stuck-detector`): Analyzes recent attempt patterns to identify circular failures or repeated unsuccessful approaches. It determines whether the system is retrying the same failing path indefinitely.
- **RollbackManager** (`../../infrastructure/scripts/rollback-manager`): Creates project checkpoints and executes state restoration when the `ROLLBACK_AND_RETRY` strategy is selected. It ensures the system can return to a known-good state before attempting a new approach.
- **RecoveryTracker** (`../../infrastructure/scripts/recovery-tracker`): Persists attempt histories to [`recovery/attempts.json`](https://github.com/SynkraAI/aiox-core/blob/main/recovery/attempts.json) for audit trails and future debugging, ensuring failure patterns are recorded across sessions.

### MasterOrchestrator Integration

The `MasterOrchestrator` in [`.aiox-core/core/orchestration/master-orchestrator.js`](https://github.com/SynkraAI/aiox-core/blob/main/.aiox-core/core/orchestration/master-orchestrator.js) instantiates the `RecoveryHandler` and delegates all failure handling to it. Lines 142–149 of the orchestrator expose configuration options (`maxRetries`, `autoEscalate`, `circularDetection`) that tune the recovery behavior according to project resiliency requirements.

## The Automatic Failure Recovery Workflow

When an epic throws an error, the **automatic failure recovery** process executes through seven distinct steps:

1. **Error Capture** – The orchestrator calls `RecoveryHandler.handleEpicFailure(epicNum, error, context)` with the failed epic number, error object, and execution context.

2. **Attempt Logging** – The handler records the failure details in its internal map and persists them via `RecoveryTracker`, creating a historical record of the error classification and attempted approaches.

3. **Stuck Detection** – The handler invokes `StuckDetector` to analyze whether the failure pattern exhibits circular behavior (same approach repeated) or exceeds consecutive failure thresholds.

4. **Strategy Selection** – Using `_classifyError` to categorize the error type and `_selectRecoveryStrategy` to evaluate attempt counts and stuck-detector output, the system chooses from five `RecoveryStrategy` options:
   - **RETRY_SAME_APPROACH** – Retries the epic immediately with identical parameters
   - **ROLLBACK_AND_RETRY** – Restores a checkpoint via `RollbackManager` then retries with a modified approach
   - **SKIP_PHASE** – Marks non-critical epics as skipped to maintain pipeline momentum
   - **ESCALATE_TO_HUMAN** – Generates an escalation report and flags the orchestrator as `BLOCKED` for manual intervention
   - **TRIGGER_RECOVERY_WORKFLOW** – Launches Epic 5, a dedicated recovery epic for complex dependency resolution

5. **Strategy Execution** – The `_executeRecoveryStrategy` method carries out the selected action, invoking infrastructure services as needed.

6. **Event Emission** – Each recovery attempt emits a `recoveryAttempt` event containing the epic number, attempt count, strategy name, and result, enabling real-time dashboard updates and agent monitoring.

7. **Result Propagation** – The handler returns a structured result object (`{ success, strategy, details }`) to the orchestrator, which updates its state machine to retry, proceed, or halt execution.

## Configuring Recovery Behavior

You enable and tune the **AIOX recovery system** through the `MasterOrchestrator` constructor:

```javascript
const { MasterOrchestrator } = require('.aiox-core/core/orchestration/master-orchestrator');

const orchestrator = new MasterOrchestrator('/path/to/project', {
  storyId: 'STORY-123',
  maxRetries: 4,          // Maximum attempts per epic
  autoRecovery: true,     // Enable automatic failure recovery
  autoEscalate: false,    // Require manual review before human escalation
  circularDetection: true // Enable circular pattern detection
});

```

For custom agents or manual intervention workflows, invoke the handler directly:

```javascript
async function recoverFromEpic(epicNum, err) {
  const result = await orchestrator.recoveryHandler.handleEpicFailure(
    epicNum,
    err,
    { approach: 'default', subtaskId: `epic-${epicNum}` }
  );
  console.log('Recovery result:', result);
}

```

Monitor recovery progress through event listeners:

```javascript
orchestrator.recoveryHandler.on('recoveryAttempt', ({ epicNum, attempt, strategy, result }) => {
  console.log(`[Recovery] Epic ${epicNum} – Attempt ${attempt} – Strategy: ${strategy}`);
  console.log('Result:', result);
});

```

## Implementing Custom Detection Logic

Override the default `StuckDetector` to implement domain-specific failure detection rules:

```javascript
// Replace the default detector with custom timeout logic
orchestrator.recoveryHandler._getStuckDetector = () => ({
  check: (attempts) => {
    const recent = attempts.slice(-3);
    const stuck = recent.every(a => /timeout/.test(a.error));
    return { stuck, reason: stuck ? 'custom_timeout_stuck' : null };
  }
});

```

This pattern allows projects to treat specific error signatures (such as three consecutive timeouts) as stuck conditions requiring immediate escalation or rollback.

## Summary

- The **AIOX recovery system** is implemented in [`.aiox-core/core/orchestration/recovery-handler.js`](https://github.com/SynkraAI/aiox-core/blob/main/.aiox-core/core/orchestration/recovery-handler.js) as the `RecoveryHandler` class, working alongside `StuckDetector`, `RollbackManager`, and `RecoveryTracker`.
- **Automatic failure recovery** follows a seven-step process: error capture, logging, stuck detection, strategy selection, execution, event emission, and result propagation.
- Five recovery strategies handle different failure modes: retry, rollback-and-retry, skip, escalate-to-human, and trigger recovery workflow (Epic 5).
- Configuration through `MasterOrchestrator` supports tuning via `maxRetries`, `autoEscalate`, and `circularDetection` parameters.
- The system emits `recoveryAttempt` events for observability and supports custom stuck-detection logic for specialized error patterns.

## Frequently Asked Questions

### What triggers the AIOX recovery system?

The system triggers when the `MasterOrchestrator` catches an error during epic execution and calls `RecoveryHandler.handleEpicFailure()`. This occurs automatically for any unhandled exception in the Spec, Execution, or QA epics when `autoRecovery` is enabled in the orchestrator configuration.

### How does the system detect circular failure patterns?

The `RecoveryHandler` consults the `StuckDetector` infrastructure module (loaded from `../../infrastructure/scripts/stuck-detector`) to analyze the `attempts` history. It identifies circular patterns when the same error type and approach repeat consecutively, or when failure counts exceed configured thresholds within a specific time window.

### Can I customize the recovery strategies?

Yes. While the five base strategies (`RETRY_SAME_APPROACH`, `ROLLBACK_AND_RETRY`, `SKIP_PHASE`, `ESCALATE_TO_HUMAN`, `TRIGGER_RECOVERY_WORKFLOW`) are defined in the `RecoveryStrategy` enumeration, you can override the `_selectRecoveryStrategy` method in a subclass of `RecoveryHandler`, or modify the stuck-detection logic via `_getStuckDetector` to influence which strategy gets selected for specific error conditions.

### What happens when all recovery attempts fail?

When the system exhausts `maxRetries` or encounters an unrecoverable error classification, it executes the `ESCALATE_TO_HUMAN` strategy. This generates a detailed escalation report containing the error history, attempted strategies, and debugging suggestions, then sets the orchestrator state to `BLOCKED`, halting automatic execution until manual intervention occurs.