How the ADE Recovery System Manages Execution Failures in SynkraAI/aios-core
The ADE Recovery System detects, classifies, and automatically resolves execution failures through the RecoveryHandler class, which selects from five strategies—retry, rollback and retry, skip, human escalation, or dedicated recovery workflows—while preventing infinite loops via circular-approach detection.
The Autonomous Development Engine (ADE) in the SynkraAI/aios-core repository orchestrates complex epic workflows (3 → 4 → 5 → 6 → 7) that can fail due to transient errors, configuration drift, or systemic issues. The ADE Recovery System provides a resilient safety net through the RecoveryHandler class located in .aios-core/core/orchestration/recovery-handler.js, which integrates tightly with the Master Orchestrator to manage failures without manual intervention.
Core Architecture of the ADE Recovery System
RecoveryHandler Class and Strategy Selection
The RecoveryHandler class serves as the central engine for failure management. It defines two critical enums: RecoveryStrategy and RecoveryResult.
The RecoveryStrategy enum (lines 33‑44 in recovery-handler.js) defines five possible actions:
RETRY_SAME_APPROACH– Re‑execute the failed epic with identical parameters.ROLLBACK_AND_RETRY– Revert state viaRollbackManagerbefore retrying.SKIP_PHASE– Bypass the current epic and proceed to the next.ESCALATE_TO_HUMAN– Block workflow and alert human operators.TRIGGER_REOVERY_WORKFLOW– Delegate to Epic 5, the dedicated recovery epic (as defined in source).
The RecoveryResult enum (lines 49‑54) reports execution outcomes: SUCCESS, FAILED, ESCALATED, or SKIPPED.
Error Classification and Stuck Detection
Before selecting a strategy, the handler classifies errors and detects circular failure patterns.
The _classifyError method (lines 20‑52) uses regex‑based heuristics to categorize errors into five types:
transient– Network timeouts, temporary resource unavailability.state– Corrupted or inconsistent internal state.configuration– Missing or invalid configuration values.dependency– External service or module failures.fatal– Unrecoverable errors requiring immediate escalation.
Simultaneously, the _checkIfStuck method invokes the StuckDetector (located in .aios-core/infrastructure/scripts/stuck-detector.js) to identify when the same approach fails repeatedly in a circular pattern. If stuck: true is returned, the handler avoids infinite loops by selecting ROLLBACK_AND_RETRY or ESCALATE_TO_HUMAN instead of simple retries.
Recovery Strategies and Execution Flow
The Five Recovery Strategies
The _selectRecoveryStrategy method (lines 42‑113) implements the decision engine. It considers:
- Retry count – Compared against
maxRetries(default configured in constructor, lines 63‑84). - Stuck status – From
StuckDetector. - Error classification – Fatal errors bypass retries.
- Epic criticality – Epic 5 (recovery epic) failures never trigger recovery to avoid infinite recursion.
- Auto‑escalation flag – When
autoEscalateis true and retries are exhausted.
The _executeRecoveryStrategy method (lines 84‑106) performs the concrete operation:
- Retry – Sets
shouldRetry: truewithout state changes. - Rollback and Retry – Invokes
_executeRollback, which delegates toRollbackManager(.aios-core/infrastructure/scripts/rollback-manager.js) to restore a known‑good checkpoint. - Skip – Marks the epic as
skippedand returnsSKIPPEDstatus. - Escalate – Sets
escalated: trueandESCALATEDstatus, blocking further automation. - Trigger Recovery Workflow – Invokes
_triggerRecoveryWorkflow, which delegates to Epic 5 via the orchestrator’sexecuteEpicmethod.
Integration with Master Orchestrator
The Master Orchestrator (.aios-core/core/orchestration/master-orchestrator.js) serves as the integration point. When an epic throws, the orchestrator calls _attemptRecovery (lines 94‑107), which guards against infinite loops (never recovers from Epic 5 failures) and forwards to RecoveryHandler.handleEpicFailure.
After the handler returns, the orchestrator processes the result (lines 117‑145):
- Updates
executionState.retryCount. - Transitions to
BLOCKEDstate if escalation occurs. - Retries the epic if
shouldRetryis true. - Proceeds to the next epic if
skippedis true.
The orchestrator also exposes getRecoveryHandler() (lines 50‑55) for external code or tests to retrieve the live handler instance.
Implementation Examples
Direct Use of RecoveryHandler
You can instantiate RecoveryHandler outside the orchestrator for custom scripts or testing:
// Example: manual recovery handling in a custom script
const { RecoveryHandler } = require('./.aios-core/core/orchestration/recovery-handler');
const handler = new RecoveryHandler({
projectRoot: process.cwd(),
storyId: 'demo-story',
maxRetries: 4,
autoEscalate: true,
});
async function runEpic(epicNum, epicFn) {
try {
await epicFn();
} catch (err) {
const result = await handler.handleEpicFailure(epicNum, err, {
approach: 'default',
});
console.log('Recovery result:', result);
if (result.shouldRetry) {
return runEpic(epicNum, epicFn); // simple retry loop
}
throw new Error('Unrecoverable failure');
}
}
The handler returns an object containing strategy, success, shouldRetry, and escalated, allowing the caller to decide whether to continue or abort.
Listening to Recovery Events
The handler emits a recoveryAttempt event after every failure processing cycle:
handler.on('recoveryAttempt', ({ epicNum, attempt, strategy, result }) => {
console.log(
`[Recovery] Epic ${epicNum} – attempt #${attempt} – strategy ${strategy} – success: ${result.success}`
);
});
This event stream integrates with monitoring tools, CI dashboards, or external alerting systems to provide real‑time visibility into ADE resilience.
Custom Strategy Extension
For specialized recovery logic, subclass RecoveryHandler to add custom strategies:
const { RecoveryHandler, RecoveryStrategy } = require('./.aios-core/core/orchestration/recovery-handler');
class MyRecoveryHandler extends RecoveryHandler {
_selectRecoveryStrategy(epicNum, error, stuckResult) {
const base = super._selectRecoveryStrategy(epicNum, error, stuckResult);
if (base === RecoveryStrategy.ROLLBACK_AND_RETRY && this._needsAlternativeToolchain(error)) {
return 'run_alternate_toolchain';
}
return base;
}
async _executeRecoveryStrategy(epicNum, strategy, error, context) {
if (strategy === 'run_alternate_toolchain') {
// custom logic here
return { success: true, shouldRetry: true, strategy };
}
return super._executeRecoveryStrategy(epicNum, strategy, error, context);
}
}
Wire the custom handler into the Master Orchestrator to override default behavior without modifying core library code.
Key Source Files and Components
| File | Role | Location |
|---|---|---|
recovery-handler.js |
Central recovery engine implementing strategy selection, attempt logging, and integration with StuckDetector and RollbackManager. | .aios-core/core/orchestration/recovery-handler.js |
master-orchestrator.js |
Orchestrates ADE epics and invokes RecoveryHandler on failures; manages retry counters and state transitions. |
.aios-core/core/orchestration/master-orchestrator.js |
index.js |
Re‑exports RecoveryHandler, RecoveryStrategy, and RecoveryResult for external consumption. |
.aios-core/core/orchestration/index.js |
stuck-detector.js |
Detects circular or repeatedly failing approaches to prevent infinite loops. | .aios-core/infrastructure/scripts/stuck-detector.js |
rollback-manager.js |
Reverts workspace state to known‑good checkpoints before retry operations. | .aios-core/infrastructure/scripts/rollback-manager.js |
recovery-tracker.js |
Persists recovery attempts for audit trails and post‑mortem debugging. | .aios-core/infrastructure/scripts/recovery-tracker.js |
Summary
The ADE Recovery System provides autonomous resilience for the SynkraAI/aios-core execution pipeline through these key mechanisms:
- Structured failure handling via the
RecoveryHandlerclass in.aios-core/core/orchestration/recovery-handler.js, which records rich metadata and classifies errors into transient, state, configuration, dependency, or fatal categories. - Stuck loop prevention using the
StuckDetectorto identify circular approach patterns before they exhaust resources. - Five distinct recovery strategies—retry, rollback and retry, skip phase, human escalation, and dedicated recovery workflow (Epic 5)—selected based on retry count, error classification, and epic criticality.
- Tight integration with the
MasterOrchestrator, which persists retry state and transitions workflows between active, blocked, or skipped states based on handler results. - Extensibility through event emissions (
recoveryAttempt) and subclassing support for custom recovery logic.
Frequently Asked Questions
How does the ADE Recovery System detect infinite retry loops?
The system delegates circular‑pattern detection to the StuckDetector class located in .aios-core/infrastructure/scripts/stuck-detector.js. The _checkIfStuck method analyzes recent failure attempts to identify when the same approach fails repeatedly in a circular pattern. If stuck: true is returned, the handler avoids infinite loops by selecting ROLLBACK_AND_RETRY or ESCALATE_TO_HUMAN instead of simple retries.
What happens when the recovery system encounters a fatal error?
Fatal errors are classified by the _classifyError method in recovery-handler.js using regex‑based heuristics that match unrecoverable conditions. When a fatal error is detected, the _selectRecoveryStrategy method bypasses retry logic and immediately returns ESCALATE_TO_HUMAN. The MasterOrchestrator then transitions the workflow to a BLOCKED state and halts further automation until human intervention resolves the underlying issue.
Can I extend the ADE Recovery System with custom recovery logic?
Yes, the RecoveryHandler class is designed for extension through standard JavaScript subclassing. You can override _selectRecoveryStrategy to inject custom decision logic that returns new strategy identifiers, then override _executeRecoveryStrategy to implement the corresponding behavior. The custom handler can be wired into the MasterOrchestrator in place of the default instance, allowing specialized recovery workflows—such as switching to alternative toolchains or invoking external remediation services—without modifying core library code.
How does the recovery system integrate with Epic 5?
Epic 5 serves as the dedicated "Recovery Epic" within the ADE workflow. When the _selectRecoveryStrategy method determines that a failure requires complex remediation—such as multi‑step state reconstruction or cross‑epic coordination—it returns TRIGGER_REOVERY_WORKFLOW. The _triggerRecoveryWorkflow method then delegates to Epic 5 via the MasterOrchestrator to execute the recovery epic with special context parameters containing the failure metadata. This design isolates complex recovery logic from standard epic execution while preventing recursive recovery attempts (the orchestrator explicitly blocks recovery for Epic 5 failures).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →