# How the OpenWork Engine Self-Heal Mechanism Recovers from Engine Failures

> Discover how the OpenWork engine self-heal mechanism recovers from failures by probing health endpoints, verifying identity, and patiently waiting for issues to resolve, ensuring session continuity.

- Repository: [Different AI/openwork](https://github.com/different-ai/openwork)
- Tags: internals
- Published: 2026-08-22

---

**The **engine self-heal mechanism** recovers from failures by continuously probing a dedicated health endpoint, verifying process identity to ensure session continuity, and gracefully waiting for transient issues to resolve without forcing process restarts.**

The `different-ai/openwork` repository implements a resilient local-first architecture where the AI engine survives network outages and resource contention through a sophisticated **self-healing protocol**. This mechanism ensures that intermittent Den connectivity issues never disrupt active workflows or require user re-authentication. By combining lightweight health checks in [`dev/scripts/dev-headless-web.ts`](https://github.com/different-ai/openwork/blob/main/dev/scripts/dev-headless-web.ts) with identity-preserving recovery logic in [`dev/helpers.ts`](https://github.com/different-ai/openwork/blob/main/dev/helpers.ts), OpenWork maintains engine availability while preserving in-memory state.

## Health Endpoint Architecture

The foundation of the **engine self-heal mechanism** rests on continuous monitoring that separates engine liveness from application functionality.

### The Engine Health Endpoint

At startup, the engine exposes a lightweight HTTP health endpoint. The client accesses this through `manifest.healthUrl` and validates responsiveness using the `probeOk` function defined in [`dev/scripts/dev-headless-web.ts`](https://github.com/different-ai/openwork/blob/main/dev/scripts/dev-headless-web.ts).

```typescript
// dev/scripts/dev-headless-web.ts
const healthOk = await probeOk(manifest.healthUrl);
if (!healthOk) {
  // Trigger recovery path rather than failing the operation
  await handleEngineRecovery();
}

```

The `probeOk` utility performs an HTTP request to the engine's health endpoint, returning `true` only when the engine responds with a success status code. This indicates the internal JavaScript runtime is responsive and ready to process AI operations.

### Lifecycle State Monitoring

Beyond simple HTTP availability, the mechanism examines the engine's internal `lifecycleState` property to distinguish between transient unavailability and terminal failure. The `waitForEngine` utility in [`dev/helpers.ts`](https://github.com/different-ai/openwork/blob/main/dev/helpers.ts) polls this state until the engine reports `"healthy"`.

```typescript
// dev/helpers.ts
export async function waitForEngine(
  app: App, 
  description: string, 
  previousEngine?: EngineInfo
) {
  const deadline = Date.now() + 30000; // 30 second timeout
  while (Date.now() < deadline) {
    const info = await app.engineDiagnostics();
    if (info.lifecycleState === "healthy") {
      return info;
    }
    await delay(500);
  }
  throw new Error(`Engine did not become healthy: ${description}`);
}

```

## Process Identity Verification

A critical aspect of the **engine self-heal mechanism** is ensuring that recovery represents true process continuity rather than a masked restart that might lose session state.

### Engine ID Persistence

When `waitForEngine` receives a `previousEngine` parameter, it validates that the recovered engine maintains the same identifier as the pre-failure instance. This prevents the client from accepting a fresh engine process as a "healed" one when the original actually crashed.

```typescript
// dev/helpers.ts – identity verification within waitForEngine
if (previousEngine && info.id !== previousEngine.id) {
  // Engine restarted – treat as new instance rather than healed
  throw new Error("Engine restarted during recovery");
}

```

This verification appears in the end-to-end test specification [`dev/evals/specs/den-intermittent-outage-recovery.e2e.test.ts`](https://github.com/different-ai/openwork/blob/main/dev/evals/specs/den-intermittent-outage-recovery.e2e.test.ts), which asserts that the engine ID remains constant across simulated outages.

## Graceful Recovery Flow

The complete recovery sequence demonstrates how OpenWork handles failures without user interruption.

### Detection and Waiting

When `probeOk` detects a failure, the client invokes `waitForEngine` with the previous engine context. This blocks subsequent operations until the engine returns to `"healthy"` status or the 30-second timeout expires.

### Identity-Preserving Recovery

Unlike traditional watchdog mechanisms that aggressively restart processes, the **engine self-heal mechanism** **waits for the existing process to recover**. This preserves in-memory state, active model contexts, and user authentication tokens. Only if the engine ID changes—indicating the original process died—does the launcher initiate a fresh startup sequence.

The e2e test validates this through multiple outage scenarios:

```typescript
// dev/evals/specs/den-intermittent-outage-recovery.e2e.test.ts
const baselineEngine = await waitForEngine(
  desktopApp, 
  "baseline healthy engine identity"
);
expect(baselineEngine.lifecycleState).toBe("healthy");

// After simulated Den outage, verify same engine instance
const outageAEngine = await waitForEngine(
  desktopApp, 
  "outage A unchanged healthy engine", 
  baselineEngine
);
expect(outageAEngine.lifecycleState).toBe("healthy");
expect(outageAEngine.id).toBe(baselineEngine.id); // Process continuity

```

## Summary

- **Health Monitoring**: The `probeOk` function in [`dev/scripts/dev-headless-web.ts`](https://github.com/different-ai/openwork/blob/main/dev/scripts/dev-headless-web.ts) polls `manifest.healthUrl` to detect engine unavailability without terminating connections.
- **Identity Verification**: The `waitForEngine` helper in [`dev/helpers.ts`](https://github.com/different-ai/openwork/blob/main/dev/helpers.ts) validates `info.id` against `previousEngine.id` to ensure true process continuity.
- **Session Preservation**: Rather than restarting, the mechanism waits for the existing process to return to `"healthy"` lifecycle state, maintaining authentication and model context.
- **Validated Resilience**: The [`den-intermittent-outage-recovery.e2e.test.ts`](https://github.com/different-ai/openwork/blob/main/den-intermittent-outage-recovery.e2e.test.ts) specification proves the engine survives Den outages without identity changes or state loss.

## Frequently Asked Questions

### How does OpenWork detect engine failures without restarting the process?

The client uses `probeOk` to check `manifest.healthUrl` before operations. When this fails, the system invokes `waitForEngine` to poll `engineDiagnostics()` until `lifecycleState` returns to `"healthy"`. Crucially, if the returned `info.id` matches `previousEngine.id`, the system confirms the original process recovered rather than being replaced.

### What happens to user sessions when the engine encounters a transient failure?

User sessions remain intact because the **engine self-heal mechanism** preserves the existing process instead of restarting it. Since `waitForEngine` validates process identity through ID comparison, in-memory session state, authentication tokens, and model contexts survive the outage without requiring re-authentication.

### Where is the engine health checking logic implemented in the OpenWork codebase?

The health probing logic resides in [`dev/scripts/dev-headless-web.ts`](https://github.com/different-ai/openwork/blob/main/dev/scripts/dev-headless-web.ts) through the `probeOk` function, while the recovery orchestration lives in [`dev/helpers.ts`](https://github.com/different-ai/openwork/blob/main/dev/helpers.ts) within the `waitForEngine` function. End-to-end validation appears in [`dev/evals/specs/den-intermittent-outage-recovery.e2e.test.ts`](https://github.com/different-ai/openwork/blob/main/dev/evals/specs/den-intermittent-outage-recovery.e2e.test.ts), which simulates network outages to verify self-healing behavior.

### Does the engine self-heal mechanism handle complete engine crashes?

Yes, though with different semantics. If `waitForEngine` detects that `info.id` differs from `previousEngine.id`, it confirms the original process died. In this scenario, the launcher initiates a fresh engine process before allowing operations to continue, ensuring availability while signaling that session state has been reset.