# How OpenWork's EnginePool Manages, Recycles, and Self-Heals Engine Processes

> Discover how OpenWork's EnginePool uses blue-green rollover to manage engine processes, safely recycle busy engines, and self-heal from crashes in seconds. Learn more.

- Repository: [Different AI/openwork](https://github.com/different-ai/openwork)
- Tags: internals
- Published: 2026-08-22

---

**OpenWork's EnginePool implements a blue-green rollover strategy in [`apps/server/src/engine-pool.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/engine-pool.ts) that safely drains busy engines during configuration changes and automatically recovers from primary engine crashes within seconds.**

The `EnginePool` class in the `different-ai/openwork` repository orchestrates the lifecycle of OpenCode engine processes, ensuring zero-downtime configuration updates and resilient operation against unexpected failures. By treating each engine instance as a **generation** with distinct states—`starting`, `primary`, `draining`, and `dead`—the pool maintains continuous availability while managing finite system resources.

## EnginePool Architecture and Core Concepts

Each engine managed by the pool represents a **generation**, tracked through metadata that includes its current status and active session ownership. The `snapshot()` method provides a concise diagnostic view of all live generations, enabling operators to inspect the current state of the pool at runtime.

Connection routing logic in `routeRequest()` determines whether incoming requests should hit the primary generation, be forwarded to a draining generation based on `sessionOwnership`, or trigger fallback behavior. This ensures that long-running sessions complete gracefully even as the pool transitions between engine versions.

## Recycling: The Drain-and-Replace Strategy

When configuration changes arrive while the current engine is busy, the pool initiates a graceful rollover rather than forcibly terminating active work.

### Spawning the Standby Generation

The `spawn()` method calls `createManagedOpencodeServer` to instantiate a new engine process with updated configuration settings. To prevent resource exhaustion during rapid configuration changes, the pool **coalesces concurrent reload requests** using `inFlight` and `pendingRollover` flags, ensuring only one standby exists at any time.

If multiple reload requests arrive simultaneously, subsequent calls wait for the in-flight rollover rather than spawning duplicate processes.

### Flip and Drain Workflow

Once the standby is healthy, the `flip()` method (lines 845-879) promotes the new generation to primary while demoting the existing engine to a `draining` state. During this phase:

- The `activeSessionsByGeneration` map tracks which sessions belong to the old generation
- `sessionOwnership` logic ensures requests for existing sessions route to the draining engine via `routeRequest()` (lines 317-352)
- New sessions exclusively target the primary generation

This approach allows zero-downtime deployments where existing connections finish naturally while new traffic flows to the updated engine.

### Graceful Termination Monitoring

The `startDrainMonitor()` method initiates a periodic timer checking `nonIdleSessionIds()` (lines 1002-1035) to determine when all sessions have completed. If sessions exceed the grace period defined by `drainTimeoutMs()`, the pool forcibly calls `abortSession()` on remaining connections and invokes `retire()` to terminate the drained process.

## Self-Healing: Recovering from Primary Engine Death

When the primary engine crashes or becomes unreachable, the `EnginePool` detects the failure and orchestrates automatic recovery without manual intervention.

### Failure Detection

The `reportRequestFailure()` method tracks connection errors across requests. After **three consecutive connection failures** (lines 248-256), the pool marks the primary as dead and triggers the recovery sequence.

### Recovery Execution

The `scheduleRecovery()` method (lines 1329-1347) enforces minimum spawn intervals to prevent thrashing, then executes `runDeadPrimaryRecovery()`:

1. **Cleanup**: Closes orphaned processes and removes the dead registration
2. **Spawn**: Calls `spawn()` to create a fresh engine instance
3. **Health Check**: Waits for the engine to become ready via `waitForHealthy()`
4. **Promotion**: Executes `flip()` to promote the new generation to primary

If the new engine fails to start, the pool keeps the orphan for a later retry and logs the error via `logger.error` (lines 1384-1405), ensuring the system remains in a recoverable state.

## Practical Implementation Examples

Create and configure an engine pool at server startup:

```typescript
import { EnginePool, createEnginePoolForConfig } from "./engine-pool.js";

const pool = createEnginePoolForConfig({
  config,
  template,
  hooks,
});

```

Request a graceful rollover after configuration changes:

```typescript
await pool.requestRollover({
  reason: "user-updated-permissions",
  workspace,
  manual: false,          // skip if fingerprint unchanged
  forceStandby: false,    // reload in place if idle
});

```

Force a standby rollout for critical updates:

```typescript
await pool.requestRollover({
  reason: "critical-security-patch",
  workspace,
  forceStandby: true,     // always roll over via standby
});

```

Inspect current pool state for debugging:

```typescript
console.log(pool.snapshot());

```

## Summary

- **Blue-green rollover**: The `EnginePool` maintains at most two generations (primary and standby) to enable zero-downtime configuration updates
- **Session-aware draining**: Active sessions complete on the old generation while new sessions route to the primary via `sessionOwnership` tracking
- **Coalesced reloads**: Concurrent `requestRollover()` calls merge into a single spawn operation to prevent resource exhaustion
- **Automatic recovery**: Three consecutive connection failures trigger `runDeadPrimaryRecovery()`, which spawns and promotes a new primary engine
- **Observability**: The `snapshot()` method and `EnginePoolLogger` integration provide visibility into generation states and lifecycle events

## Frequently Asked Questions

### How does EnginePool handle concurrent configuration reload requests?

The pool coalesces concurrent reload attempts using `inFlight` and `pendingRollover` state flags. If a rollover is already in progress when a new request arrives, the subsequent call waits for the existing operation rather than spawning additional standby processes, ensuring only one engine generation transitions at a time.

### What happens to active sessions when an engine is marked for draining?

Sessions owned by the draining generation continue running on that instance until completion. The `routeRequest()` method checks `activeSessionsByGeneration` to forward existing session requests to the draining engine, while new sessions route exclusively to the primary. If sessions exceed `drainTimeoutMs()`, the pool forcibly aborts them via `abortSession()`.

### How does the pool detect that the primary engine has died?

The `reportRequestFailure()` method tracks connection errors per request. After accumulating three consecutive connection failures (as implemented in lines 248-256 of [`engine-pool.ts`](https://github.com/different-ai/openwork/blob/main/engine-pool.ts)), the pool classifies the primary as dead and schedules recovery through `scheduleRecovery()`.

### Can the drain timeout be configured, and what happens when it expires?

Yes, the `drainTimeoutMs()` method defines the grace period for draining. When this timeout expires, the `startDrainMonitor()` routine forcibly terminates remaining sessions using `abortSession()` and immediately retires the engine via `retire()`, freeing system resources regardless of session completion status.