How OpenWork's EnginePool Manages, Recycles, and Self-Heals Engine Processes

OpenWork's EnginePool implements a blue-green rollover strategy in apps/server/src/engine-pool.ts that safely drains busy engines during configuration changes and automatically recovers from primary engine crashes within seconds.

The EnginePool class in the different-ai/openwork repository orchestrates the lifecycle of OpenCode engine processes, ensuring zero-downtime configuration updates and resilient operation against unexpected failures. By treating each engine instance as a generation with distinct states—starting, primary, draining, and dead—the pool maintains continuous availability while managing finite system resources.

EnginePool Architecture and Core Concepts

Each engine managed by the pool represents a generation, tracked through metadata that includes its current status and active session ownership. The snapshot() method provides a concise diagnostic view of all live generations, enabling operators to inspect the current state of the pool at runtime.

Connection routing logic in routeRequest() determines whether incoming requests should hit the primary generation, be forwarded to a draining generation based on sessionOwnership, or trigger fallback behavior. This ensures that long-running sessions complete gracefully even as the pool transitions between engine versions.

Recycling: The Drain-and-Replace Strategy

When configuration changes arrive while the current engine is busy, the pool initiates a graceful rollover rather than forcibly terminating active work.

Spawning the Standby Generation

The spawn() method calls createManagedOpencodeServer to instantiate a new engine process with updated configuration settings. To prevent resource exhaustion during rapid configuration changes, the pool coalesces concurrent reload requests using inFlight and pendingRollover flags, ensuring only one standby exists at any time.

If multiple reload requests arrive simultaneously, subsequent calls wait for the in-flight rollover rather than spawning duplicate processes.

Flip and Drain Workflow

Once the standby is healthy, the flip() method (lines 845-879) promotes the new generation to primary while demoting the existing engine to a draining state. During this phase:

  • The activeSessionsByGeneration map tracks which sessions belong to the old generation
  • sessionOwnership logic ensures requests for existing sessions route to the draining engine via routeRequest() (lines 317-352)
  • New sessions exclusively target the primary generation

This approach allows zero-downtime deployments where existing connections finish naturally while new traffic flows to the updated engine.

Graceful Termination Monitoring

The startDrainMonitor() method initiates a periodic timer checking nonIdleSessionIds() (lines 1002-1035) to determine when all sessions have completed. If sessions exceed the grace period defined by drainTimeoutMs(), the pool forcibly calls abortSession() on remaining connections and invokes retire() to terminate the drained process.

Self-Healing: Recovering from Primary Engine Death

When the primary engine crashes or becomes unreachable, the EnginePool detects the failure and orchestrates automatic recovery without manual intervention.

Failure Detection

The reportRequestFailure() method tracks connection errors across requests. After three consecutive connection failures (lines 248-256), the pool marks the primary as dead and triggers the recovery sequence.

Recovery Execution

The scheduleRecovery() method (lines 1329-1347) enforces minimum spawn intervals to prevent thrashing, then executes runDeadPrimaryRecovery():

  1. Cleanup: Closes orphaned processes and removes the dead registration
  2. Spawn: Calls spawn() to create a fresh engine instance
  3. Health Check: Waits for the engine to become ready via waitForHealthy()
  4. Promotion: Executes flip() to promote the new generation to primary

If the new engine fails to start, the pool keeps the orphan for a later retry and logs the error via logger.error (lines 1384-1405), ensuring the system remains in a recoverable state.

Practical Implementation Examples

Create and configure an engine pool at server startup:

import { EnginePool, createEnginePoolForConfig } from "./engine-pool.js";

const pool = createEnginePoolForConfig({
  config,
  template,
  hooks,
});

Request a graceful rollover after configuration changes:

await pool.requestRollover({
  reason: "user-updated-permissions",
  workspace,
  manual: false,          // skip if fingerprint unchanged
  forceStandby: false,    // reload in place if idle
});

Force a standby rollout for critical updates:

await pool.requestRollover({
  reason: "critical-security-patch",
  workspace,
  forceStandby: true,     // always roll over via standby
});

Inspect current pool state for debugging:

console.log(pool.snapshot());

Summary

  • Blue-green rollover: The EnginePool maintains at most two generations (primary and standby) to enable zero-downtime configuration updates
  • Session-aware draining: Active sessions complete on the old generation while new sessions route to the primary via sessionOwnership tracking
  • Coalesced reloads: Concurrent requestRollover() calls merge into a single spawn operation to prevent resource exhaustion
  • Automatic recovery: Three consecutive connection failures trigger runDeadPrimaryRecovery(), which spawns and promotes a new primary engine
  • Observability: The snapshot() method and EnginePoolLogger integration provide visibility into generation states and lifecycle events

Frequently Asked Questions

How does EnginePool handle concurrent configuration reload requests?

The pool coalesces concurrent reload attempts using inFlight and pendingRollover state flags. If a rollover is already in progress when a new request arrives, the subsequent call waits for the existing operation rather than spawning additional standby processes, ensuring only one engine generation transitions at a time.

What happens to active sessions when an engine is marked for draining?

Sessions owned by the draining generation continue running on that instance until completion. The routeRequest() method checks activeSessionsByGeneration to forward existing session requests to the draining engine, while new sessions route exclusively to the primary. If sessions exceed drainTimeoutMs(), the pool forcibly aborts them via abortSession().

How does the pool detect that the primary engine has died?

The reportRequestFailure() method tracks connection errors per request. After accumulating three consecutive connection failures (as implemented in lines 248-256 of engine-pool.ts), the pool classifies the primary as dead and schedules recovery through scheduleRecovery().

Can the drain timeout be configured, and what happens when it expires?

Yes, the drainTimeoutMs() method defines the grace period for draining. When this timeout expires, the startDrainMonitor() routine forcibly terminates remaining sessions using abortSession() and immediately retires the engine via retire(), freeing system resources regardless of session completion status.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →